Python Libraries for PDF Manipulation in the Claude Skills Repository

The ComposioHQ/awesome-claude-skills repository relies on two primary Python libraries for PDF manipulation: pypdf for reading, writing, and editing document structure—including form fields and annotations—and pdf2image for converting PDF pages into PNG raster images.

The ComposioHQ/awesome-claude-skills repository provides document processing capabilities through dedicated PDF manipulation scripts located in the document-skills/pdf/scripts/ directory. Understanding which Python libraries for PDF manipulation power these workflows helps developers extend existing functionality or integrate similar patterns into their own applications. This guide examines the specific implementation details and source code patterns used throughout the repository.

Core Python Libraries for PDF Processing

The repository implements PDF workflows using two specialized libraries that handle distinct aspects of document manipulation: structural editing and raster conversion.

pypdf for Document Structure and Form Editing

pypdf serves as the primary library for reading, writing, and manipulating PDF document structure. According to the source code in document-skills/pdf/scripts/, the implementation leverages the PdfReader and PdfWriter classes to handle form field operations, annotations, and page assembly. This pure-Python toolkit enables the repository to extract field metadata, verify fillable fields, and programmatically update PDF values without requiring additional compiled dependencies.

pdf2image for Page Rasterization

pdf2image provides the bridge between PDF vector content and raster image processing. This library wraps the Poppler utilities to convert PDF pages into PIL Image objects through the convert_from_path function. The repository utilizes this capability in document-skills/pdf/scripts/convert_pdf_to_images.py to render PDF pages as PNG images for downstream processing workflows.

Working with PDF Forms Using pypdf

The repository demonstrates sophisticated form manipulation patterns using pypdf. In document-skills/pdf/scripts/fill_fillable_fields.py, the implementation reads existing PDF structures and updates form field values programmatically by accessing the page.annotations dictionary and updating the /V key for each target field.

from pypdf import PdfReader, PdfWriter

def fill_pdf_fields(input_pdf, fields, output_pdf):
    reader = PdfReader(input_pdf)
    writer = PdfWriter()

    for page_num, page in enumerate(reader.pages):
        # Update fields on the current page

        for field_name, value in fields.get(page_num + 1, {}).items():
            if field_name in page.annotations:
                page.annotations[field_name].update({"/V": value})

        writer.add_page(page)

    with open(output_pdf, "wb") as f:
        writer.write(f)

# Example usage

fields = {1: {"Name": "Alice", "Age": "30"}}
fill_pdf_fields("form.pdf", fields, "filled_form.pdf")

Additional form-related functionality appears in document-skills/pdf/scripts/fill_pdf_form_with_annotations.py, which adds free-text annotations, and document-skills/pdf/scripts/extract_form_field_info.py, which extracts metadata about available fields.

Converting PDFs to Images with pdf2image

For workflows requiring image-based processing, document-skills/pdf/scripts/convert_pdf_to_images.py implements the pdf2image conversion pipeline. The convert_from_path function accepts a dpi parameter to control output resolution, returning a list of PIL Image objects that the script processes and saves as PNG files.

from pdf2image import convert_from_path
import os

def pdf_to_png(pdf_path, out_dir, max_dim=1000):
    images = convert_from_path(pdf_path, dpi=200)

    for i, img in enumerate(images, start=1):
        # Resize if the image is larger than max_dim

        if max(img.size) > max_dim:
            scale = max_dim / max(img.size)
            img = img.resize((int(img.width * scale), int(img.height * scale)))

        img_path = os.path.join(out_dir, f"page_{i}.png")
        img.save(img_path)
        print(f"Saved {img_path}")

# Example usage

pdf_to_png("document.pdf", "output_images")

This implementation includes optional scaling logic to ensure output images do not exceed specified maximum dimensions, processing each page returned by convert_from_path individually.

Key Implementation Files in the Repository

The PDF manipulation logic is distributed across five key files in the document-skills/pdf/scripts/ directory:

These files collectively demonstrate how the repository utilizes pypdf for PDF structure manipulation and pdf2image for rasterizing pages. No other third-party PDF libraries such as PyMuPDF, pdfminer, or ReportLab appear in the codebase.

Summary

  • The Claude Skills repository uses pypdf and pdf2image as its exclusive Python libraries for PDF manipulation, handling structural editing and image conversion respectively.
  • pypdf handles form operations through PdfReader and PdfWriter classes, enabling field extraction, validation, and value updates in files like fill_fillable_fields.py.
  • pdf2image requires Poppler external dependencies and converts PDF pages to PIL Images via convert_from_path as shown in convert_pdf_to_images.py.
  • All PDF scripts reside in document-skills/pdf/scripts/, with clear separation between form manipulation and image conversion workflows.
  • No alternative PDF libraries are implemented in the current codebase.

Frequently Asked Questions

What Python libraries does the Claude Skills repository use for PDF manipulation?

The repository uses pypdf for reading, writing, and manipulating PDF document structure—including form fields and annotations—and pdf2image for converting PDF pages into raster images. These libraries handle all PDF-related operations in the document-skills/pdf/scripts/ directory.

How does pypdf handle form field filling in the Claude Skills codebase?

The implementation in fill_fillable_fields.py uses PdfReader to load the source document, then iterates through reader.pages to access the page.annotations dictionary. It updates field values by modifying the /V key for each annotation, then writes the modified pages to a new PdfWriter instance before saving to disk.

Can I use pdf2image without installing external dependencies?

No. According to the source code analysis, pdf2image functions as a thin wrapper around the Poppler utilities, which must be installed separately on the host system. The convert_from_path function relies on these external binaries to render PDF pages as PIL Image objects.

Are there alternatives to pypdf used in the repository for PDF manipulation?

No. The ComposioHQ/awesome-claude-skills repository exclusively uses pypdf for structural PDF operations. Other common libraries like PyMuPDF, pdfminer.six, or ReportLab do not appear in any of the PDF manipulation scripts within the codebase.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →