How to Build and Use the Python Bindings for pdf-inspector with PyO3

Enable the python feature in Cargo.toml and use cargo build --features python or maturin develop to compile a native Python extension module that exposes pdf-inspector's Rust API to CPython 3.8+.

The pdf-inspector repository provides optional Python bindings powered by PyO3, a Rust crate for creating Python extension modules. These bindings let you call the library's high-performance PDF processing functions directly from Python code without leaving the interpreter. This guide covers the complete build process, installation methods, and practical usage examples based on the actual source implementation in firecrawl/pdf-inspector.

Prerequisites and Feature Configuration

The Python bindings are gated behind an optional Cargo feature to keep the core library lightweight. Before building, verify your environment meets these requirements:

  • Rust toolchain (stable channel, compatible with edition 2021)
  • Python 3.8 or later (the bindings use PyO3's abi3-py38 flag for broad compatibility)
  • pip and virtual environment (recommended for isolation)

In Cargo.toml, the python feature is declared with PyO3 configured for extension-module builds:

[dependencies]
pyo3 = { version = "0.25", features = ["extension-module", "abi3-py38"], optional = true }

[features]
python = ["pyo3"]  # Cargo.toml line 64

The extension-module flag tells PyO3 to build a shared library suitable for Python import, while abi3-py38 ensures the resulting wheel works across CPython 3.8–3.12 without recompilation (source: Cargo.toml lines 30–64).

Building the Python Bindings

You have two primary paths to build and install the extension: direct Cargo compilation or Python-native tooling with Maturin.

Method 1: Cargo Build (Manual)

Compile the shared library directly with Cargo:

cargo build --release --features python

This produces a platform-specific shared library in target/release/:

  • Linux/macOS: libpdf_inspector.so or pdf_inspector.so (renamed for import)
  • Windows: pdf_inspector.pyd

To use the compiled library, either copy it to your working directory or add target/release to PYTHONPATH.

Maturin streamlines the build-and-install workflow for Python developers:


# Install maturin if needed

pip install maturin

# Build and install into active virtual environment

maturin develop --release --features python

This single command compiles with optimizations enabled and registers the module for immediate import pdf_inspector usage. For distribution, maturin build --features python creates manylinux-compliant wheels.

Method 3: setuptools-rust

For projects integrating pdf-inspector into a larger Python package, add to pyproject.toml:

[build-system]
requires = ["setuptools", "setuptools-rust"]

[tool.setuptools-rust]
rust-extensions = [
    { path = "Cargo.toml", binding = "PyO3", args = ["--features", "python"] }
]

Then install with pip install . as with any source package.

Module Structure and Public API

The Python module is implemented in src/python.rs, which exports Rust structs as Python classes and functions via #[pyclass] and #[pyfunction] macros. The key types and their Rust source locations:

Python Name Rust Struct Module Export Description
PdfResult PyPdfResult src/python.rs lines 15–54 Complete processing output with markdown, metadata, and layout
PdfClassification PyPdfClassification src/python.rs lines 92–101 Lightweight PDF type detection result
TextItem PyTextItem src/python.rs line 165+ Text with bounding box and font information

All classes expose fields as read-only properties using #[pyo3(get)], making them introspectable like dataclasses without setters.

Core Functions for PDF Processing

The module exposes eight main functions in src/python.rs (lines 440–520+). Each has both file-path and in-memory bytes variants:

Full Processing Functions

process_pdf(path: str, pages: Optional[List[int]] = None) -> PdfResult (source) : Executes the complete pipeline: PDF parsing, text extraction, OCR analysis, layout detection, and markdown generation. The optional pages parameter limits processing to specific 1-indexed page numbers.

process_pdf_bytes(data: bytes, pages: Optional[List[int]] = None) -> PdfResult (source) : Identical functionality for PDFs already loaded in memory.

Detection and Classification

detect_pdf(path: str) -> PdfResult (source) : Runs only the detection pipeline, returning a PdfResult with pdf_type populated but without full content extraction. Faster than process_pdf when you only need classification.

detect_pdf_bytes(data: bytes) -> PdfResult (source) : In-memory variant of detect_pdf.

classify_pdf(path: str) -> PdfClassification (source) : Returns just the PdfClassification with pdf_type (values like "text_based", " scanned", "hybrid"). Fastest option for type checking.

classify_pdf_bytes(data: bytes) -> PdfClassification : In-memory variant of classify_pdf.

Plain Text Extraction

extract_text(path: str) -> str : Returns raw extracted text without markdown formatting or metadata.

extract_text_bytes(data: bytes) -> str : In-memory variant.

extract_text_with_positions(path: str, pages: Optional[List[int]] = None) -> List[TextItem] : Returns TextItem objects containing text, x0, y0, x1, y1, font_name, and font_size for precise layout reconstruction.

extract_pages_markdown(path: str) -> List[str] : Returns a list of markdown strings, one per page, preserving the document's structure hierarchy.

All functions raise PyValueError on invalid input, corrupted PDFs, or processing failures (see pyo3::exceptions::PyValueError usage throughout src/python.rs).

Practical Usage Examples

Installation Verification

import pdf_inspector

print(f"pdf-inspector Python bindings loaded: {pdf_inspector.__file__}")

Fast PDF Classification

import pdf_inspector

# Determine PDF type without expensive OCR

classification = pdf_inspector.classify_pdf("document.pdf")
print(f"Type: {classification.pdf_type}")  # "text_based", "scanned", etc.

print(f"Confidence: {classification.confidence:.2f}")

Full Document Processing

import pdf_inspector

# Process all pages with complete metadata

result = pdf_inspector.process_pdf("report.pdf")

print(f"Pages: {result.page_count}")
print(f"Detected type: {result.pdf_type}")
print(f"Has text layer: {result.has_text}")
print(f"OCR required: {result.requires_ocr}")

# Access generated markdown

print(result.markdown[:500])  # First 500 characters

# Per-page markdown (if structure preservation needed)

pages_md = pdf_inspector.extract_pages_markdown("report.pdf")
for i, page_md in enumerate(pages_md, 1):
    print(f"--- Page {i} ---")
    print(page_md[:200])

In-Memory Processing (Web Applications)

import pdf_inspector
from pathlib import Path

# Load PDF from upload, database, or network

pdf_bytes = Path("uploaded.pdf").read_bytes()

# Process without writing to filesystem

result = pdf_inspector.process_pdf_bytes(pdf_bytes, pages=[1, 2, 3])
classification = pdf_inspector.classify_pdf_bytes(pdf_bytes)

Layout-Aware Text Extraction

import pdf_inspector

# Extract text with precise positioning for custom rendering

items = pdf_inspector.extract_text_with_positions("layout.pdf", pages=[1])

for item in items[:10]:
    print(f"'{item.text[:30]}...' "
          f"at ({item.x0:.1f}, {item.y0:.1f}) "
          f"font={item.font_name} size={item.font_size}")

Error Handling

import pdf_inspector
from pyo3 import exceptions

try:
    result = pdf_inspector.process_pdf("corrupted.pdf")
except Exception as e:  # Catches PyValueError

    print(f"Processing failed: {e}")

Key Source Files Reference

File Purpose Link
src/python.rs PyO3 bindings implementation: class definitions, function exports, error handling src/python.rs
Cargo.toml Feature flags, PyO3 dependency configuration, build profiles Cargo.toml
docs/python.md Extended Python API documentation and additional examples docs/python.md

Summary

  • Enable the python feature in Cargo.toml to activate PyO3 bindings compilation
  • Use maturin develop --features python for the smoothest development workflow, or cargo build --features python for manual builds
  • Import pdf_inspector after installation; all functions are immediately available
  • Choose the right function for your latency needs: classify_pdf for fastest type detection, process_pdf for complete analysis, extract_text_with_positions for layout control
  • Handle PyValueError exceptions for robust production code

The PyO3-based bindings provide zero-copy efficiency where possible and ABI3 compatibility across Python 3.8–3.12, making pdf-inspector suitable for high-throughput document processing pipelines.

Frequently Asked Questions

What Python versions are supported by pdf-inspector's PyO3 bindings?

The bindings target CPython 3.8 and later through PyO3's abi3-py38 flag. This produces a stable ABI that works across 3.8, 3.9, 3.10, 3.11, and 3.12 without recompilation. Alternative Python implementations (PyPy, GraalPython) are not officially supported as of the current Cargo.toml configuration.

Can I use pdf-inspector in Python without building from source?

Pre-built wheels are not currently distributed on PyPI for firecrawl/pdf-inspector. You must build from source using the methods described above. The Maturin toolchain (pip install maturin; maturin develop --features python) handles compilation automatically with no manual Rust configuration required for most platforms.

How do I pass only specific pages to process_pdf?

Use the optional pages parameter with a list of 1-indexed integers: pdf_inspector.process_pdf("doc.pdf", pages=[1, 3, 5]). This limits both processing time and memory usage for large documents. The same parameter works for process_pdf_bytes and extract_text_with_positions.

What is the difference between detect_pdf and process_pdf?

detect_pdf runs only the classification pipeline to determine PDF type (text-based, scanned, hybrid) without extracting content or running OCR—optimal for routing decisions. process_pdf executes the full pipeline including markdown generation, text extraction, OCR metadata, and layout analysis. For large batch operations, use classify_pdf first to filter documents before expensive full processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →