How to Build and Use the Python Bindings for pdf-inspector with PyO3
Enable the python feature in Cargo.toml and use cargo build --features python or maturin develop to compile a native Python extension module that exposes pdf-inspector's Rust API to CPython 3.8+.
The pdf-inspector repository provides optional Python bindings powered by PyO3, a Rust crate for creating Python extension modules. These bindings let you call the library's high-performance PDF processing functions directly from Python code without leaving the interpreter. This guide covers the complete build process, installation methods, and practical usage examples based on the actual source implementation in firecrawl/pdf-inspector.
Prerequisites and Feature Configuration
The Python bindings are gated behind an optional Cargo feature to keep the core library lightweight. Before building, verify your environment meets these requirements:
- Rust toolchain (stable channel, compatible with edition 2021)
- Python 3.8 or later (the bindings use PyO3's
abi3-py38flag for broad compatibility) - pip and virtual environment (recommended for isolation)
In Cargo.toml, the python feature is declared with PyO3 configured for extension-module builds:
[dependencies]
pyo3 = { version = "0.25", features = ["extension-module", "abi3-py38"], optional = true }
[features]
python = ["pyo3"] # Cargo.toml line 64
The extension-module flag tells PyO3 to build a shared library suitable for Python import, while abi3-py38 ensures the resulting wheel works across CPython 3.8–3.12 without recompilation (source: Cargo.toml lines 30–64).
Building the Python Bindings
You have two primary paths to build and install the extension: direct Cargo compilation or Python-native tooling with Maturin.
Method 1: Cargo Build (Manual)
Compile the shared library directly with Cargo:
cargo build --release --features python
This produces a platform-specific shared library in target/release/:
- Linux/macOS:
libpdf_inspector.soorpdf_inspector.so(renamed for import) - Windows:
pdf_inspector.pyd
To use the compiled library, either copy it to your working directory or add target/release to PYTHONPATH.
Method 2: Maturin (Recommended)
Maturin streamlines the build-and-install workflow for Python developers:
# Install maturin if needed
pip install maturin
# Build and install into active virtual environment
maturin develop --release --features python
This single command compiles with optimizations enabled and registers the module for immediate import pdf_inspector usage. For distribution, maturin build --features python creates manylinux-compliant wheels.
Method 3: setuptools-rust
For projects integrating pdf-inspector into a larger Python package, add to pyproject.toml:
[build-system]
requires = ["setuptools", "setuptools-rust"]
[tool.setuptools-rust]
rust-extensions = [
{ path = "Cargo.toml", binding = "PyO3", args = ["--features", "python"] }
]
Then install with pip install . as with any source package.
Module Structure and Public API
The Python module is implemented in src/python.rs, which exports Rust structs as Python classes and functions via #[pyclass] and #[pyfunction] macros. The key types and their Rust source locations:
| Python Name | Rust Struct | Module Export | Description |
|---|---|---|---|
PdfResult |
PyPdfResult |
src/python.rs lines 15–54 |
Complete processing output with markdown, metadata, and layout |
PdfClassification |
PyPdfClassification |
src/python.rs lines 92–101 |
Lightweight PDF type detection result |
TextItem |
PyTextItem |
src/python.rs line 165+ |
Text with bounding box and font information |
All classes expose fields as read-only properties using #[pyo3(get)], making them introspectable like dataclasses without setters.
Core Functions for PDF Processing
The module exposes eight main functions in src/python.rs (lines 440–520+). Each has both file-path and in-memory bytes variants:
Full Processing Functions
process_pdf(path: str, pages: Optional[List[int]] = None) -> PdfResult (source)
: Executes the complete pipeline: PDF parsing, text extraction, OCR analysis, layout detection, and markdown generation. The optional pages parameter limits processing to specific 1-indexed page numbers.
process_pdf_bytes(data: bytes, pages: Optional[List[int]] = None) -> PdfResult (source)
: Identical functionality for PDFs already loaded in memory.
Detection and Classification
detect_pdf(path: str) -> PdfResult (source)
: Runs only the detection pipeline, returning a PdfResult with pdf_type populated but without full content extraction. Faster than process_pdf when you only need classification.
detect_pdf_bytes(data: bytes) -> PdfResult (source)
: In-memory variant of detect_pdf.
classify_pdf(path: str) -> PdfClassification (source)
: Returns just the PdfClassification with pdf_type (values like "text_based", " scanned", "hybrid"). Fastest option for type checking.
classify_pdf_bytes(data: bytes) -> PdfClassification
: In-memory variant of classify_pdf.
Plain Text Extraction
extract_text(path: str) -> str
: Returns raw extracted text without markdown formatting or metadata.
extract_text_bytes(data: bytes) -> str
: In-memory variant.
extract_text_with_positions(path: str, pages: Optional[List[int]] = None) -> List[TextItem]
: Returns TextItem objects containing text, x0, y0, x1, y1, font_name, and font_size for precise layout reconstruction.
extract_pages_markdown(path: str) -> List[str]
: Returns a list of markdown strings, one per page, preserving the document's structure hierarchy.
All functions raise PyValueError on invalid input, corrupted PDFs, or processing failures (see pyo3::exceptions::PyValueError usage throughout src/python.rs).
Practical Usage Examples
Installation Verification
import pdf_inspector
print(f"pdf-inspector Python bindings loaded: {pdf_inspector.__file__}")
Fast PDF Classification
import pdf_inspector
# Determine PDF type without expensive OCR
classification = pdf_inspector.classify_pdf("document.pdf")
print(f"Type: {classification.pdf_type}") # "text_based", "scanned", etc.
print(f"Confidence: {classification.confidence:.2f}")
Full Document Processing
import pdf_inspector
# Process all pages with complete metadata
result = pdf_inspector.process_pdf("report.pdf")
print(f"Pages: {result.page_count}")
print(f"Detected type: {result.pdf_type}")
print(f"Has text layer: {result.has_text}")
print(f"OCR required: {result.requires_ocr}")
# Access generated markdown
print(result.markdown[:500]) # First 500 characters
# Per-page markdown (if structure preservation needed)
pages_md = pdf_inspector.extract_pages_markdown("report.pdf")
for i, page_md in enumerate(pages_md, 1):
print(f"--- Page {i} ---")
print(page_md[:200])
In-Memory Processing (Web Applications)
import pdf_inspector
from pathlib import Path
# Load PDF from upload, database, or network
pdf_bytes = Path("uploaded.pdf").read_bytes()
# Process without writing to filesystem
result = pdf_inspector.process_pdf_bytes(pdf_bytes, pages=[1, 2, 3])
classification = pdf_inspector.classify_pdf_bytes(pdf_bytes)
Layout-Aware Text Extraction
import pdf_inspector
# Extract text with precise positioning for custom rendering
items = pdf_inspector.extract_text_with_positions("layout.pdf", pages=[1])
for item in items[:10]:
print(f"'{item.text[:30]}...' "
f"at ({item.x0:.1f}, {item.y0:.1f}) "
f"font={item.font_name} size={item.font_size}")
Error Handling
import pdf_inspector
from pyo3 import exceptions
try:
result = pdf_inspector.process_pdf("corrupted.pdf")
except Exception as e: # Catches PyValueError
print(f"Processing failed: {e}")
Key Source Files Reference
| File | Purpose | Link |
|---|---|---|
src/python.rs |
PyO3 bindings implementation: class definitions, function exports, error handling | src/python.rs |
Cargo.toml |
Feature flags, PyO3 dependency configuration, build profiles | Cargo.toml |
docs/python.md |
Extended Python API documentation and additional examples | docs/python.md |
Summary
- Enable the
pythonfeature inCargo.tomlto activate PyO3 bindings compilation - Use
maturin develop --features pythonfor the smoothest development workflow, orcargo build --features pythonfor manual builds - Import
pdf_inspectorafter installation; all functions are immediately available - Choose the right function for your latency needs:
classify_pdffor fastest type detection,process_pdffor complete analysis,extract_text_with_positionsfor layout control - Handle
PyValueErrorexceptions for robust production code
The PyO3-based bindings provide zero-copy efficiency where possible and ABI3 compatibility across Python 3.8–3.12, making pdf-inspector suitable for high-throughput document processing pipelines.
Frequently Asked Questions
What Python versions are supported by pdf-inspector's PyO3 bindings?
The bindings target CPython 3.8 and later through PyO3's abi3-py38 flag. This produces a stable ABI that works across 3.8, 3.9, 3.10, 3.11, and 3.12 without recompilation. Alternative Python implementations (PyPy, GraalPython) are not officially supported as of the current Cargo.toml configuration.
Can I use pdf-inspector in Python without building from source?
Pre-built wheels are not currently distributed on PyPI for firecrawl/pdf-inspector. You must build from source using the methods described above. The Maturin toolchain (pip install maturin; maturin develop --features python) handles compilation automatically with no manual Rust configuration required for most platforms.
How do I pass only specific pages to process_pdf?
Use the optional pages parameter with a list of 1-indexed integers: pdf_inspector.process_pdf("doc.pdf", pages=[1, 3, 5]). This limits both processing time and memory usage for large documents. The same parameter works for process_pdf_bytes and extract_text_with_positions.
What is the difference between detect_pdf and process_pdf?
detect_pdf runs only the classification pipeline to determine PDF type (text-based, scanned, hybrid) without extracting content or running OCR—optimal for routing decisions. process_pdf executes the full pipeline including markdown generation, text extraction, OCR metadata, and layout analysis. For large batch operations, use classify_pdf first to filter documents before expensive full processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →