# How to Build and Use the Python Bindings for pdf-inspector with PyO3

> Learn to build and use Python bindings for pdf-inspector with PyO3. Compile a native Python extension for CPython 3.8+ using Cargo build or Maturin develop.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Enable the `python` feature in Cargo.toml and use `cargo build --features python` or `maturin develop` to compile a native Python extension module that exposes pdf-inspector's Rust API to CPython 3.8+.**

The pdf-inspector repository provides optional Python bindings powered by **PyO3**, a Rust crate for creating Python extension modules. These bindings let you call the library's high-performance PDF processing functions directly from Python code without leaving the interpreter. This guide covers the complete build process, installation methods, and practical usage examples based on the actual source implementation in `firecrawl/pdf-inspector`.

## Prerequisites and Feature Configuration

The Python bindings are gated behind an optional Cargo feature to keep the core library lightweight. Before building, verify your environment meets these requirements:

- **Rust toolchain** (stable channel, compatible with edition 2021)
- **Python 3.8 or later** (the bindings use PyO3's `abi3-py38` flag for broad compatibility)
- **pip and virtual environment** (recommended for isolation)

In [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml), the `python` feature is declared with PyO3 configured for extension-module builds:

```toml
[dependencies]
pyo3 = { version = "0.25", features = ["extension-module", "abi3-py38"], optional = true }

[features]
python = ["pyo3"]  # Cargo.toml line 64

```

The `extension-module` flag tells PyO3 to build a shared library suitable for Python import, while `abi3-py38` ensures the resulting wheel works across CPython 3.8–3.12 without recompilation (source: [Cargo.toml lines 30–64](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml#L30-L64)).

## Building the Python Bindings

You have two primary paths to build and install the extension: direct Cargo compilation or Python-native tooling with Maturin.

### Method 1: Cargo Build (Manual)

Compile the shared library directly with Cargo:

```bash
cargo build --release --features python

```

This produces a platform-specific shared library in `target/release/`:
- Linux/macOS: `libpdf_inspector.so` or `pdf_inspector.so` (renamed for import)
- Windows: `pdf_inspector.pyd`

To use the compiled library, either copy it to your working directory or add `target/release` to `PYTHONPATH`.

### Method 2: Maturin (Recommended)

**Maturin** streamlines the build-and-install workflow for Python developers:

```bash

# Install maturin if needed

pip install maturin

# Build and install into active virtual environment

maturin develop --release --features python

```

This single command compiles with optimizations enabled and registers the module for immediate `import pdf_inspector` usage. For distribution, `maturin build --features python` creates `manylinux`-compliant wheels.

### Method 3: setuptools-rust

For projects integrating pdf-inspector into a larger Python package, add to [`pyproject.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/pyproject.toml):

```toml
[build-system]
requires = ["setuptools", "setuptools-rust"]

[tool.setuptools-rust]
rust-extensions = [
    { path = "Cargo.toml", binding = "PyO3", args = ["--features", "python"] }
]

```

Then install with `pip install .` as with any source package.

## Module Structure and Public API

The Python module is implemented in **[`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs)**, which exports Rust structs as Python classes and functions via `#[pyclass]` and `#[pyfunction]` macros. The key types and their Rust source locations:

| Python Name | Rust Struct | Module Export | Description |
|-------------|-------------|---------------|-------------|
| `PdfResult` | `PyPdfResult` | [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) lines 15–54 | Complete processing output with markdown, metadata, and layout |
| `PdfClassification` | `PyPdfClassification` | [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) lines 92–101 | Lightweight PDF type detection result |
| `TextItem` | `PyTextItem` | [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) line 165+ | Text with bounding box and font information |

All classes expose fields as read-only properties using `#[pyo3(get)]`, making them introspectable like `dataclasses` without setters.

## Core Functions for PDF Processing

The module exposes eight main functions in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) (lines 440–520+). Each has both file-path and in-memory `bytes` variants:

### Full Processing Functions

**`process_pdf(path: str, pages: Optional[List[int]] = None) -> PdfResult`** ([source](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs#L440-L452))
: Executes the complete pipeline: PDF parsing, text extraction, OCR analysis, layout detection, and markdown generation. The optional `pages` parameter limits processing to specific 1-indexed page numbers.

**`process_pdf_bytes(data: bytes, pages: Optional[List[int]] = None) -> PdfResult`** ([source](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs#L452-L463))
: Identical functionality for PDFs already loaded in memory.

### Detection and Classification

**`detect_pdf(path: str) -> PdfResult`** ([source](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs#L463-L470))
: Runs only the detection pipeline, returning a `PdfResult` with `pdf_type` populated but without full content extraction. Faster than `process_pdf` when you only need classification.

**`detect_pdf_bytes(data: bytes) -> PdfResult`** ([source](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs#L470-L479))
: In-memory variant of `detect_pdf`.

**`classify_pdf(path: str) -> PdfClassification`** ([source](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs#L479-L487))
: Returns just the `PdfClassification` with `pdf_type` (values like `"text_based"`, `" scanned"`, `"hybrid"`). Fastest option for type checking.

**`classify_pdf_bytes(data: bytes) -> PdfClassification`**
: In-memory variant of `classify_pdf`.

### Plain Text Extraction

**`extract_text(path: str) -> str`**
: Returns raw extracted text without markdown formatting or metadata.

**`extract_text_bytes(data: bytes) -> str`**
: In-memory variant.

**`extract_text_with_positions(path: str, pages: Optional[List[int]] = None) -> List[TextItem]`**
: Returns `TextItem` objects containing `text`, `x0`, `y0`, `x1`, `y1`, `font_name`, and `font_size` for precise layout reconstruction.

**`extract_pages_markdown(path: str) -> List[str]`**
: Returns a list of markdown strings, one per page, preserving the document's structure hierarchy.

All functions raise **`PyValueError`** on invalid input, corrupted PDFs, or processing failures (see `pyo3::exceptions::PyValueError` usage throughout [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs)).

## Practical Usage Examples

### Installation Verification

```python
import pdf_inspector

print(f"pdf-inspector Python bindings loaded: {pdf_inspector.__file__}")

```

### Fast PDF Classification

```python
import pdf_inspector

# Determine PDF type without expensive OCR

classification = pdf_inspector.classify_pdf("document.pdf")
print(f"Type: {classification.pdf_type}")  # "text_based", "scanned", etc.

print(f"Confidence: {classification.confidence:.2f}")

```

### Full Document Processing

```python
import pdf_inspector

# Process all pages with complete metadata

result = pdf_inspector.process_pdf("report.pdf")

print(f"Pages: {result.page_count}")
print(f"Detected type: {result.pdf_type}")
print(f"Has text layer: {result.has_text}")
print(f"OCR required: {result.requires_ocr}")

# Access generated markdown

print(result.markdown[:500])  # First 500 characters

# Per-page markdown (if structure preservation needed)

pages_md = pdf_inspector.extract_pages_markdown("report.pdf")
for i, page_md in enumerate(pages_md, 1):
    print(f"--- Page {i} ---")
    print(page_md[:200])

```

### In-Memory Processing (Web Applications)

```python
import pdf_inspector
from pathlib import Path

# Load PDF from upload, database, or network

pdf_bytes = Path("uploaded.pdf").read_bytes()

# Process without writing to filesystem

result = pdf_inspector.process_pdf_bytes(pdf_bytes, pages=[1, 2, 3])
classification = pdf_inspector.classify_pdf_bytes(pdf_bytes)

```

### Layout-Aware Text Extraction

```python
import pdf_inspector

# Extract text with precise positioning for custom rendering

items = pdf_inspector.extract_text_with_positions("layout.pdf", pages=[1])

for item in items[:10]:
    print(f"'{item.text[:30]}...' "
          f"at ({item.x0:.1f}, {item.y0:.1f}) "
          f"font={item.font_name} size={item.font_size}")

```

### Error Handling

```python
import pdf_inspector
from pyo3 import exceptions

try:
    result = pdf_inspector.process_pdf("corrupted.pdf")
except Exception as e:  # Catches PyValueError

    print(f"Processing failed: {e}")

```

## Key Source Files Reference

| File | Purpose | Link |
|------|---------|------|
| [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) | PyO3 bindings implementation: class definitions, function exports, error handling | [src/python.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) |
| [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml) | Feature flags, PyO3 dependency configuration, build profiles | [Cargo.toml](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml) |
| [`docs/python.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md) | Extended Python API documentation and additional examples | [docs/python.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md) |

## Summary

- **Enable the `python` feature** in [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml) to activate PyO3 bindings compilation
- **Use `maturin develop --features python`** for the smoothest development workflow, or `cargo build --features python` for manual builds
- **Import `pdf_inspector`** after installation; all functions are immediately available
- **Choose the right function** for your latency needs: `classify_pdf` for fastest type detection, `process_pdf` for complete analysis, `extract_text_with_positions` for layout control
- **Handle `PyValueError`** exceptions for robust production code

The PyO3-based bindings provide **zero-copy efficiency** where possible and **ABI3 compatibility** across Python 3.8–3.12, making pdf-inspector suitable for high-throughput document processing pipelines.

## Frequently Asked Questions

### What Python versions are supported by pdf-inspector's PyO3 bindings?

The bindings target **CPython 3.8 and later** through PyO3's `abi3-py38` flag. This produces a stable ABI that works across 3.8, 3.9, 3.10, 3.11, and 3.12 without recompilation. Alternative Python implementations (PyPy, GraalPython) are not officially supported as of the current [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml) configuration.

### Can I use pdf-inspector in Python without building from source?

Pre-built wheels are not currently distributed on PyPI for `firecrawl/pdf-inspector`. You must build from source using the methods described above. The Maturin toolchain (`pip install maturin; maturin develop --features python`) handles compilation automatically with no manual Rust configuration required for most platforms.

### How do I pass only specific pages to process_pdf?

Use the optional `pages` parameter with a list of 1-indexed integers: `pdf_inspector.process_pdf("doc.pdf", pages=[1, 3, 5])`. This limits both processing time and memory usage for large documents. The same parameter works for `process_pdf_bytes` and `extract_text_with_positions`.

### What is the difference between detect_pdf and process_pdf?

**`detect_pdf`** runs only the classification pipeline to determine PDF type (text-based, scanned, hybrid) without extracting content or running OCR—optimal for routing decisions. **`process_pdf`** executes the full pipeline including markdown generation, text extraction, OCR metadata, and layout analysis. For large batch operations, use `classify_pdf` first to filter documents before expensive full processing.