Python Libraries for Firecrawl PDF-Inspector: Official PyO3 Bindings Guide

Yes, pdf-inspector provides official Python bindings built with PyO3, available on PyPI as the pdf-inspector package.

The firecrawl/pdf-inspector repository is a high-performance Rust library for PDF text extraction and classification. To make this functionality accessible to Python developers, the project ships first-class Python bindings that expose the same core API without requiring Rust knowledge at runtime.

Installing the Python Library

pip install pdf-inspector

Pre-built wheels are available for Windows, macOS, and Linux. No Rust toolchain is required for standard installation.

Building From Source

If you need the latest unreleased features or are on an unsupported platform:

pip install maturin
git clone https://github.com/firecrawl/pdf-inspector.git
cd pdf-inspector
maturin develop  # builds and installs into current virtualenv

Source builds require a working Rust compiler.

Core Python API

The bindings expose two primary functions implemented in src/python.rs:

Function Purpose Rust Equivalent
detect_pdf(path: str) -> str Classifies PDF type detect_pdf in src/detector.rs
process_pdf(path: str, options: dict) -> dict Extracts text, tables, and converts to Markdown process_pdf_with_options in src/lib.rs

Basic Usage Example

from pdf_inspector import process_pdf, detect_pdf

# Classify the PDF type: "TextBased", "Scanned", or "Mixed"

pdf_type = detect_pdf("sample.pdf")
print(f"PDF type: {pdf_type}")

# Extract content with Markdown output

result = process_pdf(
    "sample.pdf",
    {
        "extract_text": True,
        "extract_tables": True,
        "output_format": "markdown",
    },
)

# Available in the returned dict:

#   - pages: list of pages with positional text items

#   - tables: structured table data

#   - markdown: clean Markdown string

print(result["markdown"])

Advanced Configuration Options

The options dict accepts fine-grained control over extraction behavior:

options = {
    "extract_text": True,       # Enable text extraction

    "extract_tables": True,     # Detect and structure tables

    "detect_columns": True,     # Identify multi-column layouts

    "detect_headers": True,     # Recognize header sections

    "json_output": True,        # Return JSON instead of Markdown

    "max_columns": 25,          # Cap table column detection (default: 25)

    "detect_images": False,     # Skip image extraction for faster processing

}

data = process_pdf("report.pdf", options)

# Work directly with structured output

for table in data["tables"]:
    print(f"Table title: {table['title']}")
    for row in table["rows"]:
        print(row)

Set "json_output": True when you need programmatic access to raw positional data; use "output_format": "markdown" for human-readable or LLM-friendly output.

Key Implementation Files

Understanding the source structure helps with debugging and advanced usage:

  • src/python.rs — PyO3 binding implementation; defines process_pdf and detect_pdf Python functions
  • docs/python.md — Official Python documentation and API reference
  • Cargo.toml — Declares pyo3 dependency and cdylib target configuration for wheel builds
  • src/lib.rs — Rust core containing process_pdf_with_options and public crate API
  • src/detector.rs — PDF classification logic behind detect_pdf

Integration Patterns

The Python library fits cleanly into common workflows:

Data pipelines — Process PDF batches with consistent Markdown output for downstream NLP or RAG systems.

LLM preprocessing — Convert documents to structured Markdown before chunking and embedding.

Document classification — Use detect_pdf to route text-based PDFs through fast extraction and scanned documents through OCR pipelines.

Summary

  • Install with pip install pdf-inspector — pre-built wheels require no Rust toolchain
  • Two main functions: detect_pdf() for classification, process_pdf() for extraction
  • Output formats: Markdown for readability, JSON/dict for structured data access
  • Same performance as Rust — bindings add minimal overhead via PyO3's zero-copy APIs
  • Source builds available via maturin for development or unsupported platforms

Frequently Asked Questions

What Python versions are supported?

The pdf-inspector package supports Python 3.8 through 3.12 according to the PyPI classifiers in Cargo.toml. Wheels are built for CPython on all major platforms.

Do I need to install Rust to use the Python library?

No — pre-compiled wheels from PyPI contain the native extension. You only need Rust if building from source or contributing to the bindings.

How does the Python API compare to using the Rust library directly?

The Python API mirrors the Rust public API with idiomatic conversions: Rust structs become Python dicts, Option<T> becomes optional parameters, and errors raise standard Python exceptions. Performance is nearly identical since PyO3 minimizes serialization overhead.

Can I use this with asyncio or in async applications?

The current bindings in src/python.rs expose synchronous functions. For async workflows, run process_pdf in asyncio.to_thread or loop.run_in_executor to avoid blocking the event loop during PDF processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →