Python Libraries for Firecrawl PDF-Inspector: Official PyO3 Bindings Guide
Yes, pdf-inspector provides official Python bindings built with PyO3, available on PyPI as the pdf-inspector package.
The firecrawl/pdf-inspector repository is a high-performance Rust library for PDF text extraction and classification. To make this functionality accessible to Python developers, the project ships first-class Python bindings that expose the same core API without requiring Rust knowledge at runtime.
Installing the Python Library
From PyPI (Recommended)
pip install pdf-inspector
Pre-built wheels are available for Windows, macOS, and Linux. No Rust toolchain is required for standard installation.
Building From Source
If you need the latest unreleased features or are on an unsupported platform:
pip install maturin
git clone https://github.com/firecrawl/pdf-inspector.git
cd pdf-inspector
maturin develop # builds and installs into current virtualenv
Source builds require a working Rust compiler.
Core Python API
The bindings expose two primary functions implemented in src/python.rs:
| Function | Purpose | Rust Equivalent |
|---|---|---|
detect_pdf(path: str) -> str |
Classifies PDF type | detect_pdf in src/detector.rs |
process_pdf(path: str, options: dict) -> dict |
Extracts text, tables, and converts to Markdown | process_pdf_with_options in src/lib.rs |
Basic Usage Example
from pdf_inspector import process_pdf, detect_pdf
# Classify the PDF type: "TextBased", "Scanned", or "Mixed"
pdf_type = detect_pdf("sample.pdf")
print(f"PDF type: {pdf_type}")
# Extract content with Markdown output
result = process_pdf(
"sample.pdf",
{
"extract_text": True,
"extract_tables": True,
"output_format": "markdown",
},
)
# Available in the returned dict:
# - pages: list of pages with positional text items
# - tables: structured table data
# - markdown: clean Markdown string
print(result["markdown"])
Advanced Configuration Options
The options dict accepts fine-grained control over extraction behavior:
options = {
"extract_text": True, # Enable text extraction
"extract_tables": True, # Detect and structure tables
"detect_columns": True, # Identify multi-column layouts
"detect_headers": True, # Recognize header sections
"json_output": True, # Return JSON instead of Markdown
"max_columns": 25, # Cap table column detection (default: 25)
"detect_images": False, # Skip image extraction for faster processing
}
data = process_pdf("report.pdf", options)
# Work directly with structured output
for table in data["tables"]:
print(f"Table title: {table['title']}")
for row in table["rows"]:
print(row)
Set "json_output": True when you need programmatic access to raw positional data; use "output_format": "markdown" for human-readable or LLM-friendly output.
Key Implementation Files
Understanding the source structure helps with debugging and advanced usage:
src/python.rs— PyO3 binding implementation; definesprocess_pdfanddetect_pdfPython functionsdocs/python.md— Official Python documentation and API referenceCargo.toml— Declarespyo3dependency andcdylibtarget configuration for wheel buildssrc/lib.rs— Rust core containingprocess_pdf_with_optionsand public crate APIsrc/detector.rs— PDF classification logic behinddetect_pdf
Integration Patterns
The Python library fits cleanly into common workflows:
Data pipelines — Process PDF batches with consistent Markdown output for downstream NLP or RAG systems.
LLM preprocessing — Convert documents to structured Markdown before chunking and embedding.
Document classification — Use detect_pdf to route text-based PDFs through fast extraction and scanned documents through OCR pipelines.
Summary
- Install with
pip install pdf-inspector— pre-built wheels require no Rust toolchain - Two main functions:
detect_pdf()for classification,process_pdf()for extraction - Output formats: Markdown for readability, JSON/dict for structured data access
- Same performance as Rust — bindings add minimal overhead via PyO3's zero-copy APIs
- Source builds available via
maturinfor development or unsupported platforms
Frequently Asked Questions
What Python versions are supported?
The pdf-inspector package supports Python 3.8 through 3.12 according to the PyPI classifiers in Cargo.toml. Wheels are built for CPython on all major platforms.
Do I need to install Rust to use the Python library?
No — pre-compiled wheels from PyPI contain the native extension. You only need Rust if building from source or contributing to the bindings.
How does the Python API compare to using the Rust library directly?
The Python API mirrors the Rust public API with idiomatic conversions: Rust structs become Python dicts, Option<T> becomes optional parameters, and errors raise standard Python exceptions. Performance is nearly identical since PyO3 minimizes serialization overhead.
Can I use this with asyncio or in async applications?
The current bindings in src/python.rs expose synchronous functions. For async workflows, run process_pdf in asyncio.to_thread or loop.run_in_executor to avoid blocking the event loop during PDF processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →