# Python Libraries for Firecrawl PDF-Inspector: Official PyO3 Bindings Guide

> Explore official Python bindings for Firecrawl PDF-Inspector. Discover how to use the PyO3-built pdf-inspector package for powerful PDF analysis and extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: api-reference
- Published: 2026-08-07

---

**Yes, `pdf-inspector` provides official Python bindings built with PyO3, available on PyPI as the `pdf-inspector` package.**

The `firecrawl/pdf-inspector` repository is a high-performance Rust library for PDF text extraction and classification. To make this functionality accessible to Python developers, the project ships first-class Python bindings that expose the same core API without requiring Rust knowledge at runtime.

## Installing the Python Library

### From PyPI (Recommended)

```bash
pip install pdf-inspector

```

Pre-built wheels are available for Windows, macOS, and Linux. No Rust toolchain is required for standard installation.

### Building From Source

If you need the latest unreleased features or are on an unsupported platform:

```bash
pip install maturin
git clone https://github.com/firecrawl/pdf-inspector.git
cd pdf-inspector
maturin develop  # builds and installs into current virtualenv

```

Source builds require a working Rust compiler.

## Core Python API

The bindings expose two primary functions implemented in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs):

| Function | Purpose | Rust Equivalent |
|----------|---------|---------------|
| `detect_pdf(path: str) -> str` | Classifies PDF type | `detect_pdf` in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) |
| `process_pdf(path: str, options: dict) -> dict` | Extracts text, tables, and converts to Markdown | `process_pdf_with_options` in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) |

## Basic Usage Example

```python
from pdf_inspector import process_pdf, detect_pdf

# Classify the PDF type: "TextBased", "Scanned", or "Mixed"

pdf_type = detect_pdf("sample.pdf")
print(f"PDF type: {pdf_type}")

# Extract content with Markdown output

result = process_pdf(
    "sample.pdf",
    {
        "extract_text": True,
        "extract_tables": True,
        "output_format": "markdown",
    },
)

# Available in the returned dict:

#   - pages: list of pages with positional text items

#   - tables: structured table data

#   - markdown: clean Markdown string

print(result["markdown"])

```

## Advanced Configuration Options

The `options` dict accepts fine-grained control over extraction behavior:

```python
options = {
    "extract_text": True,       # Enable text extraction

    "extract_tables": True,     # Detect and structure tables

    "detect_columns": True,     # Identify multi-column layouts

    "detect_headers": True,     # Recognize header sections

    "json_output": True,        # Return JSON instead of Markdown

    "max_columns": 25,          # Cap table column detection (default: 25)

    "detect_images": False,     # Skip image extraction for faster processing

}

data = process_pdf("report.pdf", options)

# Work directly with structured output

for table in data["tables"]:
    print(f"Table title: {table['title']}")
    for row in table["rows"]:
        print(row)

```

Set `"json_output": True` when you need programmatic access to raw positional data; use `"output_format": "markdown"` for human-readable or LLM-friendly output.

## Key Implementation Files

Understanding the source structure helps with debugging and advanced usage:

- **[`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs)** — PyO3 binding implementation; defines `process_pdf` and `detect_pdf` Python functions
- **[`docs/python.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md)** — Official Python documentation and API reference
- **[`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml)** — Declares `pyo3` dependency and `cdylib` target configuration for wheel builds
- **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** — Rust core containing `process_pdf_with_options` and public crate API
- **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** — PDF classification logic behind `detect_pdf`

## Integration Patterns

The Python library fits cleanly into common workflows:

**Data pipelines** — Process PDF batches with consistent Markdown output for downstream NLP or RAG systems.

**LLM preprocessing** — Convert documents to structured Markdown before chunking and embedding.

**Document classification** — Use `detect_pdf` to route text-based PDFs through fast extraction and scanned documents through OCR pipelines.

## Summary

- **Install with** `pip install pdf-inspector` — pre-built wheels require no Rust toolchain
- **Two main functions:** `detect_pdf()` for classification, `process_pdf()` for extraction
- **Output formats:** Markdown for readability, JSON/dict for structured data access
- **Same performance as Rust** — bindings add minimal overhead via PyO3's zero-copy APIs
- **Source builds** available via `maturin` for development or unsupported platforms

## Frequently Asked Questions

### What Python versions are supported?

The `pdf-inspector` package supports Python 3.8 through 3.12 according to the PyPI classifiers in [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml). Wheels are built for CPython on all major platforms.

### Do I need to install Rust to use the Python library?

No — pre-compiled wheels from PyPI contain the native extension. You only need Rust if building from source or contributing to the bindings.

### How does the Python API compare to using the Rust library directly?

The Python API mirrors the Rust public API with idiomatic conversions: Rust structs become Python dicts, `Option<T>` becomes optional parameters, and errors raise standard Python exceptions. Performance is nearly identical since PyO3 minimizes serialization overhead.

### Can I use this with asyncio or in async applications?

The current bindings in [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) expose synchronous functions. For async workflows, run `process_pdf` in `asyncio.to_thread` or `loop.run_in_executor` to avoid blocking the event loop during PDF processing.