Core Dependencies for pdf-inspector: Inside Firecrawl's Rust PDF Parser

The core dependencies for pdf-inspector include lopdf for PDF parsing, ttf-parser for font decoding, regex and unicode-normalization for text cleaning, thiserror and log for error handling, and optional pyo3 for Python bindings.

The firecrawl/pdf-inspector repository is a Rust library and command-line tool designed to extract structured Markdown or JSON from PDF documents. Understanding the core dependencies for pdf-inspector reveals how the project balances performance, accuracy, and cross-platform compatibility across native and WebAssembly targets.

PDF Parsing and Document Processing

The foundation of pdf-inspector rests on robust crates that handle low-level PDF structure and parallel execution.

Primary PDF Parsing with lopdf

The lopdf crate serves as the primary PDF parsing engine. It reads PDF structure, streams, and objects, enabling pdf-inspector to navigate document hierarchies and extract content streams. In src/extractor/content_stream.rs, lopdf powers the operator state machine that processes PDF commands like Tj (show text), Td (move text position), and Tm (set text matrix).

Parallel Processing via rayon

For native builds, pdf-inspector leverages rayon to enable parallel processing of PDF pages and content streams. This dependency significantly speeds up extraction on multi-core systems when processing large documents. The WASM target excludes rayon to maintain compatibility with single-threaded browser environments.

Font and Text Handling

Accurate text extraction requires sophisticated font parsing and Unicode normalization capabilities.

TrueType Font Parsing

The ttf-parser crate parses TrueType font tables to extract CID-to-Unicode mappings. This is essential in src/extractor/fonts.rs, where pdf-inspector handles font width calculations, CMap caching, and fallback mechanisms for Identity-H CID fonts embedded within PDFs.

Unicode Normalization

unicode-normalization supplies NFKC and NFKD normalization forms used during post-processing. When cleaning extracted text in src/markdown/postprocess.rs, this crate helps expand ligatures and normalize characters to ensure consistent output across different PDF generation sources.

Text Processing Utilities

Post-extraction cleaning relies on efficient regular expressions and lazy initialization patterns.

Pattern Matching with regex

The regex crate drives URL detection, hyphenation fixes, and dot-leader removal in the post-processing pipeline. Compiled regular expressions are cached using once_cell, which provides lazy-initialized globals without the overhead of lazy_static. This combination appears throughout src/markdown/postprocess.rs for tasks like detecting broken words across line endings.

Error Handling and Observability

Robust error propagation and logging are critical for a library used in both CLI and embedded contexts.

Typed Error Management

thiserror provides a lightweight derive macro for creating clear, typed error enums. Throughout src/lib.rs and the extractor modules, this enables ergonomic error handling when PDF structures are malformed or fonts are missing required tables.

Logging Infrastructure

log acts as the generic logging facade used across the codebase via macros like log::debug! and log::info!. For command-line binaries defined in src/bin/pdf2md.rs and src/bin/detect_pdf.rs, env_logger initializes logging levels from environment variables such as RUST_LOG=pdf_inspector::extractor::layout=debug.

Optional and Platform-Specific Dependencies

Python Bindings with pyo3

When the python feature is enabled, pyo3 exposes the Rust API to Python. This allows Python applications to call process_pdf_with_options directly without spawning external processes. The bindings wrap the core functionality defined in src/lib.rs, accepting parameters like PdfOptions to control JSON output and table detection.

WASM Bundling with include_dir

For WebAssembly builds, include_dir embeds bundled CMap files directly into the binary. Since browser runtimes cannot access the filesystem, this dependency ensures font mapping tables are available when src/extractor/fonts.rs processes embedded CID fonts in WASM environments.

Practical Usage Examples

Basic Library Usage

Extract PDF content to structured Markdown using the Rust API:

use pdf_inspector::process_pdf_with_options;
use pdf_inspector::options::PdfOptions;

fn main() -> Result<(), pdf_inspector::Error> {
    let opts = PdfOptions {
        json: true,
        detect_tables: true,
        ..Default::default()
    };
    
    let output = process_pdf_with_options("example.pdf", &opts)?;
    println!("{}", output);
    Ok(())
}

Command-Line Interface

Run the pdf2md binary with debug logging enabled:

RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- path/to/document.pdf --json

Python Integration

With the python feature compiled, use the library from Python:

import pdf_inspector

markdown = pdf_inspector.process_pdf("example.pdf", json=False)
print(markdown)

Summary

  • PDF parsing relies on lopdf with optional rayon parallelism for native targets.
  • Font handling combines ttf-parser for TrueType tables and unicode-normalization for character standardization.
  • Text post-processing uses regex and once_cell for efficient pattern matching and lazy initialization.
  • Error handling implements thiserror for typed errors, while log and env_logger provide observability.
  • Cross-platform support includes pyo3 for Python bindings and include_dir for WASM resource embedding.

Frequently Asked Questions

What crate does pdf-inspector use to parse PDF files?

pdf-inspector uses the lopdf crate as its primary PDF parsing library. According to the firecrawl/pdf-inspector source code, lopdf handles reading PDF structure, streams, and objects, while pdf-inspector implements higher-level extraction logic in modules like src/extractor/content_stream.rs.

How does pdf-inspector handle parallel processing?

The project uses rayon to enable parallel processing of PDF pages and content streams on native targets. This dependency is excluded from WASM builds to maintain browser compatibility. The parallelization occurs during the extraction phase orchestrated in src/extractor/mod.rs.

Can pdf-inspector be used from Python?

Yes, when compiled with the python feature, pdf-inspector exposes Rust functions to Python via pyo3. This allows Python scripts to import the library and call process_pdf() or process_pdf_with_options() directly without requiring a separate CLI subprocess.

Why does pdf-inspector need ttf-parser?

ttf-parser is required to decode TrueType font tables embedded within PDFs. In src/extractor/fonts.rs, this crate extracts CID-to-Unicode mappings necessary for converting internal font glyph IDs to readable Unicode text, particularly for documents using embedded subset fonts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →