# Core Dependencies for pdf-inspector: Inside Firecrawl's Rust PDF Parser

> Discover the core dependencies for pdf-inspector, Firecrawl's Rust PDF parser. Learn about lopdf, ttf-parser, regex, and more for efficient PDF analysis.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: internals
- Published: 2026-08-04

---

**The core dependencies for pdf-inspector include `lopdf` for PDF parsing, `ttf-parser` for font decoding, `regex` and `unicode-normalization` for text cleaning, `thiserror` and `log` for error handling, and optional `pyo3` for Python bindings.**

The `firecrawl/pdf-inspector` repository is a Rust library and command-line tool designed to extract structured Markdown or JSON from PDF documents. Understanding the core dependencies for pdf-inspector reveals how the project balances performance, accuracy, and cross-platform compatibility across native and WebAssembly targets.

## PDF Parsing and Document Processing

The foundation of pdf-inspector rests on robust crates that handle low-level PDF structure and parallel execution.

### Primary PDF Parsing with lopdf

The **`lopdf`** crate serves as the primary PDF parsing engine. It reads PDF structure, streams, and objects, enabling pdf-inspector to navigate document hierarchies and extract content streams. In [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs), `lopdf` powers the operator state machine that processes PDF commands like `Tj` (show text), `Td` (move text position), and `Tm` (set text matrix).

### Parallel Processing via rayon

For native builds, pdf-inspector leverages **`rayon`** to enable parallel processing of PDF pages and content streams. This dependency significantly speeds up extraction on multi-core systems when processing large documents. The WASM target excludes `rayon` to maintain compatibility with single-threaded browser environments.

## Font and Text Handling

Accurate text extraction requires sophisticated font parsing and Unicode normalization capabilities.

### TrueType Font Parsing

The **`ttf-parser`** crate parses TrueType font tables to extract CID-to-Unicode mappings. This is essential in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs), where pdf-inspector handles font width calculations, CMap caching, and fallback mechanisms for Identity-H CID fonts embedded within PDFs.

### Unicode Normalization

**`unicode-normalization`** supplies NFKC and NFKD normalization forms used during post-processing. When cleaning extracted text in [`src/markdown/postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/postprocess.rs), this crate helps expand ligatures and normalize characters to ensure consistent output across different PDF generation sources.

## Text Processing Utilities

Post-extraction cleaning relies on efficient regular expressions and lazy initialization patterns.

### Pattern Matching with regex

The **`regex`** crate drives URL detection, hyphenation fixes, and dot-leader removal in the post-processing pipeline. Compiled regular expressions are cached using **`once_cell`**, which provides lazy-initialized globals without the overhead of `lazy_static`. This combination appears throughout [`src/markdown/postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/postprocess.rs) for tasks like detecting broken words across line endings.

## Error Handling and Observability

Robust error propagation and logging are critical for a library used in both CLI and embedded contexts.

### Typed Error Management

**`thiserror`** provides a lightweight derive macro for creating clear, typed error enums. Throughout [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) and the extractor modules, this enables ergonomic error handling when PDF structures are malformed or fonts are missing required tables.

### Logging Infrastructure

**`log`** acts as the generic logging facade used across the codebase via macros like `log::debug!` and `log::info!`. For command-line binaries defined in [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) and [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs), **`env_logger`** initializes logging levels from environment variables such as `RUST_LOG=pdf_inspector::extractor::layout=debug`.

## Optional and Platform-Specific Dependencies

### Python Bindings with pyo3

When the `python` feature is enabled, **`pyo3`** exposes the Rust API to Python. This allows Python applications to call `process_pdf_with_options` directly without spawning external processes. The bindings wrap the core functionality defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), accepting parameters like `PdfOptions` to control JSON output and table detection.

### WASM Bundling with include_dir

For WebAssembly builds, **`include_dir`** embeds bundled CMap files directly into the binary. Since browser runtimes cannot access the filesystem, this dependency ensures font mapping tables are available when [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) processes embedded CID fonts in WASM environments.

## Practical Usage Examples

### Basic Library Usage

Extract PDF content to structured Markdown using the Rust API:

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::options::PdfOptions;

fn main() -> Result<(), pdf_inspector::Error> {
    let opts = PdfOptions {
        json: true,
        detect_tables: true,
        ..Default::default()
    };
    
    let output = process_pdf_with_options("example.pdf", &opts)?;
    println!("{}", output);
    Ok(())
}

```

### Command-Line Interface

Run the `pdf2md` binary with debug logging enabled:

```bash
RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- path/to/document.pdf --json

```

### Python Integration

With the `python` feature compiled, use the library from Python:

```python
import pdf_inspector

markdown = pdf_inspector.process_pdf("example.pdf", json=False)
print(markdown)

```

## Summary

- **PDF parsing** relies on `lopdf` with optional `rayon` parallelism for native targets.
- **Font handling** combines `ttf-parser` for TrueType tables and `unicode-normalization` for character standardization.
- **Text post-processing** uses `regex` and `once_cell` for efficient pattern matching and lazy initialization.
- **Error handling** implements `thiserror` for typed errors, while `log` and `env_logger` provide observability.
- **Cross-platform support** includes `pyo3` for Python bindings and `include_dir` for WASM resource embedding.

## Frequently Asked Questions

### What crate does pdf-inspector use to parse PDF files?

pdf-inspector uses the **`lopdf`** crate as its primary PDF parsing library. According to the `firecrawl/pdf-inspector` source code, `lopdf` handles reading PDF structure, streams, and objects, while pdf-inspector implements higher-level extraction logic in modules like [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs).

### How does pdf-inspector handle parallel processing?

The project uses **`rayon`** to enable parallel processing of PDF pages and content streams on native targets. This dependency is excluded from WASM builds to maintain browser compatibility. The parallelization occurs during the extraction phase orchestrated in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs).

### Can pdf-inspector be used from Python?

Yes, when compiled with the `python` feature, pdf-inspector exposes Rust functions to Python via **`pyo3`**. This allows Python scripts to import the library and call `process_pdf()` or `process_pdf_with_options()` directly without requiring a separate CLI subprocess.

### Why does pdf-inspector need ttf-parser?

**`ttf-parser`** is required to decode TrueType font tables embedded within PDFs. In [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs), this crate extracts CID-to-Unicode mappings necessary for converting internal font glyph IDs to readable Unicode text, particularly for documents using embedded subset fonts.