# What Is the Underlying Technology Used by PDF‑Inspector for PDF Parsing?

> Discover the Rust crate lopdf powering PDF-Inspector for efficient PDF parsing. Explore its low-level document handling, object lookup, and stream decompression capabilities.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-04

---

**PDF‑Inspector uses the Rust crate `lopdf` as its core PDF‑parsing engine, providing low‑level document representation, object lookup, stream decompression, and syntax handling.**

PDF‑Inspector, an open‑source Rust library from the Firecrawl project, depends entirely on `lopdf` to read and interpret PDF files. This crate supplies the foundational document model—objects, dictionaries, and streams—upon which PDF‑Inspector builds its detection, extraction, and markdown conversion pipelines.

## How PDF‑Inspector Integrates `lopdf` at the Core

`lopdf` is imported directly in the library's entry point and serves as the sole interface to PDF file structure.

In [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), the crate brings in `lopdf::Document` as the primary abstraction:

```rust
use lopdf::Document;

```

This `Document` type is the root object passed through every subsequent operation. PDF‑Inspector never parses raw PDF bytes directly; instead, it delegates all byte‑level work to `lopdf` and operates on the resulting structured objects.

## PDF Type Detection via `lopdf::Document`

The detection module ([`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)) determines whether a PDF is text‑based, scanned, image‑based, or mixed. It receives a `Document` instance from helper functions `crate::load_document_from_path` and `crate::load_document_from_mem`.

The `detect_pdf_type` and `detect_pdf_type_mem` functions navigate this `Document` to analyze page resources, content streams, and embedded fonts. Because `lopdf` has already resolved object references and decompressed streams, PDF‑Inspector can focus purely on heuristic analysis.

## Content Stream Parsing with Raw `lopdf` Streams

Text and graphics operators—`Tj`, `TJ`, `Td`, and others—are processed in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs). PDF‑Inspector extracts the raw byte streams via `lopdf` methods, then parses the operator sequences itself.

This split responsibility is deliberate: `lopdf` handles PDF‑specific encoding, filters (FlateDecode, LZWDecode, etc.), and object cross‑references, while PDF‑Inspector implements the higher‑level operator semantics for text positioning and extraction.

## Font and Encoding Handling Through `lopdf` Dictionaries

Character mapping relies heavily on `lopdf` dictionary access. In [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs), PDF‑Inspector reads `/ToUnicode` CMap streams and CID‑to‑GID mappings by traversing font dictionaries retrieved via `lopdf` APIs.

Without `lopdf`'s ability to resolve indirect objects and decode embedded streams, extracting accurate Unicode text from arbitrary PDF encodings would require reimplementing substantial portions of the PDF specification.

## Tagged‑PDF Structure Extraction

Accessibility and semantic structure trees are parsed in [`src/structure_tree.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/structure_tree.rs) using `lopdf::Dictionary` objects. The `StructTreeRoot` and parent‑tree hierarchies are navigated through `lopdf`'s map‑like dictionary interface, letting PDF‑Inspector reconstruct document semantics without manual object numbering or cross‑reference table handling.

## Practical Usage Examples

### Detect a PDF's Content Type

```rust
use pdf_inspector::{detect_pdf_type, PdfType};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = detect_pdf_type("example.pdf")?;
    match result.pdf_type {
        PdfType::TextBased => println!("Extractable text found"),
        PdfType::Scanned   => println!("Only images – OCR needed"),
        PdfType::ImageBased | PdfType::Mixed => println!("Mixed content"),
    }
    Ok(())
}

```

### Full Processing Pipeline

```rust
use pdf_inspector::{process_pdf, PdfOptions, ProcessMode};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let opts = PdfOptions::new()
        .mode(ProcessMode::Full);

    let res = pdf_inspector::process_pdf_with_options("report.pdf", opts)?;
    println!("Pages: {}", res.page_count);
    if let Some(md) = res.markdown {
        println!("Markdown output:\n{}", md);
    }
    Ok(())
}

```

### Extract Text with Positional Data

```rust
use pdf_inspector::extractor::extract_text_with_positions;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let items = extract_text_with_positions("paper.pdf")?;
    for item in items {
        println!("\"{}\" at ({}, {})", item.text, item.x, item.y);
    }
    Ok(())
}

```

## Key Source Files Demonstrating `lopdf` Usage

- **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** — Public API entry point; loads `lopdf::Document` and orchestrates detection, extraction, and markdown generation.
- **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** — PDF‑type detection operating on `lopdf::Document` structures.
- **[`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs)** — Content‑stream operator parsing from `lopdf`‑provided streams.
- **[`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)** — `/ToUnicode` CMap and font encoding resolution via `lopdf` dictionaries.
- **[`src/structure_tree.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/structure_tree.rs)** — Tagged‑PDF structure tree navigation using `lopdf` objects.

## Summary

- **Core engine**: `lopdf` Rust crate handles all PDF parsing, object resolution, and stream decompression.
- **Layered architecture**: PDF‑Inspector adds detection heuristics, text extraction, layout analysis, and markdown conversion on top.
- **Key dependency**: Every module—[`detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detector.rs), [`content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/content_stream.rs), [`tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/tounicode.rs), [`structure_tree.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/structure_tree.rs)—receives data through `lopdf::Document` or `lopdf::Dictionary` objects.
- **No redundant parsing**: PDF‑Inspector does not reimplement PDF syntax; it relies entirely on `lopdf` for specification compliance.

## Frequently Asked Questions

### What Rust crate does PDF‑Inspector use for PDF parsing?

PDF‑Inspector uses **`lopdf`**, a mature Rust library for reading and manipulating PDF documents. As shown in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), `lopdf::Document` is the central type passed through all internal operations.

### Does PDF‑Inspector parse PDF files from scratch?

No. PDF‑Inspector delegates all low‑level PDF parsing—including cross‑reference tables, object streams, and filter decompression—to `lopdf`. The codebase only implements higher‑level analysis and extraction logic.

### Why was `lopdf` chosen over other Rust PDF libraries?

`lopdf` provides a lightweight, zero‑copy‑friendly representation of PDF objects with minimal dependencies. Its dictionary‑and‑stream model maps cleanly to PDF‑Inspector's needs for text extraction, font handling, and structure tree navigation without imposing a specific high‑level API.

### Can I use PDF‑Inspector with PDFs that `lopdf` cannot parse?

PDF‑Inspector's capabilities are bounded by `lopdf` parsing support. If `lopdf` fails to load a malformed or heavily obfuscated PDF, PDF‑Inspector will surface that error through its `PdfError` type during `load_document_from_path` or `load_document_from_mem` calls.