# How pdf-inspector Parses PDF Files: A Technical Deep Dive into the Rust Pipeline

> Discover how pdf-inspector parses PDF files using a Rust pipeline for type detection, content extraction with lopdf, and layout processing. Transform PDFs into structured text and Markdown.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-04

---

**pdf-inspector parses PDF files through a three-stage Rust pipeline—type detection, low-level content-stream extraction with `lopdf`, and sophisticated layout post-processing—to transform binary documents into structured text items, rectangles, and Markdown output.**

The `firecrawl/pdf-inspector` repository implements a complete PDF parsing engine that bridges the gap between raw binary content and structured data. Understanding how pdf-inspector parses PDF files reveals a state-machine architecture capable of handling complex font encodings, graphics transformations, and multi-column layouts. The parser processes documents through distinct detection, extraction, and analysis phases to produce clean, machine-readable output.

## Stage 1: PDF Type Detection and Loading (src/detector.rs)

Before extraction begins, pdf-inspector classifies the document to determine processing strategy. The `detect_pdf_type` function in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) loads the file using `lopdf` into a `Document` struct, then runs a heuristic analysis on sampled pages.

The detection routine scans content streams via `analyze_page_content` to count text operators (`Tj`, `TJ`), image references (`Do`), and vector outlines. It specifically checks for problematic fonts—such as Identity-H/V fonts lacking ToUnicode CMaps or Type 3 fonts without Unicode mappings—and calculates text-to-image ratios. Based on thresholds defined in `DetectionConfig`, it returns a `PdfTypeResult` classifying the PDF as **TextBased**, **Scanned**, **ImageBased**, or **Mixed**, along with per-page OCR recommendations.

```rust
pub fn detect_pdf_type<P: AsRef<Path>>(path: P) -> Result<PdfTypeResult, PdfError> {
    // 1️⃣ Validate the file
    crate::validate_pdf_file(&path)?;
    // 2️⃣ Load once (shared for detection & extraction)
    let (doc, page_count) = crate::load_document_from_path(&path)?;
    // 3️⃣ Run heuristic analysis on a sampled set of pages
    detect_from_document(&doc, page_count, &DetectionConfig::default())
}

```

## Stage 2: Content-Stream Parsing and Text Extraction (src/extractor/content_stream.rs)

The core parsing logic resides in `extract_page_text_items` within [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs). This function implements a state-machine that walks every page’s content stream, tracking the graphics state—including the Current Transformation Matrix (CTM), text matrix, font, and spacing parameters.

The parser strips PDF comments from raw bytes, decodes the stream into operators using `Content::decode`, and iterates through operations:

- **Text operators** (`Tj`, `TJ`, `'`, `"`): Decode operands using the font’s ToUnicode CMap or fallback strategies and emit `TextItem` structs with device-space coordinates.
- **State operators** (`BT`, `ET`, `Tf`, `Td`, `TD`): Manage text block boundaries, font selection, and positioning.
- **Drawing operators** (`re`, path construction): Collect rectangles and lines for table detection.
- **XObjects** (`Do`): Handle Image XObjects (emitting placeholders with bounding boxes computed by `image_bbox_from_ctm`) and recursively process Form XObjects via `extract_form_xobject_text`.

```rust
pub(crate) fn extract_page_text_items(
    doc: &Document,
    page_id: ObjectId,
    page_num: u32,
    font_cmaps: &FontCMaps,
    include_invisible: bool,
    style_cache: &mut FontStyleCache,
) -> Result<(PageExtraction, bool, bool), PdfError> {
    // 1️⃣ Get the raw content bytes and strip PDF comments
    let content_data = doc.get_page_content(page_id)?;
    let content_data = strip_pdf_comments(&content_data);

    // 2️⃣ Decode the stream into a list of operators
    let content = Content::decode(&content_data)?;

    // 3️⃣ Iterate over all operators
    for op in &content.operations {
        match op.operator.as_str() {
            "BT" => { /* begin text block */ }
            "ET" => { /* end text block */ }
            "Tf" => { /* set current font */ }
            "Tj" => {
                // Decode operand → Unicode text using ToUnicode or fallback
                if let Some(text) = extract_text_from_operand(...) {
                    items.push(TextItem { … });
                }
            }
            "Do" => { /* image or form XObject */ }
            "re" => { /* rectangle – collect for table detection */ }
            // … other operators update graphics state
        }
    }
    Ok(((items, rects, lines), has_gid_fonts, coords_rotated))
}

```

### Font Handling and Unicode Decoding

Font decoding occurs through the `tounicode` module ([`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)). The `build_font_encodings` function constructs per-font `Encoding` objects from PDF *Differences* arrays. For Unicode mapping, the parser first attempts ToUnicode CMap parsing; if unavailable, it falls back to CID-to-Unicode heuristics or embedded TrueType/OpenType cmap tables. This ensures accurate text extraction even from documents with legacy or malformed font encodings.

### Graphics State and Coordinate Transformation

The parser maintains precise positional data through matrix operations. Functions like `multiply_matrices` and `rise_adjusted` in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) apply the CTM and text rise (`Ts`) values to transform glyph positions into device-space coordinates stored in each `TextItem`.

### Marked-Content and ActualText

The state-machine captures marked-content blocks (`BDC`/`EMC`) to extract *ActualText* entries—invisible text layers often embedded in scanned PDFs—without rendering their glyphs, preserving semantic content while ignoring display artifacts.

## Stage 3: Post-Processing and Layout Analysis (src/extractor/mod.rs)

After raw extraction, `extract_positioned_text_impl` orchestrates post-processing to reconstruct logical document structure:

- **Clipping**: Items are clipped to the page’s CropBox or MediaBox using `get_page_box`.
- **Letter-spacing correction**: `fix_letterspaced_items` repairs artificial spacing introduced by PDF tracking.
- **Text merging**: `merge_text_items` combines adjacent glyphs into words while respecting font size and style; `merge_subscript_items` handles chemical formulas and footnote markers.
- **Layout detection**: Functions in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) identify columns (`detect_columns`), newspaper layouts (`is_newspaper_layout`), and table structures using the `tables` module’s rectangle-, line-, and heuristic-based detectors.

The final output populates a `PdfProcessResult` struct containing the detected PDF type, optional Markdown representation, processing metadata, and lists of pages requiring OCR.

## Working with the pdf-inspector API

The public API in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) exposes high-level functions for common use cases:

```rust
use pdf_inspector::{process_pdf, extract_text, extract_text_with_positions};

fn main() -> Result<(), pdf_inspector::PdfError> {
    // Full processing: type detection → extraction → Markdown
    let result = process_pdf("reports/annual_report.pdf")?;
    println!("Detected type: {:?} (confidence {:.2})", result.pdf_type, result.confidence);
    if let Some(md) = result.markdown {
        println!("--- Markdown output ---\n{}", md);
    }

    // Fast detection only (no text extraction)
    let info = pdf_inspector::detect_pdf("scanned/document.pdf")?;
    println!("PDF type: {:?}, pages needing OCR: {:?}", info.pdf_type, info.pages_needing_ocr);

    // Get raw positioned text for custom layout work
    let items = extract_text_with_positions("tables/mixed_layout.pdf")?;
    for item in items.iter().take(5) {
        println!("Page {} – ({:.1},{:.1}) '{}' [{:.1}pt]",
                 item.page, item.x, item.y, item.text, item.font_size);
    }

    Ok(())
}

```

Key source files implementing this pipeline include:

- **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)**: PDF type classification and OCR recommendation logic.
- **[`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs)**: State-machine parser for PDF content streams.
- **[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)**: Orchestration, clipping, text merging, and layout helpers.
- **[`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)**: ToUnicode CMap parsing and font fallback strategies.
- **[`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs)**: Column, newspaper, and table detection algorithms.
- **[`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs)**: Core data structures including `TextItem`, `PdfRect`, and `PdfProcessResult`.

## Summary

- **pdf-inspector** uses a three-stage pipeline: type detection via `detect_pdf_type`, content extraction via `extract_page_text_items`, and layout reconstruction through post-processing functions.
- The parser relies on `lopdf` for document loading but implements a custom state-machine to handle PDF operators, graphics state matrices, and font encoding complexities.
- Advanced features include ToUnicode CMap parsing with fallback heuristics, recursive Form XObject processing, and sophisticated layout detection for tables and columns.
- The API provides both high-level Markdown generation via `process_pdf` and low-level positioned text access via `extract_text_with_positions` for custom processing needs.

## Frequently Asked Questions

### What Rust library does pdf-inspector use to read PDF files?

pdf-inspector uses the `lopdf` crate to load PDF files into a `Document` struct and access page content streams. However, it implements a custom state-machine parser in [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) to walk operators and extract text, rather than relying solely on `lopdf`'s built-in text extraction.

### How does pdf-inspector handle fonts without Unicode mappings?

When encountering fonts lacking ToUnicode CMaps—such as Identity-H/V encodings or Type 3 fonts—the parser in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) applies fallback strategies including CID-to-Unicode heuristics and embedded TrueType/OpenType cmap table analysis. If decoding fails, the system flags `has_encoding_issues` in the results.

### Can pdf-inspector detect whether a PDF requires OCR processing?

Yes, the `detect_pdf_type` function analyzes content streams for text operator density versus image presence and checks for problematic font configurations. It classifies documents as **TextBased**, **Scanned**, **ImageBased**, or **Mixed**, returning a `pages_needing_ocr` vector indicating which pages lack extractable text layers.

### How does pdf-inspector reconstruct tables and multi-column layouts?

After extracting raw text items, pdf-inspector runs post-processing functions including `merge_text_items` for word reconstruction and `detect_columns` for layout analysis. The [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) module and `src/tables/` subdirectory implement rectangle detection, line analysis, and heuristic algorithms to identify table structures and newspaper-style column flows.