How pdf-inspector Parses PDF Files: A Technical Deep Dive into the Rust Pipeline

pdf-inspector parses PDF files through a three-stage Rust pipeline—type detection, low-level content-stream extraction with lopdf, and sophisticated layout post-processing—to transform binary documents into structured text items, rectangles, and Markdown output.

The firecrawl/pdf-inspector repository implements a complete PDF parsing engine that bridges the gap between raw binary content and structured data. Understanding how pdf-inspector parses PDF files reveals a state-machine architecture capable of handling complex font encodings, graphics transformations, and multi-column layouts. The parser processes documents through distinct detection, extraction, and analysis phases to produce clean, machine-readable output.

Stage 1: PDF Type Detection and Loading (src/detector.rs)

Before extraction begins, pdf-inspector classifies the document to determine processing strategy. The detect_pdf_type function in src/detector.rs loads the file using lopdf into a Document struct, then runs a heuristic analysis on sampled pages.

The detection routine scans content streams via analyze_page_content to count text operators (Tj, TJ), image references (Do), and vector outlines. It specifically checks for problematic fonts—such as Identity-H/V fonts lacking ToUnicode CMaps or Type 3 fonts without Unicode mappings—and calculates text-to-image ratios. Based on thresholds defined in DetectionConfig, it returns a PdfTypeResult classifying the PDF as TextBased, Scanned, ImageBased, or Mixed, along with per-page OCR recommendations.

pub fn detect_pdf_type<P: AsRef<Path>>(path: P) -> Result<PdfTypeResult, PdfError> {
    // 1️⃣ Validate the file
    crate::validate_pdf_file(&path)?;
    // 2️⃣ Load once (shared for detection & extraction)
    let (doc, page_count) = crate::load_document_from_path(&path)?;
    // 3️⃣ Run heuristic analysis on a sampled set of pages
    detect_from_document(&doc, page_count, &DetectionConfig::default())
}

Stage 2: Content-Stream Parsing and Text Extraction (src/extractor/content_stream.rs)

The core parsing logic resides in extract_page_text_items within src/extractor/content_stream.rs. This function implements a state-machine that walks every page’s content stream, tracking the graphics state—including the Current Transformation Matrix (CTM), text matrix, font, and spacing parameters.

The parser strips PDF comments from raw bytes, decodes the stream into operators using Content::decode, and iterates through operations:

  • Text operators (Tj, TJ, ', "): Decode operands using the font’s ToUnicode CMap or fallback strategies and emit TextItem structs with device-space coordinates.
  • State operators (BT, ET, Tf, Td, TD): Manage text block boundaries, font selection, and positioning.
  • Drawing operators (re, path construction): Collect rectangles and lines for table detection.
  • XObjects (Do): Handle Image XObjects (emitting placeholders with bounding boxes computed by image_bbox_from_ctm) and recursively process Form XObjects via extract_form_xobject_text.
pub(crate) fn extract_page_text_items(
    doc: &Document,
    page_id: ObjectId,
    page_num: u32,
    font_cmaps: &FontCMaps,
    include_invisible: bool,
    style_cache: &mut FontStyleCache,
) -> Result<(PageExtraction, bool, bool), PdfError> {
    // 1️⃣ Get the raw content bytes and strip PDF comments
    let content_data = doc.get_page_content(page_id)?;
    let content_data = strip_pdf_comments(&content_data);

    // 2️⃣ Decode the stream into a list of operators
    let content = Content::decode(&content_data)?;

    // 3️⃣ Iterate over all operators
    for op in &content.operations {
        match op.operator.as_str() {
            "BT" => { /* begin text block */ }
            "ET" => { /* end text block */ }
            "Tf" => { /* set current font */ }
            "Tj" => {
                // Decode operand → Unicode text using ToUnicode or fallback
                if let Some(text) = extract_text_from_operand(...) {
                    items.push(TextItem { … });
                }
            }
            "Do" => { /* image or form XObject */ }
            "re" => { /* rectangle – collect for table detection */ }
            // … other operators update graphics state
        }
    }
    Ok(((items, rects, lines), has_gid_fonts, coords_rotated))
}

Font Handling and Unicode Decoding

Font decoding occurs through the tounicode module (src/tounicode.rs). The build_font_encodings function constructs per-font Encoding objects from PDF Differences arrays. For Unicode mapping, the parser first attempts ToUnicode CMap parsing; if unavailable, it falls back to CID-to-Unicode heuristics or embedded TrueType/OpenType cmap tables. This ensures accurate text extraction even from documents with legacy or malformed font encodings.

Graphics State and Coordinate Transformation

The parser maintains precise positional data through matrix operations. Functions like multiply_matrices and rise_adjusted in src/extractor/mod.rs apply the CTM and text rise (Ts) values to transform glyph positions into device-space coordinates stored in each TextItem.

Marked-Content and ActualText

The state-machine captures marked-content blocks (BDC/EMC) to extract ActualText entries—invisible text layers often embedded in scanned PDFs—without rendering their glyphs, preserving semantic content while ignoring display artifacts.

Stage 3: Post-Processing and Layout Analysis (src/extractor/mod.rs)

After raw extraction, extract_positioned_text_impl orchestrates post-processing to reconstruct logical document structure:

  • Clipping: Items are clipped to the page’s CropBox or MediaBox using get_page_box.
  • Letter-spacing correction: fix_letterspaced_items repairs artificial spacing introduced by PDF tracking.
  • Text merging: merge_text_items combines adjacent glyphs into words while respecting font size and style; merge_subscript_items handles chemical formulas and footnote markers.
  • Layout detection: Functions in src/extractor/layout.rs identify columns (detect_columns), newspaper layouts (is_newspaper_layout), and table structures using the tables module’s rectangle-, line-, and heuristic-based detectors.

The final output populates a PdfProcessResult struct containing the detected PDF type, optional Markdown representation, processing metadata, and lists of pages requiring OCR.

Working with the pdf-inspector API

The public API in src/lib.rs exposes high-level functions for common use cases:

use pdf_inspector::{process_pdf, extract_text, extract_text_with_positions};

fn main() -> Result<(), pdf_inspector::PdfError> {
    // Full processing: type detection → extraction → Markdown
    let result = process_pdf("reports/annual_report.pdf")?;
    println!("Detected type: {:?} (confidence {:.2})", result.pdf_type, result.confidence);
    if let Some(md) = result.markdown {
        println!("--- Markdown output ---\n{}", md);
    }

    // Fast detection only (no text extraction)
    let info = pdf_inspector::detect_pdf("scanned/document.pdf")?;
    println!("PDF type: {:?}, pages needing OCR: {:?}", info.pdf_type, info.pages_needing_ocr);

    // Get raw positioned text for custom layout work
    let items = extract_text_with_positions("tables/mixed_layout.pdf")?;
    for item in items.iter().take(5) {
        println!("Page {} – ({:.1},{:.1}) '{}' [{:.1}pt]",
                 item.page, item.x, item.y, item.text, item.font_size);
    }

    Ok(())
}

Key source files implementing this pipeline include:

Summary

  • pdf-inspector uses a three-stage pipeline: type detection via detect_pdf_type, content extraction via extract_page_text_items, and layout reconstruction through post-processing functions.
  • The parser relies on lopdf for document loading but implements a custom state-machine to handle PDF operators, graphics state matrices, and font encoding complexities.
  • Advanced features include ToUnicode CMap parsing with fallback heuristics, recursive Form XObject processing, and sophisticated layout detection for tables and columns.
  • The API provides both high-level Markdown generation via process_pdf and low-level positioned text access via extract_text_with_positions for custom processing needs.

Frequently Asked Questions

What Rust library does pdf-inspector use to read PDF files?

pdf-inspector uses the lopdf crate to load PDF files into a Document struct and access page content streams. However, it implements a custom state-machine parser in src/extractor/content_stream.rs to walk operators and extract text, rather than relying solely on lopdf's built-in text extraction.

How does pdf-inspector handle fonts without Unicode mappings?

When encountering fonts lacking ToUnicode CMaps—such as Identity-H/V encodings or Type 3 fonts—the parser in src/tounicode.rs applies fallback strategies including CID-to-Unicode heuristics and embedded TrueType/OpenType cmap table analysis. If decoding fails, the system flags has_encoding_issues in the results.

Can pdf-inspector detect whether a PDF requires OCR processing?

Yes, the detect_pdf_type function analyzes content streams for text operator density versus image presence and checks for problematic font configurations. It classifies documents as TextBased, Scanned, ImageBased, or Mixed, returning a pages_needing_ocr vector indicating which pages lack extractable text layers.

How does pdf-inspector reconstruct tables and multi-column layouts?

After extracting raw text items, pdf-inspector runs post-processing functions including merge_text_items for word reconstruction and detect_columns for layout analysis. The src/extractor/layout.rs module and src/tables/ subdirectory implement rectangle detection, line analysis, and heuristic algorithms to identify table structures and newspaper-style column flows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →