How pdf-inspector Parses PDF Files: A Technical Deep Dive into the Rust Pipeline
pdf-inspector parses PDF files through a three-stage Rust pipeline—type detection, low-level content-stream extraction with lopdf, and sophisticated layout post-processing—to transform binary documents into structured text items, rectangles, and Markdown output.
The firecrawl/pdf-inspector repository implements a complete PDF parsing engine that bridges the gap between raw binary content and structured data. Understanding how pdf-inspector parses PDF files reveals a state-machine architecture capable of handling complex font encodings, graphics transformations, and multi-column layouts. The parser processes documents through distinct detection, extraction, and analysis phases to produce clean, machine-readable output.
Stage 1: PDF Type Detection and Loading (src/detector.rs)
Before extraction begins, pdf-inspector classifies the document to determine processing strategy. The detect_pdf_type function in src/detector.rs loads the file using lopdf into a Document struct, then runs a heuristic analysis on sampled pages.
The detection routine scans content streams via analyze_page_content to count text operators (Tj, TJ), image references (Do), and vector outlines. It specifically checks for problematic fonts—such as Identity-H/V fonts lacking ToUnicode CMaps or Type 3 fonts without Unicode mappings—and calculates text-to-image ratios. Based on thresholds defined in DetectionConfig, it returns a PdfTypeResult classifying the PDF as TextBased, Scanned, ImageBased, or Mixed, along with per-page OCR recommendations.
pub fn detect_pdf_type<P: AsRef<Path>>(path: P) -> Result<PdfTypeResult, PdfError> {
// 1️⃣ Validate the file
crate::validate_pdf_file(&path)?;
// 2️⃣ Load once (shared for detection & extraction)
let (doc, page_count) = crate::load_document_from_path(&path)?;
// 3️⃣ Run heuristic analysis on a sampled set of pages
detect_from_document(&doc, page_count, &DetectionConfig::default())
}
Stage 2: Content-Stream Parsing and Text Extraction (src/extractor/content_stream.rs)
The core parsing logic resides in extract_page_text_items within src/extractor/content_stream.rs. This function implements a state-machine that walks every page’s content stream, tracking the graphics state—including the Current Transformation Matrix (CTM), text matrix, font, and spacing parameters.
The parser strips PDF comments from raw bytes, decodes the stream into operators using Content::decode, and iterates through operations:
- Text operators (
Tj,TJ,',"): Decode operands using the font’s ToUnicode CMap or fallback strategies and emitTextItemstructs with device-space coordinates. - State operators (
BT,ET,Tf,Td,TD): Manage text block boundaries, font selection, and positioning. - Drawing operators (
re, path construction): Collect rectangles and lines for table detection. - XObjects (
Do): Handle Image XObjects (emitting placeholders with bounding boxes computed byimage_bbox_from_ctm) and recursively process Form XObjects viaextract_form_xobject_text.
pub(crate) fn extract_page_text_items(
doc: &Document,
page_id: ObjectId,
page_num: u32,
font_cmaps: &FontCMaps,
include_invisible: bool,
style_cache: &mut FontStyleCache,
) -> Result<(PageExtraction, bool, bool), PdfError> {
// 1️⃣ Get the raw content bytes and strip PDF comments
let content_data = doc.get_page_content(page_id)?;
let content_data = strip_pdf_comments(&content_data);
// 2️⃣ Decode the stream into a list of operators
let content = Content::decode(&content_data)?;
// 3️⃣ Iterate over all operators
for op in &content.operations {
match op.operator.as_str() {
"BT" => { /* begin text block */ }
"ET" => { /* end text block */ }
"Tf" => { /* set current font */ }
"Tj" => {
// Decode operand → Unicode text using ToUnicode or fallback
if let Some(text) = extract_text_from_operand(...) {
items.push(TextItem { … });
}
}
"Do" => { /* image or form XObject */ }
"re" => { /* rectangle – collect for table detection */ }
// … other operators update graphics state
}
}
Ok(((items, rects, lines), has_gid_fonts, coords_rotated))
}
Font Handling and Unicode Decoding
Font decoding occurs through the tounicode module (src/tounicode.rs). The build_font_encodings function constructs per-font Encoding objects from PDF Differences arrays. For Unicode mapping, the parser first attempts ToUnicode CMap parsing; if unavailable, it falls back to CID-to-Unicode heuristics or embedded TrueType/OpenType cmap tables. This ensures accurate text extraction even from documents with legacy or malformed font encodings.
Graphics State and Coordinate Transformation
The parser maintains precise positional data through matrix operations. Functions like multiply_matrices and rise_adjusted in src/extractor/mod.rs apply the CTM and text rise (Ts) values to transform glyph positions into device-space coordinates stored in each TextItem.
Marked-Content and ActualText
The state-machine captures marked-content blocks (BDC/EMC) to extract ActualText entries—invisible text layers often embedded in scanned PDFs—without rendering their glyphs, preserving semantic content while ignoring display artifacts.
Stage 3: Post-Processing and Layout Analysis (src/extractor/mod.rs)
After raw extraction, extract_positioned_text_impl orchestrates post-processing to reconstruct logical document structure:
- Clipping: Items are clipped to the page’s CropBox or MediaBox using
get_page_box. - Letter-spacing correction:
fix_letterspaced_itemsrepairs artificial spacing introduced by PDF tracking. - Text merging:
merge_text_itemscombines adjacent glyphs into words while respecting font size and style;merge_subscript_itemshandles chemical formulas and footnote markers. - Layout detection: Functions in
src/extractor/layout.rsidentify columns (detect_columns), newspaper layouts (is_newspaper_layout), and table structures using thetablesmodule’s rectangle-, line-, and heuristic-based detectors.
The final output populates a PdfProcessResult struct containing the detected PDF type, optional Markdown representation, processing metadata, and lists of pages requiring OCR.
Working with the pdf-inspector API
The public API in src/lib.rs exposes high-level functions for common use cases:
use pdf_inspector::{process_pdf, extract_text, extract_text_with_positions};
fn main() -> Result<(), pdf_inspector::PdfError> {
// Full processing: type detection → extraction → Markdown
let result = process_pdf("reports/annual_report.pdf")?;
println!("Detected type: {:?} (confidence {:.2})", result.pdf_type, result.confidence);
if let Some(md) = result.markdown {
println!("--- Markdown output ---\n{}", md);
}
// Fast detection only (no text extraction)
let info = pdf_inspector::detect_pdf("scanned/document.pdf")?;
println!("PDF type: {:?}, pages needing OCR: {:?}", info.pdf_type, info.pages_needing_ocr);
// Get raw positioned text for custom layout work
let items = extract_text_with_positions("tables/mixed_layout.pdf")?;
for item in items.iter().take(5) {
println!("Page {} – ({:.1},{:.1}) '{}' [{:.1}pt]",
item.page, item.x, item.y, item.text, item.font_size);
}
Ok(())
}
Key source files implementing this pipeline include:
src/detector.rs: PDF type classification and OCR recommendation logic.src/extractor/content_stream.rs: State-machine parser for PDF content streams.src/extractor/mod.rs: Orchestration, clipping, text merging, and layout helpers.src/tounicode.rs: ToUnicode CMap parsing and font fallback strategies.src/extractor/layout.rs: Column, newspaper, and table detection algorithms.src/types.rs: Core data structures includingTextItem,PdfRect, andPdfProcessResult.
Summary
- pdf-inspector uses a three-stage pipeline: type detection via
detect_pdf_type, content extraction viaextract_page_text_items, and layout reconstruction through post-processing functions. - The parser relies on
lopdffor document loading but implements a custom state-machine to handle PDF operators, graphics state matrices, and font encoding complexities. - Advanced features include ToUnicode CMap parsing with fallback heuristics, recursive Form XObject processing, and sophisticated layout detection for tables and columns.
- The API provides both high-level Markdown generation via
process_pdfand low-level positioned text access viaextract_text_with_positionsfor custom processing needs.
Frequently Asked Questions
What Rust library does pdf-inspector use to read PDF files?
pdf-inspector uses the lopdf crate to load PDF files into a Document struct and access page content streams. However, it implements a custom state-machine parser in src/extractor/content_stream.rs to walk operators and extract text, rather than relying solely on lopdf's built-in text extraction.
How does pdf-inspector handle fonts without Unicode mappings?
When encountering fonts lacking ToUnicode CMaps—such as Identity-H/V encodings or Type 3 fonts—the parser in src/tounicode.rs applies fallback strategies including CID-to-Unicode heuristics and embedded TrueType/OpenType cmap table analysis. If decoding fails, the system flags has_encoding_issues in the results.
Can pdf-inspector detect whether a PDF requires OCR processing?
Yes, the detect_pdf_type function analyzes content streams for text operator density versus image presence and checks for problematic font configurations. It classifies documents as TextBased, Scanned, ImageBased, or Mixed, returning a pages_needing_ocr vector indicating which pages lack extractable text layers.
How does pdf-inspector reconstruct tables and multi-column layouts?
After extracting raw text items, pdf-inspector runs post-processing functions including merge_text_items for word reconstruction and detect_columns for layout analysis. The src/extractor/layout.rs module and src/tables/ subdirectory implement rectangle detection, line analysis, and heuristic algorithms to identify table structures and newspaper-style column flows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →