How pdf-inspector Handles Different PDF Layouts: Detection and Extraction Pipeline
pdf-inspector handles different PDF layouts through a three-stage pipeline that classifies document types (TextBased, Scanned, ImageBased, or Mixed), detects column structures using histogram analysis and XY-cut algorithms, and applies tiered table detection strategies to extract structured Markdown regardless of visual complexity.
The firecrawl/pdf-inspector repository implements a sophisticated layout analysis engine that determines the optimal extraction strategy for any PDF document. Understanding how pdf-inspector handles different PDF layouts reveals a tightly-coupled detection and processing pipeline that adapts to everything from single-column text to multi-column newspapers and complex tabular data.
PDF Type Detection: Classifying Document Structure
Before processing any content, pdf-inspector analyzes the document to determine its fundamental type. This classification drives all subsequent layout decisions.
The PdfType Classification System
The core classification logic resides in src/detector.rs, where the PdfType enum defines four distinct categories: TextBased, Scanned, ImageBased, and Mixed【/tmp/instagit_detl3mqu/src/detector.rs#L13-L23】. The detect_from_document function executes the analysis by sampling a subset of pages and examining content streams for text operators, image density, vector-drawn text patterns, and font characteristics【/tmp/instagit_detl3mqu/src/detector.rs#L81-L136】.
The detector counts unique characters, measures image density, and identifies problematic font configurations such as Identity-H/V encodings without ToUnicode CMaps or documents using Type 3 fonts exclusively. These checks determine whether standard text extraction will succeed or if OCR is required.
Newspaper-Style Layout Detection
For dense multi-column publications, pdf-inspector employs a specialized heuristic that identifies "newspaper-style" layouts. This detection looks for very high text-operator counts combined with a low font-change-to-text-operator ratio (approximately 0.02–0.06)【/tmp/instagit_detl3mqu/src/detector.rs#L36-L50】. When these conditions match, the system flags the document as requiring OCR and prepares for interleaved column reading order processing.
OCR Recommendations and Confidence Scoring
The detection phase outputs a PdfTypeResult structure containing the classification, a confidence score, and a per-page breakdown of why OCR is needed—whether due to vector text rendering, scanned images, or undecodable fonts【/tmp/instagit_detl3mqu/src/detector.rs#L44-L68】. This granular reporting allows selective OCR processing only on specific pages rather than entire documents.
Column and Reading-Order Detection
Once classified as text-based, documents undergo layout analysis to determine proper reading order across complex column arrangements.
Histogram-Based Column Detection
The detect_columns function in src/extractor/layout.rs constructs a horizontal occupancy histogram using 2-point bins to identify whitespace "valleys" that represent gutters between columns【/tmp/instagit_detl3mqu/src/extractor/layout.rs#L16-L31】. To prevent false positives from titles or full-width figures, the algorithm discards spanning items wider than 60% of the page width【/tmp/instagit_detl3mqu/src/extractor/layout.rs#L60-L66】.
Handling Justified Text with Relative-Valley Algorithm
When justified text fills potential gutters, the histogram approach fails to find clean valleys. In these cases, pdf-inspector falls back to a relative-valley algorithm that searches for local minima with sufficient contrast between adjacent bins【/tmp/instagit_detl3mqu/src/extractor/layout.rs#L84-L99】. This method detects subtle column divisions even when text spans nearly the entire page width.
XY-Cut Fallback Strategy
If histogram methods prove insufficient, the system employs a simplified single-level XY-cut algorithm. This approach searches for the largest horizontal gap between item edges while enforcing minimum vertical overlap and column-item count thresholds【/tmp/instagit_detl3mqu/src/extractor/layout.rs#L22-L38】. The XY-cut serves as a robust last resort for irregular layouts that defeat statistical detection methods.
Reading Order Logic for Multi-Column Layouts
Validated valleys convert into ColumnRegion objects that dictate extraction flow. Standard documents process columns left-to-right, while newspaper-classified documents use a top-to-bottom interleaved reading order. This distinction ensures that multi-column news articles read sequentially across rows rather than completing each column individually.
Table Detection Across Layout Variations
pdf-inspector runs three independent table-extraction pipelines, each targeting different visual cues in the source document.
Rect-Based Detection (Primary Strategy)
The primary detection strategy in src/tables/detect_rects.rs clusters axis-aligned rectangles representing table cells using a union-find algorithm. This method excels at capturing dense, grid-like tables with clear visual boundaries and serves as the first-line approach for well-structured tabular data.
Line-Based Detection (Secondary Strategy)
When rectangle detection fails, the system falls back to src/tables/detect_lines.rs, which builds horizontal and vertical line grids from vector line operators. This strategy catches tables defined by ruling lines rather than background rectangles, common in financial reports and academic papers.
Heuristic Detection (Fallback Strategy)
For loosely-structured tables without clear borders, src/tables/detect_heuristic.rs implements gap-histogram analysis combined with font-size clustering. This permissive approach identifies tabular arrangements through whitespace patterns and typographic consistency, capturing tables that lack explicit graphical boundaries.
The three strategies execute in sequence, with the first successful detection winning. This tiered approach guarantees extraction of rigidly formatted tables before attempting more speculative heuristics.
Markdown Conversion Pipeline
After layout resolution and table extraction, the src/markdown/convert.rs module orchestrates final output generation. The converter walks the ordered list of TextLine objects, applying content classification (headers, lists, code blocks, captions) and post-processing fixes for hyphenation, dot-leaders, and URL reconstruction. This final stage respects the column reading order and table structures established in previous phases.
Practical Examples
Detect document type and identify pages requiring OCR:
pdf2md --json --detect-type sample.pdf
# → { "pdf_type":"Mixed", "ocr_recommended":true, "pages_needing_ocr":[1,5,12] }
Extract full markdown with automatic column and table handling:
pdf2md sample.pdf > output.md
Run only the column detector for layout debugging:
cargo run --bin detect-pdf -- --columns sample.pdf
Programmatically access type detection in Rust:
use pdf_inspector::detector::{detect_pdf_type, PdfTypeResult};
fn main() -> Result<(), pdf_inspector::PdfError> {
let result: PdfTypeResult = detect_pdf_type("sample.pdf")?;
println!("PDF type: {:?}, OCR needed on pages {:?}", result.pdf_type, result.pages_needing_ocr);
Ok(())
}
Summary
- pdf-inspector classifies PDFs into TextBased, Scanned, ImageBased, or Mixed types using
src/detector.rsbefore extraction begins. - The newspaper heuristic identifies dense multi-column layouts through text-operator density and font-change ratios, triggering specialized reading-order logic.
- Column detection combines histogram analysis, relative-valley algorithms, and XY-cut fallbacks to handle everything from simple single-column text to complex multi-column magazines.
- Three-tiered table detection (rect-based, line-based, heuristic) ensures extraction of tabular data regardless of border visibility or formatting quality.
- The pipeline outputs structured Markdown while preserving document semantics through
src/markdown/convert.rs.
Frequently Asked Questions
How does pdf-inspector determine if a PDF needs OCR?
The detect_from_document function in src/detector.rs analyzes sampled pages for image density, undecodable fonts (Identity-H/V without ToUnicode CMaps), Type 3 font usage, and vector-drawn text patterns【/tmp/instagit_detl3mqu/src/detector.rs#L81-L136】. Pages lacking extractable text operators or containing high image-to-text ratios are flagged in the PdfTypeResult.pages_needing_ocr list, enabling selective OCR processing only where necessary.
What is the "newspaper-style" heuristic in pdf-inspector?
This heuristic detects dense multi-column publications by measuring the ratio of font-change operators to total text operators. When text-operator counts are extremely high but the font-change ratio falls between 0.02 and 0.06, the system classifies the document as newspaper-style【/tmp/instagit_detl3mqu/src/detector.rs#L36-L50】. This triggers both OCR recommendations and interleaved top-to-bottom reading order for column extraction.
How does pdf-inspector handle tables without visible borders?
When rect-based and line-based detection fail to find graphical boundaries, the system falls back to the heuristic detector in src/tables/detect_heuristic.rs. This analyzer constructs gap histograms and examines font-size consistency to identify loosely-structured tables that rely on whitespace alignment rather than ruling lines or cell backgrounds.
Can pdf-inspector process mixed PDFs containing both text and scanned images?
Yes. The Mixed PdfType classification specifically handles documents combining extractable text pages with scanned image pages. The detector outputs per-page OCR recommendations, allowing the extraction pipeline to process text pages natively while routing only specific pages (such as those containing embedded scanned images or undecodable fonts) through OCR engines【/tmp/instagit_detl3mqu/src/detector.rs#L44-L68】.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →