# How pdf-inspector Handles Different PDF Layouts: Detection and Extraction Pipeline

> Discover how pdf-inspector's three-stage pipeline handles diverse PDF layouts, from text-based to scanned, for accurate data extraction and structured Markdown output.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: internals
- Published: 2026-08-04

---

**pdf-inspector handles different PDF layouts through a three-stage pipeline that classifies document types (TextBased, Scanned, ImageBased, or Mixed), detects column structures using histogram analysis and XY-cut algorithms, and applies tiered table detection strategies to extract structured Markdown regardless of visual complexity.**

The `firecrawl/pdf-inspector` repository implements a sophisticated layout analysis engine that determines the optimal extraction strategy for any PDF document. Understanding how pdf-inspector handles different PDF layouts reveals a tightly-coupled detection and processing pipeline that adapts to everything from single-column text to multi-column newspapers and complex tabular data.

## PDF Type Detection: Classifying Document Structure

Before processing any content, pdf-inspector analyzes the document to determine its fundamental type. This classification drives all subsequent layout decisions.

### The PdfType Classification System

The core classification logic resides in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), where the `PdfType` enum defines four distinct categories: **TextBased**, **Scanned**, **ImageBased**, and **Mixed**【/tmp/instagit_detl3mqu/src/detector.rs#L13-L23】. The `detect_from_document` function executes the analysis by sampling a subset of pages and examining content streams for text operators, image density, vector-drawn text patterns, and font characteristics【/tmp/instagit_detl3mqu/src/detector.rs#L81-L136】.

The detector counts unique characters, measures image density, and identifies problematic font configurations such as Identity-H/V encodings without ToUnicode CMaps or documents using Type 3 fonts exclusively. These checks determine whether standard text extraction will succeed or if OCR is required.

### Newspaper-Style Layout Detection

For dense multi-column publications, pdf-inspector employs a specialized heuristic that identifies "newspaper-style" layouts. This detection looks for very high text-operator counts combined with a low font-change-to-text-operator ratio (approximately 0.02–0.06)【/tmp/instagit_detl3mqu/src/detector.rs#L36-L50】. When these conditions match, the system flags the document as requiring OCR and prepares for interleaved column reading order processing.

### OCR Recommendations and Confidence Scoring

The detection phase outputs a `PdfTypeResult` structure containing the classification, a confidence score, and a per-page breakdown of why OCR is needed—whether due to vector text rendering, scanned images, or undecodable fonts【/tmp/instagit_detl3mqu/src/detector.rs#L44-L68】. This granular reporting allows selective OCR processing only on specific pages rather than entire documents.

## Column and Reading-Order Detection

Once classified as text-based, documents undergo layout analysis to determine proper reading order across complex column arrangements.

### Histogram-Based Column Detection

The `detect_columns` function in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) constructs a horizontal occupancy histogram using 2-point bins to identify whitespace "valleys" that represent gutters between columns【/tmp/instagit_detl3mqu/src/extractor/layout.rs#L16-L31】. To prevent false positives from titles or full-width figures, the algorithm discards spanning items wider than 60% of the page width【/tmp/instagit_detl3mqu/src/extractor/layout.rs#L60-L66】.

### Handling Justified Text with Relative-Valley Algorithm

When justified text fills potential gutters, the histogram approach fails to find clean valleys. In these cases, pdf-inspector falls back to a **relative-valley** algorithm that searches for local minima with sufficient contrast between adjacent bins【/tmp/instagit_detl3mqu/src/extractor/layout.rs#L84-L99】. This method detects subtle column divisions even when text spans nearly the entire page width.

### XY-Cut Fallback Strategy

If histogram methods prove insufficient, the system employs a simplified single-level XY-cut algorithm. This approach searches for the largest horizontal gap between item edges while enforcing minimum vertical overlap and column-item count thresholds【/tmp/instagit_detl3mqu/src/extractor/layout.rs#L22-L38】. The XY-cut serves as a robust last resort for irregular layouts that defeat statistical detection methods.

### Reading Order Logic for Multi-Column Layouts

Validated valleys convert into `ColumnRegion` objects that dictate extraction flow. Standard documents process columns left-to-right, while newspaper-classified documents use a **top-to-bottom interleaved** reading order. This distinction ensures that multi-column news articles read sequentially across rows rather than completing each column individually.

## Table Detection Across Layout Variations

pdf-inspector runs three independent table-extraction pipelines, each targeting different visual cues in the source document.

### Rect-Based Detection (Primary Strategy)

The primary detection strategy in [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) clusters axis-aligned rectangles representing table cells using a union-find algorithm. This method excels at capturing dense, grid-like tables with clear visual boundaries and serves as the first-line approach for well-structured tabular data.

### Line-Based Detection (Secondary Strategy)

When rectangle detection fails, the system falls back to [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs), which builds horizontal and vertical line grids from vector line operators. This strategy catches tables defined by ruling lines rather than background rectangles, common in financial reports and academic papers.

### Heuristic Detection (Fallback Strategy)

For loosely-structured tables without clear borders, [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) implements gap-histogram analysis combined with font-size clustering. This permissive approach identifies tabular arrangements through whitespace patterns and typographic consistency, capturing tables that lack explicit graphical boundaries.

The three strategies execute in sequence, with the first successful detection winning. This tiered approach guarantees extraction of rigidly formatted tables before attempting more speculative heuristics.

## Markdown Conversion Pipeline

After layout resolution and table extraction, the [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) module orchestrates final output generation. The converter walks the ordered list of `TextLine` objects, applying content classification (headers, lists, code blocks, captions) and post-processing fixes for hyphenation, dot-leaders, and URL reconstruction. This final stage respects the column reading order and table structures established in previous phases.

## Practical Examples

Detect document type and identify pages requiring OCR:

```bash
pdf2md --json --detect-type sample.pdf

# → { "pdf_type":"Mixed", "ocr_recommended":true, "pages_needing_ocr":[1,5,12] }

```

Extract full markdown with automatic column and table handling:

```bash
pdf2md sample.pdf > output.md

```

Run only the column detector for layout debugging:

```bash
cargo run --bin detect-pdf -- --columns sample.pdf

```

Programmatically access type detection in Rust:

```rust
use pdf_inspector::detector::{detect_pdf_type, PdfTypeResult};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result: PdfTypeResult = detect_pdf_type("sample.pdf")?;
    println!("PDF type: {:?}, OCR needed on pages {:?}", result.pdf_type, result.pages_needing_ocr);
    Ok(())
}

```

## Summary

- **pdf-inspector** classifies PDFs into TextBased, Scanned, ImageBased, or Mixed types using [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) before extraction begins.
- The **newspaper heuristic** identifies dense multi-column layouts through text-operator density and font-change ratios, triggering specialized reading-order logic.
- **Column detection** combines histogram analysis, relative-valley algorithms, and XY-cut fallbacks to handle everything from simple single-column text to complex multi-column magazines.
- **Three-tiered table detection** (rect-based, line-based, heuristic) ensures extraction of tabular data regardless of border visibility or formatting quality.
- The pipeline outputs structured Markdown while preserving document semantics through [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs).

## Frequently Asked Questions

### How does pdf-inspector determine if a PDF needs OCR?

The `detect_from_document` function in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) analyzes sampled pages for image density, undecodable fonts (Identity-H/V without ToUnicode CMaps), Type 3 font usage, and vector-drawn text patterns【/tmp/instagit_detl3mqu/src/detector.rs#L81-L136】. Pages lacking extractable text operators or containing high image-to-text ratios are flagged in the `PdfTypeResult.pages_needing_ocr` list, enabling selective OCR processing only where necessary.

### What is the "newspaper-style" heuristic in pdf-inspector?

This heuristic detects dense multi-column publications by measuring the ratio of font-change operators to total text operators. When text-operator counts are extremely high but the font-change ratio falls between 0.02 and 0.06, the system classifies the document as newspaper-style【/tmp/instagit_detl3mqu/src/detector.rs#L36-L50】. This triggers both OCR recommendations and interleaved top-to-bottom reading order for column extraction.

### How does pdf-inspector handle tables without visible borders?

When rect-based and line-based detection fail to find graphical boundaries, the system falls back to the heuristic detector in [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs). This analyzer constructs gap histograms and examines font-size consistency to identify loosely-structured tables that rely on whitespace alignment rather than ruling lines or cell backgrounds.

### Can pdf-inspector process mixed PDFs containing both text and scanned images?

Yes. The **Mixed** `PdfType` classification specifically handles documents combining extractable text pages with scanned image pages. The detector outputs per-page OCR recommendations, allowing the extraction pipeline to process text pages natively while routing only specific pages (such as those containing embedded scanned images or undecodable fonts) through OCR engines【/tmp/instagit_detl3mqu/src/detector.rs#L44-L68】.