How pdf-inspector Detects Columns and Layouts in Multi-Column PDFs

pdf-inspector uses a histogram-based geometric analysis in src/extractor/layout.rs to identify column boundaries by projecting text baselines onto the X-axis and analyzing gaps, then classifies layouts as newspaper or tabular.

Multi-column PDFs present unique challenges for text extraction because standard reading order can flow vertically within columns or horizontally across rows. The pdf-inspector library, developed by Firecrawl, solves this by parsing each page into geometric TextItem objects and applying specialized layout detection algorithms that preserve document structure.

The Core Algorithm: Horizontal Projection Histograms

The column detection process begins by treating text items as geometric entities with baseline coordinates rather than mere character streams.

Building the X-Axis Histogram

In src/extractor/layout.rs, the detect_columns function projects all text baselines onto the horizontal axis. The algorithm creates a histogram with default buckets of approximately 2 pt, counting how many TextItem objects occupy each horizontal position. This projection reveals the spatial distribution of text across the page width.

Identifying Column Boundaries

The algorithm identifies column separators by locating valleys in the histogram. A gap qualifies as a column boundary when it exceeds the configurable threshold COLUMN_GAP_MIN (approximately 40 pt) and is surrounded by sufficiently populated peaks. Each detected column is stored as a ColumnRegion struct containing x_min and x_max coordinates, sorted left-to-right for downstream processing.

Handling Spanning Elements and Edge Cases

Real-world PDFs contain elements that cross column boundaries, such as titles and page numbers, which require special handling to prevent false column detection.

Pre-Masking Spanning Lines

Before finalizing column assignments, detect_columns removes "spanning lines" that cross detected column gaps. The algorithm masks these lines when their height falls below a column-aware threshold and their width crosses a potential gap. This prevents wide headings from corrupting the histogram and creating false column boundaries.

ColumnRegion Struct and Assignment

Each ColumnRegion tracks its horizontal boundaries. The assign_to_best_overlap function assigns every TextItem to the column with which it shares the greatest horizontal overlap. Items spanning multiple columns are identified as either split candidates or preserved as spanning elements based on surrounding context.

Layout Classification: Newspaper vs. Tabular

After establishing column boundaries, pdf-inspector classifies the page layout type to determine proper reading order reconstruction.

The is_newspaper_layout function examines per_column_lines to distinguish between layout types:

  • Newspaper layouts: Characterized by asymmetric column density (greater than 60% of items in a single column) or configurations with side annotations. Text flows vertically within each column before moving to the next.
  • Tabular layouts: Multiple columns of comparable density where text may flow horizontally across rows. These boundaries feed into tables::mod::try_build_table_from_columns for borderless table reconstruction.

Integration in the Extraction Pipeline

The column detection integrates into the main processing workflow through process_pdf_with_options in src/lib.rs. For every page, the extractor:

  1. Builds the raw text stream into layout_items
  2. Invokes detect_columns to generate column metadata
  3. Stores results in PageInfo.pages_with_columns (defined in src/types.rs)
  4. Passes column information to downstream modules for reading-order reconstruction and table detection

Practical Implementation Examples

You can interact with the column detection logic both as a library and via command line.

Using the Rust library:

use pdf_inspector::process_pdf_with_options;
use pdf_inspector::extractor::detect_columns;

fn analyze_pdf_columns() -> Result<(), Box<dyn std::error::Error>> {
    let pdf_bytes = std::fs::read("research-paper.pdf")?;
    let mut opts = pdf_inspector::ProcessOptions::default();
    let result = process_pdf_with_options(&pdf_bytes, &mut opts)?;
    
    for (page_no, page) in result.pages.iter().enumerate() {
        // `page.layout_items` contains the TextItem baselines
        let cols = detect_columns(&page.layout_items, page_no as u32, false);
        println!("Page {}: {} columns detected", page_no + 1, cols.len());
    }
    Ok(())
}

Using the CLI:


# Extract Markdown with column-aware layout handling

pdf2md research-paper.pdf

# View detected column boundaries in JSON format

pdf2md research-paper.pdf --json | jq '.pages[].columns'

Summary

  • Geometric histogram analysis: pdf-inspector projects text baselines onto the X-axis in src/extractor/layout.rs to identify column gaps using configurable thresholds.
  • Robust edge case handling: Spanning lines and wide headings are pre-masked to prevent false column boundaries before assign_to_best_overlap distributes items into ColumnRegion structs.
  • Layout classification: The system distinguishes newspaper (vertical flow) from tabular (horizontal flow) layouts using density analysis of per_column_lines.
  • Pipeline integration: Column detection runs automatically within process_pdf_with_options, storing results in PageInfo.pages_with_columns for table reconstruction and Markdown conversion.

Frequently Asked Questions

What file contains the main column detection logic?

The primary implementation resides in src/extractor/layout.rs, which exports the detect_columns function. This file handles histogram generation, gap analysis, and the ColumnRegion struct definitions.

How does pdf-inspector handle headings that span multiple columns?

Before building the final column model, the algorithm pre-masks "spanning lines" in detect_columns. Wide headings that cross column gaps are temporarily removed from the histogram analysis if their height falls below a column-aware threshold, preventing them from creating false column boundaries while preserving them for later text assignment.

What is the difference between newspaper and tabular layout detection?

Newspaper layouts exhibit asymmetric column density (typically >60% of content in one column) and follow vertical reading order within each column. Tabular layouts show comparable density across multiple columns and potentially flow horizontally, triggering the table detection pipeline via tables::mod::try_build_table_from_columns.

Can I configure the minimum gap threshold for column detection?

Yes, the COLUMN_GAP_MIN constant (default approximately 40 pt) defines the minimum horizontal gap required to register as a column separator. While the public API in process_pdf_with_options uses sensible defaults, the underlying detect_columns function accepts parameters that influence gap detection sensitivity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →