How PDF-Inspector Detects Newspaper-Style Layouts in PDF Documents

Newspaper-style layout detection in PDF-Inspector analyzes text flow geometry across columns to distinguish independent prose flows from tabular data, using line density ratios and vertical collision tests in src/extractor/layout.rs.

The firecrawl/pdf-inspector repository implements sophisticated document understanding logic to extract structured content from PDF files. When processing multi-column documents, accurate newspaper-style layout detection ensures that side-by-side text blocks are read in the correct order—either sequentially down each column or interleaved row-by-row for tables.

The Core Algorithm in is_newspaper_layout

The detection logic resides in the private function is_newspaper_layout within src/extractor/layout.rs. This function receives grouped text lines per column and evaluates four distinct stages to classify the layout.

Stage 1: Basic Sanity Checks

Before analysis begins, the algorithm verifies that the page contains genuine multi-column content. The code requires at least two columns and a minimum line density to make reliable decisions.

Specifically, the function checks that per_column_lines.len() >= 2 and that every column contains at least 5 lines. If either condition fails, the function immediately returns false, defaulting to tabular processing.

if per_column_lines.len() < 2 { return false; }
let min_lines = …; let max_lines = …;
if min_lines < 5 { return false; }

Stage 2: Sidebar Detection

When a page contains exactly two columns with fewer than 15 lines each, the algorithm tests for a sidebar pattern common in academic papers and reports. Sidebars contain marginal notes or annotations that should be treated as independent prose flows rather than table cells.

The detection requires five simultaneous conditions:

  • Width ratio of the narrower to wider column must be less than 0.5
  • Line balance (min_lines / max_lines) must be less than 0.35
  • The wider column must contain at least 20 lines
  • The narrow column must be at least 160 pt wide
  • The average Y-gap of the narrow column must be at least 2.5× the gap of the wide column
if columns.len() == 2 && per_column_lines.len() == 2 {
    let width_ratio = w0.min(w1) / w0.max(w1);
    let line_balance = min_lines as f32 / max_lines as f32;
    if width_ratio < 0.50 && line_balance < 0.35 && max_lines >= 20 && narrow_width >= 160.0 {
        // compare avg Y‑gaps …
        if narrow_gap / wide_gap >= 2.5 { return true; }
    }
}

Stage 3: Balanced Dense Columns

For columns containing substantial text (15 or more lines each), the algorithm applies a simple balance ratio test. If min_lines / max_lines > 0.7, the layout is classified as newspaper regardless of vertical alignment. This reflects the typical newspaper layout where columns contain comparable amounts of text flowing independently.

let balance_ratio = min_lines as f32 / max_lines as f32;
if balance_ratio > 0.7 { return true; }

Stage 4: Y-Collision Fallback

When columns are unbalanced or fail the density test, the algorithm performs a vertical collision analysis on the smallest column. For each line in this column, it searches other columns for lines whose baselines fall within ±5 pt (y_tol = 5.0).

If more than 50% of lines in the smallest column vertically align with lines in other columns (collisions / smallest.len() > 0.5), the layout is deemed tabular. Conversely, a low collision ratio indicates independent prose columns typical of newspaper layouts.

let y_tol = 5.0;
for line in smallest {
    for (ci, col) in per_column_lines.iter().enumerate() { … }
}
let ratio = collisions as f32 / smallest.len() as f32;
ratio > 0.5

Pipeline Integration and Reading Order

The newspaper detection integrates into the extraction pipeline at multiple points in src/extractor/mod.rs and src/tables/mod.rs. After detect_columns builds horizontal occupancy histograms and group_into_lines organizes text items, the system calls is_newspaper_layout to determine reading order.

When the function returns true, the extractor processes columns sequentially (complete column 1, then column 2). When false, it applies Y-interleaved ordering suitable for tables. The src/markdown/convert.rs module consumes this decision to emit correctly ordered Markdown output.

Practical Implementation Example

The following Rust pattern demonstrates how to invoke the detection logic after column extraction:

use pdf_inspector::extractor::{detect_columns, group_into_lines, is_newspaper_layout};

// 1. Extract raw text items from the PDF (already available as `items`).
let lines = group_into_lines(items.clone());

// 2. Detect column regions on each page.
let columns = detect_columns(&items, page_number, page_has_table);

// 3. Split lines per column (the function `split_column_stragglers` helps here).
let per_column_lines = split_lines_by_columns(&lines, &columns);

// 4. Ask the layout engine whether the page is newspaper style.
if is_newspaper_layout(&per_column_lines, &columns) {
    println!("Read column 1 → column 2 (newspaper order)");
} else {
    println!("Interleave columns by Y (tabular order)");
}

Note that split_lines_by_columns represents internal logic similar to the extractor's grouping mechanism that feeds is_newspaper_layout.

Summary

  • Newspaper-style layout detection in firecrawl/pdf-inspector resides in src/extractor/layout.rs within the is_newspaper_layout function.
  • The algorithm requires at least two columns with five lines each before analysis begins.
  • Sidebar detection handles asymmetrical two-column layouts common in academic papers by comparing width ratios and Y-gaps.
  • Balanced dense columns (15+ lines with >0.7 line ratio) automatically classify as newspaper layouts.
  • The Y-collision test uses a ±5 pt tolerance to detect vertically aligned text indicative of tables, with a 0.5 threshold determining the final classification.
  • Detection results drive reading order decisions in src/markdown/convert.rs, ensuring proper sequential or interleaved text extraction.

Frequently Asked Questions

What is the difference between newspaper-style and tabular layouts in PDFs?

Newspaper-style layouts contain independent prose flows where text reads sequentially down one column before continuing to the next. Tabular layouts organize content in rows where horizontally adjacent cells relate to each other, requiring interleaved reading order. PDF-Inspector distinguishes these by analyzing vertical alignment patterns and line density ratios.

How does the Y-collision test distinguish tables from newspaper columns?

The Y-collision test checks if lines in different columns share the same vertical baseline within ±5 points. Tables exhibit high vertical alignment between columns (many collisions), while newspaper columns show independent vertical flow (few collisions). When collisions exceed 50% of lines, the layout is classified as tabular.

What are the minimum requirements for a layout to be considered newspaper-style?

The layout must contain at least two columns, each with a minimum of five lines of text. Without these thresholds, the algorithm returns false and defaults to tabular processing. Additional requirements vary by detection stage, such as the 15-line minimum for balanced column tests or specific width ratios for sidebar detection.

Where is the newspaper detection logic used in the extraction pipeline?

The is_newspaper_layout function is called after column detection and line grouping in src/extractor/mod.rs. Its boolean result determines whether src/markdown/convert.rs processes columns sequentially or applies Y-interleaved ordering. Complementary document-level heuristics also exist in src/detector.rs for OCR recommendations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →