# How PDF-Inspector Detects Newspaper-Style Layouts in PDF Documents

> Discover how PDF-Inspector uses geometry, density ratios, and collision tests to detect newspaper-style layouts in PDFs. Understand text flow analysis for accurate document parsing.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: internals
- Published: 2026-08-08

---

**Newspaper-style layout detection in PDF-Inspector analyzes text flow geometry across columns to distinguish independent prose flows from tabular data, using line density ratios and vertical collision tests in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs).**

The `firecrawl/pdf-inspector` repository implements sophisticated document understanding logic to extract structured content from PDF files. When processing multi-column documents, accurate **newspaper-style layout detection** ensures that side-by-side text blocks are read in the correct order—either sequentially down each column or interleaved row-by-row for tables.

## The Core Algorithm in `is_newspaper_layout`

The detection logic resides in the private function `is_newspaper_layout` within [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs). This function receives grouped text lines per column and evaluates four distinct stages to classify the layout.

### Stage 1: Basic Sanity Checks

Before analysis begins, the algorithm verifies that the page contains genuine multi-column content. The code requires at least two columns and a minimum line density to make reliable decisions.

Specifically, the function checks that `per_column_lines.len() >= 2` and that every column contains at least **5 lines**. If either condition fails, the function immediately returns `false`, defaulting to tabular processing.

```rust
if per_column_lines.len() < 2 { return false; }
let min_lines = …; let max_lines = …;
if min_lines < 5 { return false; }

```

### Stage 2: Sidebar Detection

When a page contains exactly two columns with fewer than 15 lines each, the algorithm tests for a **sidebar pattern** common in academic papers and reports. Sidebars contain marginal notes or annotations that should be treated as independent prose flows rather than table cells.

The detection requires five simultaneous conditions:

- **Width ratio** of the narrower to wider column must be less than 0.5
- **Line balance** (`min_lines / max_lines`) must be less than 0.35
- The wider column must contain at least 20 lines
- The narrow column must be at least 160 pt wide
- The average Y-gap of the narrow column must be at least **2.5×** the gap of the wide column

```rust
if columns.len() == 2 && per_column_lines.len() == 2 {
    let width_ratio = w0.min(w1) / w0.max(w1);
    let line_balance = min_lines as f32 / max_lines as f32;
    if width_ratio < 0.50 && line_balance < 0.35 && max_lines >= 20 && narrow_width >= 160.0 {
        // compare avg Y‑gaps …
        if narrow_gap / wide_gap >= 2.5 { return true; }
    }
}

```

### Stage 3: Balanced Dense Columns

For columns containing substantial text (15 or more lines each), the algorithm applies a simple balance ratio test. If `min_lines / max_lines > 0.7`, the layout is classified as newspaper regardless of vertical alignment. This reflects the typical newspaper layout where columns contain comparable amounts of text flowing independently.

```rust
let balance_ratio = min_lines as f32 / max_lines as f32;
if balance_ratio > 0.7 { return true; }

```

### Stage 4: Y-Collision Fallback

When columns are unbalanced or fail the density test, the algorithm performs a **vertical collision analysis** on the smallest column. For each line in this column, it searches other columns for lines whose baselines fall within **±5 pt** (`y_tol = 5.0`).

If more than 50% of lines in the smallest column vertically align with lines in other columns (`collisions / smallest.len() > 0.5`), the layout is deemed tabular. Conversely, a low collision ratio indicates independent prose columns typical of newspaper layouts.

```rust
let y_tol = 5.0;
for line in smallest {
    for (ci, col) in per_column_lines.iter().enumerate() { … }
}
let ratio = collisions as f32 / smallest.len() as f32;
ratio > 0.5

```

## Pipeline Integration and Reading Order

The newspaper detection integrates into the extraction pipeline at multiple points in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) and [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs). After `detect_columns` builds horizontal occupancy histograms and `group_into_lines` organizes text items, the system calls `is_newspaper_layout` to determine reading order.

When the function returns `true`, the extractor processes columns sequentially (complete column 1, then column 2). When `false`, it applies Y-interleaved ordering suitable for tables. The [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) module consumes this decision to emit correctly ordered Markdown output.

## Practical Implementation Example

The following Rust pattern demonstrates how to invoke the detection logic after column extraction:

```rust
use pdf_inspector::extractor::{detect_columns, group_into_lines, is_newspaper_layout};

// 1. Extract raw text items from the PDF (already available as `items`).
let lines = group_into_lines(items.clone());

// 2. Detect column regions on each page.
let columns = detect_columns(&items, page_number, page_has_table);

// 3. Split lines per column (the function `split_column_stragglers` helps here).
let per_column_lines = split_lines_by_columns(&lines, &columns);

// 4. Ask the layout engine whether the page is newspaper style.
if is_newspaper_layout(&per_column_lines, &columns) {
    println!("Read column 1 → column 2 (newspaper order)");
} else {
    println!("Interleave columns by Y (tabular order)");
}

```

Note that `split_lines_by_columns` represents internal logic similar to the extractor's grouping mechanism that feeds `is_newspaper_layout`.

## Summary

- **Newspaper-style layout detection** in `firecrawl/pdf-inspector` resides in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) within the `is_newspaper_layout` function.
- The algorithm requires at least two columns with five lines each before analysis begins.
- **Sidebar detection** handles asymmetrical two-column layouts common in academic papers by comparing width ratios and Y-gaps.
- **Balanced dense columns** (15+ lines with >0.7 line ratio) automatically classify as newspaper layouts.
- The **Y-collision test** uses a ±5 pt tolerance to detect vertically aligned text indicative of tables, with a 0.5 threshold determining the final classification.
- Detection results drive reading order decisions in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs), ensuring proper sequential or interleaved text extraction.

## Frequently Asked Questions

### What is the difference between newspaper-style and tabular layouts in PDFs?

Newspaper-style layouts contain independent prose flows where text reads sequentially down one column before continuing to the next. Tabular layouts organize content in rows where horizontally adjacent cells relate to each other, requiring interleaved reading order. PDF-Inspector distinguishes these by analyzing vertical alignment patterns and line density ratios.

### How does the Y-collision test distinguish tables from newspaper columns?

The Y-collision test checks if lines in different columns share the same vertical baseline within ±5 points. Tables exhibit high vertical alignment between columns (many collisions), while newspaper columns show independent vertical flow (few collisions). When collisions exceed 50% of lines, the layout is classified as tabular.

### What are the minimum requirements for a layout to be considered newspaper-style?

The layout must contain at least two columns, each with a minimum of five lines of text. Without these thresholds, the algorithm returns `false` and defaults to tabular processing. Additional requirements vary by detection stage, such as the 15-line minimum for balanced column tests or specific width ratios for sidebar detection.

### Where is the newspaper detection logic used in the extraction pipeline?

The `is_newspaper_layout` function is called after column detection and line grouping in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs). Its boolean result determines whether [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) processes columns sequentially or applies Y-interleaved ordering. Complementary document-level heuristics also exist in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) for OCR recommendations.