# How pdf‑inspector Distinguishes Multi‑Column Newspaper Layouts from Tabular Layouts

> Discover how pdf-inspector differentiates multi-column newspaper layouts from tabular data using its two-stage detection and classification pipeline. Learn about prose density, fill ratios, and column balance.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-06

---

**pdf‑inspector uses a two-stage pipeline: first detecting column boundaries via horizontal projection histograms and XY-cut heuristics, then classifying the result as newspaper or tabular based on prose density, line-fill ratios, and line-count balance across columns.**

The firecrawl/pdf‑inspector repository implements sophisticated layout analysis to differentiate newspaper-style multi-column prose from true data tables. Understanding this distinction matters because reading order differs dramatically—newspaper columns are read sequentially top-to-bottom, while tabular cells must be interleaved by row position. The implementation in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) combines geometric analysis with content heuristics to make this determination.

## Column Detection via Horizontal Projection Histograms

The first stage identifies candidate column boundaries using the `detect_columns` function in [[`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs#L21-L124).

### Filtering Wide Elements

The engine begins by excluding text items that exceed **60% of page width**. This prevents full-width titles, headers, or paragraphs from obscuring the gutter valleys between columns.

### Absolute Valley Detection

Empty regions—histogram bins with low text counts—become candidate boundaries when they satisfy three conditions:

- **Width threshold**: The gap must exceed `MIN_GUTTER_WIDTH`
- **Edge exclusion**: The valley cannot be too close to page margins
- **Significant depth**: The bin count must drop substantially from adjacent peaks

### Relative Valley Fallback

When justified text eliminates completely empty gutters, the `find_relative_valleys` helper locates local minima that dip significantly below surrounding peaks. This catches layouts where spacing is compressed but columns remain visually distinct.

### XY-Cut Asymmetric Handling

For irregular layouts with sidebars or uneven columns, `try_xy_cut_split` searches for the largest horizontal gap between item edges. This handles cases where histogram-based methods fail due to non-uniform column widths.

## Newspaper vs. Tabular Classification

After column detection, [`is_newspaper_layout`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs#L1673-L1726) applies content-based heuristics to distinguish prose from data tables.

### Line-Level Grouping

Items within each column are aggregated into rough lines using Y-proximity thresholds via `group_into_lines_with_thresholds`. This establishes the fundamental unit for subsequent analysis.

### Prose Density Evaluation

The `columns_have_prose` function implements the core classification logic:

| Metric | Threshold | Purpose |
|--------|-----------|---------|
| `LINE_FILL_THRESHOLD` | 0.45 | Minimum width ratio for a "full line" |
| `MIN_PROSE_RATIO` | 0.40 | Minimum proportion of full lines to qualify as prose |
| `MAX_AVG_ITEMS_PER_LINE` | 3.5 | Maximum average items per line for prose |

A column qualifies as prose-rich when **more than 40% of its lines span at least 45% of column width** **and** the average items per line stays below 3.5. Tables typically fail this test due to fragmented cells and high item counts per line.

### Balance and Width Heuristics

**Balanced line counts**: Two or more columns with line counts within 30% of each other indicate parallel newspaper columns.

**Column-width ratio**: Very narrow sidebars still classify as newspaper when the primary column exceeds 40% of page width and the narrow column passes prose density checks.

**Tabular exclusion**: High `avg_items_per_line` values and low full-line ratios cause `columns_have_prose` to return `false`, routing the layout through table-specific processing.

## CLI and Library Usage

### Command-Line Extraction

```bash
pdf2md --json my_newspaper.pdf > out.json

```

The JSON output includes a `reading_order` field set to `"newspaper"` when `is_newspaper_layout` returns `true`, or `"tabular"` for table-dominated pages.

### Programmatic Classification

```rust
use pdf_inspector::extractor::{detect_columns, is_newspaper_layout};
use pdf_inspector::types::TextItem;

/// Classify a page's layout based on its text items.
fn classify_page(items: &[TextItem], page: u32, has_table: bool) {
    // Detect column boundaries
    let columns = detect_columns(items, page, has_table);
    
    // Group items into lines per column
    let per_column_lines = pdf_inspector::extractor
        ::group_into_lines_with_thresholds(&columns, items);
    
    // Apply newspaper vs. tabular classification
    let newspaper = is_newspaper_layout(&per_column_lines, &columns);
    
    println!(
        "Page {} is {}",
        page,
        if newspaper { "newspaper" } else { "tabular" }
    );
}

```

This API enables custom pipelines to inspect intermediate classification decisions.

## Integration with Table Detection Pipeline

The classification result propagates to [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs), which selects appropriate reading order:

- **Newspaper layouts**: Columns emitted sequentially (left column complete, then right column)
- **Tabular layouts**: Cells interleaved by Y-position to preserve row relationships

The final Markdown serialization in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) embeds this decision in output metadata.

## Summary

- **Horizontal projection histograms** in `detect_columns` locate column boundaries through empty valleys and relative minima, with XY-cut fallback for asymmetric layouts
- **Prose density metrics** (`LINE_FILL_THRESHOLD`, `MIN_PROSE_RATIO`, `MAX_AVG_ITEMS_PER_LINE`) separate continuous text from fragmented table cells
- **Balance heuristics** compare line counts and column widths to confirm newspaper-style parallelism
- **Classification output** drives reading-order selection downstream in the table detection pipeline

## Frequently Asked Questions

### What threshold values control newspaper detection in pdf‑inspector?

pdf‑inspector uses three primary constants in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs): `LINE_FILL_THRESHOLD = 0.45` defines a "full" line, `MIN_PROSE_RATIO = 0.40` requires 40% full lines for prose classification, and `MAX_AVG_ITEMS_PER_LINE = 3.5` caps fragment density. These values are hardcoded based on empirical document analysis.

### How does pdf‑inspector handle justified text without clear gutters?

The `find_relative_valleys` function implements a relative-valley fallback that identifies local histogram minima dipping significantly below adjacent peaks. This catches compressed newspaper layouts where inter-column spacing is minimized for aesthetic reasons.

### Can the newspaper detection be disabled for specific documents?

The detection runs automatically per-page, but the `has_table` parameter in `detect_columns` influences heuristic sensitivity. For forced classification, direct API access through `is_newspaper_layout` allows manual override after custom preprocessing.

### Where does the final reading order decision get applied?

[`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs) consumes the boolean result from `is_newspaper_layout`. When `true`, it preserves sequential column reading; when `false`, it interleaves content by Y-coordinate. The [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) module serializes this choice into the output's `reading_order` field.