How pdf‑inspector Distinguishes Multi‑Column Newspaper Layouts from Tabular Layouts
pdf‑inspector uses a two-stage pipeline: first detecting column boundaries via horizontal projection histograms and XY-cut heuristics, then classifying the result as newspaper or tabular based on prose density, line-fill ratios, and line-count balance across columns.
The firecrawl/pdf‑inspector repository implements sophisticated layout analysis to differentiate newspaper-style multi-column prose from true data tables. Understanding this distinction matters because reading order differs dramatically—newspaper columns are read sequentially top-to-bottom, while tabular cells must be interleaved by row position. The implementation in src/extractor/layout.rs combines geometric analysis with content heuristics to make this determination.
Column Detection via Horizontal Projection Histograms
The first stage identifies candidate column boundaries using the detect_columns function in [src/extractor/layout.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs#L21-L124).
Filtering Wide Elements
The engine begins by excluding text items that exceed 60% of page width. This prevents full-width titles, headers, or paragraphs from obscuring the gutter valleys between columns.
Absolute Valley Detection
Empty regions—histogram bins with low text counts—become candidate boundaries when they satisfy three conditions:
- Width threshold: The gap must exceed
MIN_GUTTER_WIDTH - Edge exclusion: The valley cannot be too close to page margins
- Significant depth: The bin count must drop substantially from adjacent peaks
Relative Valley Fallback
When justified text eliminates completely empty gutters, the find_relative_valleys helper locates local minima that dip significantly below surrounding peaks. This catches layouts where spacing is compressed but columns remain visually distinct.
XY-Cut Asymmetric Handling
For irregular layouts with sidebars or uneven columns, try_xy_cut_split searches for the largest horizontal gap between item edges. This handles cases where histogram-based methods fail due to non-uniform column widths.
Newspaper vs. Tabular Classification
After column detection, is_newspaper_layout applies content-based heuristics to distinguish prose from data tables.
Line-Level Grouping
Items within each column are aggregated into rough lines using Y-proximity thresholds via group_into_lines_with_thresholds. This establishes the fundamental unit for subsequent analysis.
Prose Density Evaluation
The columns_have_prose function implements the core classification logic:
| Metric | Threshold | Purpose |
|---|---|---|
LINE_FILL_THRESHOLD |
0.45 | Minimum width ratio for a "full line" |
MIN_PROSE_RATIO |
0.40 | Minimum proportion of full lines to qualify as prose |
MAX_AVG_ITEMS_PER_LINE |
3.5 | Maximum average items per line for prose |
A column qualifies as prose-rich when more than 40% of its lines span at least 45% of column width and the average items per line stays below 3.5. Tables typically fail this test due to fragmented cells and high item counts per line.
Balance and Width Heuristics
Balanced line counts: Two or more columns with line counts within 30% of each other indicate parallel newspaper columns.
Column-width ratio: Very narrow sidebars still classify as newspaper when the primary column exceeds 40% of page width and the narrow column passes prose density checks.
Tabular exclusion: High avg_items_per_line values and low full-line ratios cause columns_have_prose to return false, routing the layout through table-specific processing.
CLI and Library Usage
Command-Line Extraction
pdf2md --json my_newspaper.pdf > out.json
The JSON output includes a reading_order field set to "newspaper" when is_newspaper_layout returns true, or "tabular" for table-dominated pages.
Programmatic Classification
use pdf_inspector::extractor::{detect_columns, is_newspaper_layout};
use pdf_inspector::types::TextItem;
/// Classify a page's layout based on its text items.
fn classify_page(items: &[TextItem], page: u32, has_table: bool) {
// Detect column boundaries
let columns = detect_columns(items, page, has_table);
// Group items into lines per column
let per_column_lines = pdf_inspector::extractor
::group_into_lines_with_thresholds(&columns, items);
// Apply newspaper vs. tabular classification
let newspaper = is_newspaper_layout(&per_column_lines, &columns);
println!(
"Page {} is {}",
page,
if newspaper { "newspaper" } else { "tabular" }
);
}
This API enables custom pipelines to inspect intermediate classification decisions.
Integration with Table Detection Pipeline
The classification result propagates to src/tables/mod.rs, which selects appropriate reading order:
- Newspaper layouts: Columns emitted sequentially (left column complete, then right column)
- Tabular layouts: Cells interleaved by Y-position to preserve row relationships
The final Markdown serialization in src/markdown/convert.rs embeds this decision in output metadata.
Summary
- Horizontal projection histograms in
detect_columnslocate column boundaries through empty valleys and relative minima, with XY-cut fallback for asymmetric layouts - Prose density metrics (
LINE_FILL_THRESHOLD,MIN_PROSE_RATIO,MAX_AVG_ITEMS_PER_LINE) separate continuous text from fragmented table cells - Balance heuristics compare line counts and column widths to confirm newspaper-style parallelism
- Classification output drives reading-order selection downstream in the table detection pipeline
Frequently Asked Questions
What threshold values control newspaper detection in pdf‑inspector?
pdf‑inspector uses three primary constants in src/extractor/layout.rs: LINE_FILL_THRESHOLD = 0.45 defines a "full" line, MIN_PROSE_RATIO = 0.40 requires 40% full lines for prose classification, and MAX_AVG_ITEMS_PER_LINE = 3.5 caps fragment density. These values are hardcoded based on empirical document analysis.
How does pdf‑inspector handle justified text without clear gutters?
The find_relative_valleys function implements a relative-valley fallback that identifies local histogram minima dipping significantly below adjacent peaks. This catches compressed newspaper layouts where inter-column spacing is minimized for aesthetic reasons.
Can the newspaper detection be disabled for specific documents?
The detection runs automatically per-page, but the has_table parameter in detect_columns influences heuristic sensitivity. For forced classification, direct API access through is_newspaper_layout allows manual override after custom preprocessing.
Where does the final reading order decision get applied?
src/tables/mod.rs consumes the boolean result from is_newspaper_layout. When true, it preserves sequential column reading; when false, it interleaves content by Y-coordinate. The src/markdown/convert.rs module serializes this choice into the output's reading_order field.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →