# How pdf-inspector Detects Columns and Layouts in Multi-Column PDFs

> Discover how pdf-inspector uses histogram-based geometric analysis to detect columns and layouts in multi-column PDFs. Learn its newspaper and tabular classification methods.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-08

---

**pdf-inspector uses a histogram-based geometric analysis in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) to identify column boundaries by projecting text baselines onto the X-axis and analyzing gaps, then classifies layouts as newspaper or tabular.**

Multi-column PDFs present unique challenges for text extraction because standard reading order can flow vertically within columns or horizontally across rows. The **pdf-inspector** library, developed by Firecrawl, solves this by parsing each page into geometric `TextItem` objects and applying specialized layout detection algorithms that preserve document structure.

## The Core Algorithm: Horizontal Projection Histograms

The column detection process begins by treating text items as geometric entities with baseline coordinates rather than mere character streams.

### Building the X-Axis Histogram

In [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs), the `detect_columns` function projects all text baselines onto the horizontal axis. The algorithm creates a histogram with default buckets of approximately **2 pt**, counting how many `TextItem` objects occupy each horizontal position. This projection reveals the spatial distribution of text across the page width.

### Identifying Column Boundaries

The algorithm identifies column separators by locating valleys in the histogram. A gap qualifies as a column boundary when it exceeds the configurable threshold `COLUMN_GAP_MIN` (approximately **40 pt**) and is surrounded by sufficiently populated peaks. Each detected column is stored as a `ColumnRegion` struct containing `x_min` and `x_max` coordinates, sorted left-to-right for downstream processing.

## Handling Spanning Elements and Edge Cases

Real-world PDFs contain elements that cross column boundaries, such as titles and page numbers, which require special handling to prevent false column detection.

### Pre-Masking Spanning Lines

Before finalizing column assignments, `detect_columns` removes "spanning lines" that cross detected column gaps. The algorithm masks these lines when their height falls below a column-aware threshold and their width crosses a potential gap. This prevents wide headings from corrupting the histogram and creating false column boundaries.

### ColumnRegion Struct and Assignment

Each `ColumnRegion` tracks its horizontal boundaries. The `assign_to_best_overlap` function assigns every `TextItem` to the column with which it shares the greatest horizontal overlap. Items spanning multiple columns are identified as either split candidates or preserved as spanning elements based on surrounding context.

## Layout Classification: Newspaper vs. Tabular

After establishing column boundaries, pdf-inspector classifies the page layout type to determine proper reading order reconstruction.

The `is_newspaper_layout` function examines `per_column_lines` to distinguish between layout types:

- **Newspaper layouts**: Characterized by asymmetric column density (greater than 60% of items in a single column) or configurations with side annotations. Text flows vertically within each column before moving to the next.
- **Tabular layouts**: Multiple columns of comparable density where text may flow horizontally across rows. These boundaries feed into `tables::mod::try_build_table_from_columns` for borderless table reconstruction.

## Integration in the Extraction Pipeline

The column detection integrates into the main processing workflow through `process_pdf_with_options` in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs). For every page, the extractor:

1. Builds the raw text stream into `layout_items`
2. Invokes `detect_columns` to generate column metadata
3. Stores results in `PageInfo.pages_with_columns` (defined in [`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs))
4. Passes column information to downstream modules for reading-order reconstruction and table detection

## Practical Implementation Examples

You can interact with the column detection logic both as a library and via command line.

Using the Rust library:

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::extractor::detect_columns;

fn analyze_pdf_columns() -> Result<(), Box<dyn std::error::Error>> {
    let pdf_bytes = std::fs::read("research-paper.pdf")?;
    let mut opts = pdf_inspector::ProcessOptions::default();
    let result = process_pdf_with_options(&pdf_bytes, &mut opts)?;
    
    for (page_no, page) in result.pages.iter().enumerate() {
        // `page.layout_items` contains the TextItem baselines
        let cols = detect_columns(&page.layout_items, page_no as u32, false);
        println!("Page {}: {} columns detected", page_no + 1, cols.len());
    }
    Ok(())
}

```

Using the CLI:

```bash

# Extract Markdown with column-aware layout handling

pdf2md research-paper.pdf

# View detected column boundaries in JSON format

pdf2md research-paper.pdf --json | jq '.pages[].columns'

```

## Summary

- **Geometric histogram analysis**: pdf-inspector projects text baselines onto the X-axis in [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) to identify column gaps using configurable thresholds.
- **Robust edge case handling**: Spanning lines and wide headings are pre-masked to prevent false column boundaries before `assign_to_best_overlap` distributes items into `ColumnRegion` structs.
- **Layout classification**: The system distinguishes newspaper (vertical flow) from tabular (horizontal flow) layouts using density analysis of `per_column_lines`.
- **Pipeline integration**: Column detection runs automatically within `process_pdf_with_options`, storing results in `PageInfo.pages_with_columns` for table reconstruction and Markdown conversion.

## Frequently Asked Questions

### What file contains the main column detection logic?

The primary implementation resides in **[`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs)**, which exports the `detect_columns` function. This file handles histogram generation, gap analysis, and the `ColumnRegion` struct definitions.

### How does pdf-inspector handle headings that span multiple columns?

Before building the final column model, the algorithm pre-masks "spanning lines" in `detect_columns`. Wide headings that cross column gaps are temporarily removed from the histogram analysis if their height falls below a column-aware threshold, preventing them from creating false column boundaries while preserving them for later text assignment.

### What is the difference between newspaper and tabular layout detection?

**Newspaper layouts** exhibit asymmetric column density (typically >60% of content in one column) and follow vertical reading order within each column. **Tabular layouts** show comparable density across multiple columns and potentially flow horizontally, triggering the table detection pipeline via `tables::mod::try_build_table_from_columns`.

### Can I configure the minimum gap threshold for column detection?

Yes, the `COLUMN_GAP_MIN` constant (default approximately 40 pt) defines the minimum horizontal gap required to register as a column separator. While the public API in `process_pdf_with_options` uses sensible defaults, the underlying `detect_columns` function accepts parameters that influence gap detection sensitivity.