# How LiteParse Performs Multi-Column Layout Detection in PDF Documents

> LiteParse detects multi-column PDF layouts by identifying large horizontal gaps between text lines. Learn how this method ensures accurate document parsing and content extraction.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-06

---

**LiteParse detects multi-column layouts by scanning text lines for oversized horizontal gaps that straddle the page midpoint, counting those gaps within each block, and refusing flowing-text classification when two or more column-gap lines are found.**

LiteParse, the open-source PDF parsing library from run-llama, reconstructs document geometry during a projection phase that determines whether a text block is a single flowing paragraph or a structured multi-column region. Its approach to multi-column layout detection in PDF documents operates entirely on positional data extracted from PDFium, using median character dimensions as a baseline for inter-item gap analysis.

## The Projection Pipeline: From Raw Text to Line Groups

The detection process begins in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) by extracting raw text items from PDFium. Each **ProjectedTextItem** carries exact `x`, `y`, `width`, `height`, and text string values.

Next, the `form_lines` function groups items into horizontal lines. Items with vertical positions within a tolerance derived from the median character height are snapped to the same line. The resulting line vector is then sorted by `x` coordinate to establish left-to-right reading order.

Before gap analysis, `compute_median_textbox_size` calculates the median character width (`median_width`). This value serves as the baseline for what constitutes a normal inter-word gap. Any gap significantly larger than this median becomes a candidate column separator.

## Identifying Column Gaps with `line_has_column_gap`

The core multi-column detection logic lives in the `line_has_column_gap` function inside [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) at approximately line 1244. As implemented in run-llama/liteparse, this function inspects every adjacent pair of text items on a single line.

A gap is flagged as a column separator when two conditions are met:

- The horizontal gap between items exceeds `median_width * 2.0`.
- The pair of items straddles the page midpoint.

When both conditions hold, the function returns `true`, signaling that the line contains a column gap.

```rust
fn line_has_column_gap(
    line: &[ProjectedTextItem],
    median_width: f32,
    page_width: f32,
) -> bool {
    // projection.rs ~L1244
    // Checks every adjacent pair of items on the line.
    // Returns true if the horizontal gap > median_width * 2.0
    // and the pair straddles the page midpoint.
}

```

This lightweight geometric check requires no explicit column markup from the PDF.

## The Flowing-Text Block Heuristic

Individual column-gap lines are not enough to classify an entire block as multi-column. Inside `is_flowing_text_block` in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) around line 1262, LiteParse tallies how many lines within a block satisfy `line_has_column_gap`.

The parser applies a strict threshold:

```rust
if column_gap_lines >= 2 {
    // multi-column block – do NOT treat as flowing text
    return false;
}

```

This heuristic ensures that an occasional wide line does not trigger false multi-column classification. If fewer than two lines exhibit column gaps, and other flow-text constraints—such as anchor count and line width ratio—are satisfied, the block is treated as flowing text. Otherwise, it remains a structured block with implicit columns.

## Anchor Extraction and Layout Filtering

LiteParse further refines layout decisions through an anchor system. The `extract_block_anchors` function at approximately line 1010 in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) generates anchors for left, right, and center positions of each non-rotated bounding box. Rotated items are deliberately ignored to prevent spurious column anchors.

Subsequent filtering steps—`delta_min_filter`, `intercept_filter`, and `try_align_floating`—clean these anchor maps. The anchor pipeline feeds into the flow-text detector, helping distinguish between genuine flowing paragraphs and columnar structures. Supporting type definitions such as `AnchorMap` reside in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs), while tunable constants including `FLOWING_COLUMN_GAP_MULTIPLIER` are defined in [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs). The overall orchestration happens in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs), which invokes the projection step and passes the resulting layout to output formatters.

## Rendering Multi-Column Output

When a block is identified as multi-column, the original `x` coordinates are preserved in the output. No explicit `"column"` field is added to the JSON; instead, the column structure is implicit in the geometry. Items belonging to the left column carry smaller `x` values, while right-column items carry larger values.

## Parsing Multi-Column PDFs in Practice

### Rust

The following Rust example uses the core `LiteParse` library to parse a multi-column PDF and print the structured result:

```rust
use liteparse::LiteParse;
use serde_json::to_string_pretty;
use tokio::fs::File;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    // Initialise the parser with default configuration
    let parser = LiteParse::default();

    // Parse a PDF file (asynchronously)
    let result = parser.parse_file("samples/multi_column.pdf").await?;

    // Serialize the structured result (includes x-coordinates that reveal columns)
    let json = to_string_pretty(&result)?;
    println!("{}", json);
    Ok(())
}

```

The printed JSON contains a `pages → items` array where each item retains its `x` coordinate, allowing downstream consumers to infer column membership from the preserved geometry.

### TypeScript

The Node.js binding exposes the same projection logic. Here is the equivalent TypeScript usage:

```typescript
import { LiteParse } from "liteparse";

(async () => {
  const parser = new LiteParse(); // default config
  const result = await parser.parseFile("samples/multi_column.pdf");
  console.log(JSON.stringify(result, null, 2));
})();

```

The resulting JavaScript object mirrors the Rust structure, maintaining the column-aware `x` positions extracted during the projection phase.

## Summary

- LiteParse performs multi-column layout detection in PDF documents during a geometric projection phase that analyzes `ProjectedTextItem` coordinates from PDFium.
- The `line_has_column_gap` function in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) detects candidate column breaks by looking for gaps wider than twice the median character width that cross the page midpoint.
- The `is_flowing_text_block` function rejects flowing-text classification when at least two lines in a block contain such gaps.
- Anchor extraction at `extract_block_anchors` (~line 1010) and subsequent filters ignore rotated items to avoid false anchors.
- Detected multi-column blocks retain their original `x` coordinates in the output JSON, making column structure implicit in the geometry.

## Frequently Asked Questions

### How does LiteParse distinguish between a wide word gap and an actual column boundary?

LiteParse uses the median character width as a baseline. A gap must exceed `median_width * 2.0` and straddle the page midpoint to qualify as a column boundary. This dual requirement filters out ordinary wide spaces that do not represent true layout divisions.

### Where is the multi-column detection threshold configured?

The gap multiplier and related tuning constants are defined in [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs). For example, `FLOWING_COLUMN_GAP_MULTIPLIER` controls the factor applied to the median width when `line_has_column_gap` evaluates inter-item distances.

### Why does LiteParse require two or more column-gap lines before classifying a block as multi-column?

The `is_flowing_text_block` heuristic demands at least two column-gap lines to avoid misclassifying blocks that contain a single unusually wide line. This conservative threshold improves robustness against sporadic layout anomalies.

### Are rotated text blocks handled differently during column detection?

Yes. The `extract_block_anchors` pipeline ignores rotated bounding boxes when generating left, right, and center anchors. This prevents rotated text from introducing spurious anchors that could distort the flow-text vs. multi-column decision.