How LiteParse Performs Multi-Column Layout Detection in PDF Documents

LiteParse detects multi-column layouts by scanning text lines for oversized horizontal gaps that straddle the page midpoint, counting those gaps within each block, and refusing flowing-text classification when two or more column-gap lines are found.

LiteParse, the open-source PDF parsing library from run-llama, reconstructs document geometry during a projection phase that determines whether a text block is a single flowing paragraph or a structured multi-column region. Its approach to multi-column layout detection in PDF documents operates entirely on positional data extracted from PDFium, using median character dimensions as a baseline for inter-item gap analysis.

The Projection Pipeline: From Raw Text to Line Groups

The detection process begins in crates/liteparse/src/projection.rs by extracting raw text items from PDFium. Each ProjectedTextItem carries exact x, y, width, height, and text string values.

Next, the form_lines function groups items into horizontal lines. Items with vertical positions within a tolerance derived from the median character height are snapped to the same line. The resulting line vector is then sorted by x coordinate to establish left-to-right reading order.

Before gap analysis, compute_median_textbox_size calculates the median character width (median_width). This value serves as the baseline for what constitutes a normal inter-word gap. Any gap significantly larger than this median becomes a candidate column separator.

Identifying Column Gaps with line_has_column_gap

The core multi-column detection logic lives in the line_has_column_gap function inside projection.rs at approximately line 1244. As implemented in run-llama/liteparse, this function inspects every adjacent pair of text items on a single line.

A gap is flagged as a column separator when two conditions are met:

  • The horizontal gap between items exceeds median_width * 2.0.
  • The pair of items straddles the page midpoint.

When both conditions hold, the function returns true, signaling that the line contains a column gap.

fn line_has_column_gap(
    line: &[ProjectedTextItem],
    median_width: f32,
    page_width: f32,
) -> bool {
    // projection.rs ~L1244
    // Checks every adjacent pair of items on the line.
    // Returns true if the horizontal gap > median_width * 2.0
    // and the pair straddles the page midpoint.
}

This lightweight geometric check requires no explicit column markup from the PDF.

The Flowing-Text Block Heuristic

Individual column-gap lines are not enough to classify an entire block as multi-column. Inside is_flowing_text_block in projection.rs around line 1262, LiteParse tallies how many lines within a block satisfy line_has_column_gap.

The parser applies a strict threshold:

if column_gap_lines >= 2 {
    // multi-column block – do NOT treat as flowing text
    return false;
}

This heuristic ensures that an occasional wide line does not trigger false multi-column classification. If fewer than two lines exhibit column gaps, and other flow-text constraints—such as anchor count and line width ratio—are satisfied, the block is treated as flowing text. Otherwise, it remains a structured block with implicit columns.

Anchor Extraction and Layout Filtering

LiteParse further refines layout decisions through an anchor system. The extract_block_anchors function at approximately line 1010 in projection.rs generates anchors for left, right, and center positions of each non-rotated bounding box. Rotated items are deliberately ignored to prevent spurious column anchors.

Subsequent filtering steps—delta_min_filter, intercept_filter, and try_align_floating—clean these anchor maps. The anchor pipeline feeds into the flow-text detector, helping distinguish between genuine flowing paragraphs and columnar structures. Supporting type definitions such as AnchorMap reside in crates/liteparse/src/types.rs, while tunable constants including FLOWING_COLUMN_GAP_MULTIPLIER are defined in crates/liteparse/src/config.rs. The overall orchestration happens in crates/liteparse/src/parser.rs, which invokes the projection step and passes the resulting layout to output formatters.

Rendering Multi-Column Output

When a block is identified as multi-column, the original x coordinates are preserved in the output. No explicit "column" field is added to the JSON; instead, the column structure is implicit in the geometry. Items belonging to the left column carry smaller x values, while right-column items carry larger values.

Parsing Multi-Column PDFs in Practice

Rust

The following Rust example uses the core LiteParse library to parse a multi-column PDF and print the structured result:

use liteparse::LiteParse;
use serde_json::to_string_pretty;
use tokio::fs::File;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    // Initialise the parser with default configuration
    let parser = LiteParse::default();

    // Parse a PDF file (asynchronously)
    let result = parser.parse_file("samples/multi_column.pdf").await?;

    // Serialize the structured result (includes x-coordinates that reveal columns)
    let json = to_string_pretty(&result)?;
    println!("{}", json);
    Ok(())
}

The printed JSON contains a pages → items array where each item retains its x coordinate, allowing downstream consumers to infer column membership from the preserved geometry.

TypeScript

The Node.js binding exposes the same projection logic. Here is the equivalent TypeScript usage:

import { LiteParse } from "liteparse";

(async () => {
  const parser = new LiteParse(); // default config
  const result = await parser.parseFile("samples/multi_column.pdf");
  console.log(JSON.stringify(result, null, 2));
})();

The resulting JavaScript object mirrors the Rust structure, maintaining the column-aware x positions extracted during the projection phase.

Summary

  • LiteParse performs multi-column layout detection in PDF documents during a geometric projection phase that analyzes ProjectedTextItem coordinates from PDFium.
  • The line_has_column_gap function in crates/liteparse/src/projection.rs detects candidate column breaks by looking for gaps wider than twice the median character width that cross the page midpoint.
  • The is_flowing_text_block function rejects flowing-text classification when at least two lines in a block contain such gaps.
  • Anchor extraction at extract_block_anchors (~line 1010) and subsequent filters ignore rotated items to avoid false anchors.
  • Detected multi-column blocks retain their original x coordinates in the output JSON, making column structure implicit in the geometry.

Frequently Asked Questions

How does LiteParse distinguish between a wide word gap and an actual column boundary?

LiteParse uses the median character width as a baseline. A gap must exceed median_width * 2.0 and straddle the page midpoint to qualify as a column boundary. This dual requirement filters out ordinary wide spaces that do not represent true layout divisions.

Where is the multi-column detection threshold configured?

The gap multiplier and related tuning constants are defined in crates/liteparse/src/config.rs. For example, FLOWING_COLUMN_GAP_MULTIPLIER controls the factor applied to the median width when line_has_column_gap evaluates inter-item distances.

Why does LiteParse require two or more column-gap lines before classifying a block as multi-column?

The is_flowing_text_block heuristic demands at least two column-gap lines to avoid misclassifying blocks that contain a single unusually wide line. This conservative threshold improves robustness against sporadic layout anomalies.

Are rotated text blocks handled differently during column detection?

Yes. The extract_block_anchors pipeline ignores rotated bounding boxes when generating left, right, and center anchors. This prevents rotated text from introducing spurious anchors that could distort the flow-text vs. multi-column decision.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →