How LiteParse Performs Multi-Column Layout Detection in PDF Documents
LiteParse detects multi-column layouts by scanning text lines for oversized horizontal gaps that straddle the page midpoint, counting those gaps within each block, and refusing flowing-text classification when two or more column-gap lines are found.
LiteParse, the open-source PDF parsing library from run-llama, reconstructs document geometry during a projection phase that determines whether a text block is a single flowing paragraph or a structured multi-column region. Its approach to multi-column layout detection in PDF documents operates entirely on positional data extracted from PDFium, using median character dimensions as a baseline for inter-item gap analysis.
The Projection Pipeline: From Raw Text to Line Groups
The detection process begins in crates/liteparse/src/projection.rs by extracting raw text items from PDFium. Each ProjectedTextItem carries exact x, y, width, height, and text string values.
Next, the form_lines function groups items into horizontal lines. Items with vertical positions within a tolerance derived from the median character height are snapped to the same line. The resulting line vector is then sorted by x coordinate to establish left-to-right reading order.
Before gap analysis, compute_median_textbox_size calculates the median character width (median_width). This value serves as the baseline for what constitutes a normal inter-word gap. Any gap significantly larger than this median becomes a candidate column separator.
Identifying Column Gaps with line_has_column_gap
The core multi-column detection logic lives in the line_has_column_gap function inside projection.rs at approximately line 1244. As implemented in run-llama/liteparse, this function inspects every adjacent pair of text items on a single line.
A gap is flagged as a column separator when two conditions are met:
- The horizontal gap between items exceeds
median_width * 2.0. - The pair of items straddles the page midpoint.
When both conditions hold, the function returns true, signaling that the line contains a column gap.
fn line_has_column_gap(
line: &[ProjectedTextItem],
median_width: f32,
page_width: f32,
) -> bool {
// projection.rs ~L1244
// Checks every adjacent pair of items on the line.
// Returns true if the horizontal gap > median_width * 2.0
// and the pair straddles the page midpoint.
}
This lightweight geometric check requires no explicit column markup from the PDF.
The Flowing-Text Block Heuristic
Individual column-gap lines are not enough to classify an entire block as multi-column. Inside is_flowing_text_block in projection.rs around line 1262, LiteParse tallies how many lines within a block satisfy line_has_column_gap.
The parser applies a strict threshold:
if column_gap_lines >= 2 {
// multi-column block – do NOT treat as flowing text
return false;
}
This heuristic ensures that an occasional wide line does not trigger false multi-column classification. If fewer than two lines exhibit column gaps, and other flow-text constraints—such as anchor count and line width ratio—are satisfied, the block is treated as flowing text. Otherwise, it remains a structured block with implicit columns.
Anchor Extraction and Layout Filtering
LiteParse further refines layout decisions through an anchor system. The extract_block_anchors function at approximately line 1010 in projection.rs generates anchors for left, right, and center positions of each non-rotated bounding box. Rotated items are deliberately ignored to prevent spurious column anchors.
Subsequent filtering steps—delta_min_filter, intercept_filter, and try_align_floating—clean these anchor maps. The anchor pipeline feeds into the flow-text detector, helping distinguish between genuine flowing paragraphs and columnar structures. Supporting type definitions such as AnchorMap reside in crates/liteparse/src/types.rs, while tunable constants including FLOWING_COLUMN_GAP_MULTIPLIER are defined in crates/liteparse/src/config.rs. The overall orchestration happens in crates/liteparse/src/parser.rs, which invokes the projection step and passes the resulting layout to output formatters.
Rendering Multi-Column Output
When a block is identified as multi-column, the original x coordinates are preserved in the output. No explicit "column" field is added to the JSON; instead, the column structure is implicit in the geometry. Items belonging to the left column carry smaller x values, while right-column items carry larger values.
Parsing Multi-Column PDFs in Practice
Rust
The following Rust example uses the core LiteParse library to parse a multi-column PDF and print the structured result:
use liteparse::LiteParse;
use serde_json::to_string_pretty;
use tokio::fs::File;
#[tokio::main]
async fn main() -> anyhow::Result<()> {
// Initialise the parser with default configuration
let parser = LiteParse::default();
// Parse a PDF file (asynchronously)
let result = parser.parse_file("samples/multi_column.pdf").await?;
// Serialize the structured result (includes x-coordinates that reveal columns)
let json = to_string_pretty(&result)?;
println!("{}", json);
Ok(())
}
The printed JSON contains a pages → items array where each item retains its x coordinate, allowing downstream consumers to infer column membership from the preserved geometry.
TypeScript
The Node.js binding exposes the same projection logic. Here is the equivalent TypeScript usage:
import { LiteParse } from "liteparse";
(async () => {
const parser = new LiteParse(); // default config
const result = await parser.parseFile("samples/multi_column.pdf");
console.log(JSON.stringify(result, null, 2));
})();
The resulting JavaScript object mirrors the Rust structure, maintaining the column-aware x positions extracted during the projection phase.
Summary
- LiteParse performs multi-column layout detection in PDF documents during a geometric projection phase that analyzes
ProjectedTextItemcoordinates from PDFium. - The
line_has_column_gapfunction incrates/liteparse/src/projection.rsdetects candidate column breaks by looking for gaps wider than twice the median character width that cross the page midpoint. - The
is_flowing_text_blockfunction rejects flowing-text classification when at least two lines in a block contain such gaps. - Anchor extraction at
extract_block_anchors(~line 1010) and subsequent filters ignore rotated items to avoid false anchors. - Detected multi-column blocks retain their original
xcoordinates in the output JSON, making column structure implicit in the geometry.
Frequently Asked Questions
How does LiteParse distinguish between a wide word gap and an actual column boundary?
LiteParse uses the median character width as a baseline. A gap must exceed median_width * 2.0 and straddle the page midpoint to qualify as a column boundary. This dual requirement filters out ordinary wide spaces that do not represent true layout divisions.
Where is the multi-column detection threshold configured?
The gap multiplier and related tuning constants are defined in crates/liteparse/src/config.rs. For example, FLOWING_COLUMN_GAP_MULTIPLIER controls the factor applied to the median width when line_has_column_gap evaluates inter-item distances.
Why does LiteParse require two or more column-gap lines before classifying a block as multi-column?
The is_flowing_text_block heuristic demands at least two column-gap lines to avoid misclassifying blocks that contain a single unusually wide line. This conservative threshold improves robustness against sporadic layout anomalies.
Are rotated text blocks handled differently during column detection?
Yes. The extract_block_anchors pipeline ignores rotated bounding boxes when generating left, right, and center anchors. This prevents rotated text from introducing spurious anchors that could distort the flow-text vs. multi-column decision.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →