How LiteParse's Grid Projection Algorithm Reconstructs Multi-Column Layouts

LiteParse's grid projection algorithm converts PDFium's flat text stream into structured columns by detecting median character dimensions, identifying column gaps, building anchor maps from text positions, and snapping items to a shared grid while filtering outliers and attaching floating elements.

LiteParse transforms the unstructured stream of text items provided by PDFium into a semantic, column-aware representation through a sophisticated grid-projection pipeline. This algorithm, implemented primarily in crates/liteparse/src/projection.rs, reconstructs the original visual layout by detecting column boundaries, establishing spatial anchors, and distinguishing between flowing paragraphs and structured multi-column text.

The Pipeline Stages of LiteParse's Grid Projection Algorithm

The reconstruction process operates through a series of discrete stages that progressively refine the text geometry from raw coordinates to structured output.

Computing Scale Factors with Median Box Sizes

The pipeline begins by establishing a baseline scale using compute_median_textbox_size (lines 31-84 in projection.rs). This function analyzes the bounding boxes of all text items to calculate the median character width and line height, which serve as the fundamental units for all subsequent geometric calculations and gap detection thresholds.

Forming Visual Lines from PDF Items

Next, form_lines (lines 95-140) sorts items by their y-coordinates—snapped to a grid—and merges them into visual lines. This step organizes the flat input stream into horizontal groups that represent actual text lines as they appear to readers, creating the foundation for block-level analysis.

Detecting Column Gaps

The algorithm identifies potential column boundaries through line_has_column_gap (lines 1396-1410). This function examines the horizontal gaps between consecutive items on a line and flags the line as a column separator if a gap exceeds 2× the median character width and the items on either side straddle the page's horizontal midpoint.

Distinguishing Flowing Text from Structured Blocks

Using is_flowing_text_block (lines 1447-1460), the system determines whether a block of lines represents flowing single-column text or a structured multi-column layout. If a block contains two or more column-gap lines, it is treated as structured; otherwise, the algorithm merges the lines into a single flowing paragraph.

Building and Cleaning Anchor Maps

For structured blocks, extract_block_anchors (lines 1569-1594) builds three hash maps—anchor_left, anchor_right, and anchor_center—recording the x-positions of every non-rotated bounding box at quarter-point precision. These anchors are then filtered: delta_min_filter removes isolated anchors with no vertical neighbors within a threshold, while intercept_filter (lines 1600-1645) prunes anchors that are visually crossed by other text items.

Aligning Floating Items and Propagating Anchors

The algorithm handles stray text through try_align_floating (lines 1660-1690), which attempts to attach unattached items to nearby anchors on adjacent lines if they fall within a configurable margin. Simultaneously, update_forward_anchor_right_bound (lines 1628-1638) performs a forward-scanning pass to propagate the rightmost snap of each column, ensuring consistent column widths across the document.

Rendering the Final Column Layout

Finally, the rendering stage uses detect_and_render_flowing_lines and render_line_as_flowing_text (lines 1464-1510) to output the text. Items snap to the nearest anchor in the cleaned maps, with flowing text blocks rendered as single paragraphs and structured blocks maintaining their column separation while preserving original inter-item spacing for column separators.

Key Implementation Details in projection.rs

The core logic resides in crates/liteparse/src/projection.rs, which defines the entire transformation from raw PDF data to structured output. The algorithm relies on BoxMeta and ProjectedTextItem structures defined in crates/liteparse/src/types.rs to maintain geometric metadata throughout the pipeline. Configuration options in crates/liteparse/src/config.rs control whether OCR and column detection are enabled, though the grid projection runs by default for all PDF parsing operations.

Practical Usage Examples

The grid-projection step is transparent to the end user; the output lines are already ordered according to the reconstructed column layout.

Rust – parsing a multi-column PDF:

use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let cfg = LiteParseConfig::default();
    let parser = LiteParse::new(cfg)?;
    
    // input.pdf contains multi-column content
    let result = parser.parse_file("input.pdf")?;
    
    // result.lines respects the column layout detected by the projection algorithm
    for (i, line) in result.lines.iter().enumerate() {
        println!("Line {}: {}", i + 1, line.text);
    }
    Ok(())
}

Python – using the Python bindings:

from liteparse import LiteParse

parser = LiteParse()
result = parser.parse_file("input.pdf")

# result.lines contains dictionaries with text and bbox fields

for i, line in enumerate(result.lines, start=1):
    print(f"Line {i}: {line['text']}")

Summary

  • LiteParse's grid projection algorithm uses median box sizes to establish geometric scale factors for all layout calculations.
  • Column detection relies on identifying gaps greater than twice the median character width that cross the page center.
  • Anchor maps (left, right, center) create a reference grid for snapping text items to consistent x-positions.
  • Floating items are recovered through proximity-based alignment to prevent orphaned text from breaking column structure.
  • Flowing text blocks are automatically detected and rendered as single paragraphs, while structured blocks maintain column boundaries.

Frequently Asked Questions

What determines whether a text block is treated as multi-column?

A block is considered structured (multi-column) if is_flowing_text_block detects at least two lines containing column gaps. Blocks with fewer than two column-gap lines are rendered as flowing single-column text, preserving the narrative flow of the document.

How does the algorithm handle text that doesn't align with established anchors?

Unanchored items are processed by try_align_floating in projection.rs (lines 1660-1690), which attempts to attach them to nearby anchors on adjacent lines within a configurable margin. This prevents headers, footers, or marginal notes from disrupting the main column structure.

What is the threshold for detecting a column separator?

line_has_column_gap (lines 1396-1410) identifies a column separator when a gap between consecutive items exceeds 2× the median character width and the items on either side of the gap cross the page's horizontal midpoint.

Where is the grid projection algorithm implemented?

The entire pipeline is implemented in crates/liteparse/src/projection.rs, with data structures defined in crates/liteparse/src/types.rs and high-level orchestration handled by crates/liteparse/src/parser.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →