Understanding the Grid Projection Algorithm in LiteParse's projection.rs

The grid projection algorithm reconstructs PDF page layouts by converting raw text boxes into a logical grid of lines and columns through a ten-step pipeline involving median size estimation, rotation normalization, and anchor-based spatial alignment.

The grid projection algorithm is the core layout engine in the run-llama/liteparse repository, transforming PDFium's raw coordinate extractions into structured, readable text flows. Implemented in crates/liteparse/src/projection.rs, this algorithm bridges low-level PDF coordinates with human-readable document structures by quantizing spatial relationships into a stable grid.

The Grid Projection Pipeline Overview

The algorithm processes text boxes through a deterministic sequence of filtering, snapping, and alignment stages. Each stage refines the spatial metadata until the final output represents clean, column-aware text lines.

The pipeline begins with compute_median_textbox_size (lines 31-88) to establish baseline tolerances, then proceeds through rotation handling, noise removal, and geometric clustering. The final result is a Vec<Vec<ProjectedTextItem>> representing the logical line grid of the document.

Step 1: Median Size Estimation

Before any geometric operations, the algorithm establishes robust baseline measurements. The compute_median_textbox_size function calculates the median character width and height across all text boxes on the page.

These median values serve as the foundation for tolerance calculations used in later snapping and merging operations. By using median rather than mean values, the algorithm remains resistant to outliers from oversized headers or tiny footnotes.

Step 2: Rotation Normalisation

PDFs often contain rotated text elements (90°, 180°, or 270°). The handle_rotation_reading_order function (lines 113-210) detects these rotations and rewrites their coordinates so subsequent pipeline stages can treat all text as upright.

For 180° rotated text, the function reorders items by X-position to maintain correct reading order. For 90° and 270° rotations, it repositions coordinates while preserving the relative grouping of rotated elements, ensuring vertical labels don't disrupt horizontal line detection.

Step 3: Margin-Line Cleanup

Spurious line numbers frequently appear in page margins and can break column detection algorithms. The clean_projected_items function (lines 332-384) identifies and removes these marginal artifacts before they contaminate the spatial analysis.

Step 4: Line Formation and Merging

The form_lines function (lines 86-200) performs two critical operations:

  • Snapping: Y-coordinates are snapped to a grid using y_sort_tolerance (derived from median height), ensuring that items visually aligned on the same horizontal line receive identical sort keys.
  • Merging: Adjacent boxes representing the same word—separated by tiny gaps or overlapping slightly—are merged into single ProjectedTextItem instances. This eliminates fragmentation caused by PDF font subsetting or kerning variations.

Step 5: Block Segmentation

Pages are divided into logical blocks using segment_blocks (lines 62-105). This function detects double-blank lines—vertical gaps larger than the standard line spacing—to isolate distinct content regions such as paragraphs, tables, and diagrams.

Step 6: Anchor Extraction

For each block, extract_block_anchors (lines 107-144) builds three anchor maps tracking left, right, and center alignment positions. Anchors are quantized using quarter-point keys (anchor_key = (x * 4).round()), creating a finite set of alignment columns that tolerate minor positional variations.

Rotated items are excluded from anchor extraction to prevent misalignment of vertical text with horizontal anchors.

Step 7: Anchor Filtering

Raw anchors often contain noise from isolated punctuation or marginalia. Two filters refine the anchor set:

  • Delta-Min Filter: Implemented in delta_min_filter (lines 449-489), this removes anchors that are vertically isolated—appearing on only one or two lines without vertical continuity.
  • Intercept Filter: The intercept_filter function (lines 499-545) discards anchors that would be crossed by other text lines, preventing false column detection when text actually flows through the proposed anchor position.

Step 8: Floating-Box Alignment

Not all text boxes align with major anchors. The try_align_floating function (lines 552-627) attempts to snap these floating items to the nearest anchor on adjacent lines, using configurable margins from FLOATING_SPACES and COLUMN_SPACES defined in crates/liteparse/src/config.rs.

Step 9: Flowing-Text Detection

The is_flowing_text_block function (lines 626-672) classifies each block as either flowing prose or structured layout. It analyzes anchor counts, line-wide coverage ratios, and column-gap heuristics to distinguish between single-column paragraphs and multi-column tables.

Step 10: Rendering

The final stage produces output based on block classification:

  • Flowing Blocks: Rendered as plain text using render_flowing_block and render_line_as_flowing_text (lines 680-778), with indentation calculated from the median character width.
  • Structured Blocks: Retain their quantized column anchors and export as JSON with explicit spatial coordinates (x, y, width, height) for each box.

Integration with the Parser

The parser.rs file orchestrates the entire process. It first calls extract_raw_items to obtain the raw PDFium output, then passes this data to project_items which invokes the grid projection pipeline in the sequence described above.

The resulting ParseResult contains the fully resolved Vec<Vec<ProjectedTextItem>> structure, where each inner vector represents a logical line of text with precise spatial metadata preserved from the original PDF coordinates.

Usage Examples

The following examples demonstrate how to invoke LiteParse. In all cases, the Rust core executes the grid projection algorithm in projection.rs to produce the structured output.

Rust

use liteparse::{LiteParse, LiteParseConfig};

let cfg = LiteParseConfig::default();
let parser = LiteParse::new(cfg);
let result = parser.parse_path("example.pdf")?;

// result.lines contains the grid produced by projection.rs
for line in result.lines {
    let text = line.iter()
                   .map(|item| &item.item.text)
                   .collect::<Vec<_>>()
                   .join(" ");
    println!("{text}");
}

Python

from liteparse import LiteParse

parser = LiteParse()
result = parser.parse_path("example.pdf")

# result.lines is a list of lists of dicts mirroring ProjectedTextItem

for line in result.lines:
    txt = " ".join(item["text"] for item in line)
    print(txt)

Node.js

import { LiteParse } from "liteparse";

(async () => {
  const parser = new LiteParse();
  const res = await parser.parsePath("example.pdf");
  for (const line of res.lines) {
    console.log(line.map((b) => b.item.text).join(" "));
  }
})();

Summary

The grid projection algorithm in projection.rs transforms raw PDF extractions into structured document layouts through these key mechanisms:

  • Median-based tolerances established by compute_median_textbox_size provide robust baseline measurements resistant to outliers
  • Rotation handling via handle_rotation_reading_order normalizes 90°, 180°, and 270° text to maintain reading order
  • Quantized anchor extraction using quarter-point keys enables stable column detection across minor positional variations
  • Dual filtering through delta_min_filter and intercept_filter removes false column markers caused by isolated text or crossing lines
  • Block classification distinguishes flowing prose from structured layouts, enabling appropriate rendering strategies

Frequently Asked Questions

What is the purpose of the quarter-point quantization in anchor extraction?

The quarter-point quantization (anchor_key = (x * 4).round()) converts continuous PDF coordinates into discrete alignment buckets. This technique groups text boxes that are visually aligned but may vary by small fractions of a point due to font rendering differences, creating stable column anchors without requiring exact coordinate matches.

How does the algorithm handle rotated text in PDFs?

The handle_rotation_reading_order function detects rotation angles and transforms coordinates accordingly. For 180° rotations, it reorders items by X-position to restore correct reading order. For 90° and 270° rotations, it repositions coordinates while maintaining the items' grouping, ensuring vertical labels don't interfere with horizontal line detection in subsequent pipeline stages.

What distinguishes flowing text blocks from structured blocks?

The is_flowing_text_block function analyzes anchor density, line-wide coverage ratios, and column-gap patterns. Flowing text blocks typically show consistent left or justified alignment with a single dominant anchor, while structured blocks exhibit multiple regular anchors with consistent gaps between them, indicating tabular or multi-column layouts.

Where are the tolerance values for snapping and merging configured?

Tolerance values originate from the median character dimensions computed by compute_median_textbox_size, with specific multipliers and thresholds defined in crates/liteparse/src/config.rs. Key parameters include FLOATING_SPACES for floating-box alignment margins and COLUMN_SPACES for column-gap detection, both configurable through LiteParseConfig.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →