Grid Projection Algorithm in projection.rs: How LiteParse Reconstructs PDF Layouts
The grid projection algorithm in crates/liteparse/src/projection.rs transforms raw PDF text boxes into a structured grid of lines, columns, and blocks through a multi-stage pipeline that includes median-based sizing, rotation normalisation, anchor filtering, and flowing-text detection.
LiteParse, an open-source PDF parser from the run-llama/liteparse repository, recovers the visual structure of documents by turning PDFium output into readable, structured text. The grid projection algorithm lives entirely in crates/liteparse/src/projection.rs and serves as the engine that quantises spatial coordinates, filters layout noise, and resolves ambiguous multi-column or rotated text. Downstream consumers receive either clean plain-text paragraphs or precise JSON blocks with explicit spatial metadata.
Preprocessing: Median Sizes, Rotation, and Margin Noise
Before the algorithm builds the logical grid, it normalises the raw text boxes extracted from the PDF page.
Median Size Estimation (compute_median_textbox_size)
At lines 31–88 in projection.rs, compute_median_textbox_size calculates the median character width and height across all items on the page. These median values become the baseline tolerances for later snapping, merging, and column-detection stages.
Rotation Normalisation (handle_rotation_reading_order)
Found around lines 113–210, this function detects text boxes rotated at 90°, 180°, or 270° and rewrites their coordinates so the remaining pipeline treats every item as upright. It also groups rotated items to preserve reading order, and for 180° text it re-orders items by X‑position to maintain logical flow.
Margin-Line Cleanup (clean_projected_items)
At lines 332–384, clean_projected_items removes spurious line numbers and artifacts that appear only in page margins. Eliminating this margin noise early prevents false anchors from breaking column detection later.
Core Grid Formation: Lines, Blocks, and Column Anchors
Once preprocessing is complete, the algorithm forms the actual grid by clustering items into lines, segmenting blocks, and extracting stable column anchors.
Line Formation and Merging (form_lines)
Between lines 86–200, form_lines executes two critical tasks. First, it snaps Y‑coordinates to a grid using y_sort_tolerance so items that belong on the same visual line receive identical sort keys. Second, it merges adjacent boxes that effectively represent the same word—cases with tiny gaps or overlapping extents—into a single ProjectedTextItem.
Block Segmentation (segment_blocks)
Around lines 62–105, segment_blocks divides the page into logical blocks by looking for double-blank-line gaps. This segmentation isolates paragraphs, tables, diagrams, and other distinct regions before column analysis begins.
Anchor Extraction (extract_block_anchors)
At lines 107–144, the algorithm constructs three anchor maps per block: left, right, and centre. Each map records X‑positions as quantised quarter-point keys using anchor_key = (x * 4).round(), creating stable column markers for alignment. Rotated items are intentionally skipped during this extraction.
Anchor Filtering (delta_min_filter and intercept_filter)
Raw anchors are refined in two passes to remove unstable markers. The delta_min_filter at lines 449–489 drops anchors that are isolated vertically, while the intercept_filter at lines 499–545 discards anchors that are crossed by other text. Together these filters ensure only genuine column boundaries survive into the final grid.
Layout Resolution: Floating Boxes and Text Flow
After the core grid is established, the algorithm resolves ambiguous items and determines whether each block should be treated as prose or structured data.
Floating-Box Alignment (try_align_floating)
At lines 552–627, try_align_floating handles items that do not match any existing anchor. The function attempts to snap these floating boxes to the nearest anchor on the line directly above or below, provided the distance falls within a configurable margin. This catches headers, footnotes, and inset labels that would otherwise remain orphaned.
Flowing-Text Detection (is_flowing_text_block)
Found at lines 626–672, this function classifies each block as either a flowing paragraph or a structured layout such as a multi-column table. It evaluates anchor counts, line-wide ratios, and column-gap heuristics to make the determination.
Rendering Flowing and Structured Blocks
The final stage dispatches each block to a renderer based on its classification.
Plain-Text and JSON Output
Inside lines 680–778, render_flowing_block and render_line_as_flowing_text convert flowing paragraphs into plain text with indentation derived from the median width. Structured blocks retain their column anchors and are serialised as JSON objects containing explicit x, y, width, and height fields for every box.
Using the Grid Projection Algorithm from Rust, Python, and Node.js
Although projection.rs performs the heavy lifting automatically, you trigger the pipeline by calling parse_path on a LiteParse instance. The examples below consume the result.lines field, which contains the resolved Vec<Vec<ProjectedTextItem>> grid.
use liteparse::{LiteParse, LiteParseConfig};
let cfg = LiteParseConfig::default();
let parser = LiteParse::new(cfg);
let result = parser.parse_path("example.pdf")?;
// `result.lines` is the line-grid produced by `projection.rs`
for line in result.lines {
let text = line.iter()
.map(|item| &item.item.text)
.collect::<Vec<_>>()
.join(" ");
println!("{text}");
}
from liteparse import LiteParse
parser = LiteParse()
result = parser.parse_path("example.pdf")
# `result.lines` is a list of lists of dicts (each dict mirrors ProjectedTextItem)
for line in result.lines:
txt = " ".join(item["text"] for item in line)
print(txt)
import { LiteParse } from "liteparse";
(async () => {
const parser = new LiteParse();
const res = await parser.parsePath("example.pdf");
for (const line of res.lines) {
console.log(line.map((b) => b.item.text).join(" "));
}
})();
How the Pipeline Fits Together
According to the run-llama/liteparse source code, parser.rs extracts raw items through extract_raw_items and then hands the list to project_items, which orchestrates the functions in projection.rs in the order described above. The resulting ParseResult contains fully resolved lines and spatial metadata ready for downstream processing.
Supporting files include:
crates/liteparse/src/types.rs— DefinesProjectedTextItem,TextItem, and the spatial structures used throughout the projection stage.crates/liteparse/src/parser.rs— Orchestrates the overall parsing flow and forwards raw items to the projection layer.crates/liteparse/src/config.rs— Holds default tolerances such asFLOATING_SPACES,COLUMN_SPACES, and flow thresholds.
Summary
- The grid projection algorithm in
crates/liteparse/src/projection.rsrebuilds PDF layouts through a 10-step pipeline of normalisation, snapping, anchoring, and rendering. - Median-based tolerances computed by
compute_median_textbox_sizeprovide robust defaults that adapt to each page. - Rotation handling in
handle_rotation_reading_orderensures 90°, 180°, and 270° text is re-oriented before grid formation. - Anchor extraction and filtering quantise X‑positions to quarter-points and discard false markers, enabling reliable column detection.
- Flowing-text detection decides whether a block becomes indented plain text or structured JSON with explicit coordinates.
Frequently Asked Questions
What is the grid projection algorithm in projection.rs?
The grid projection algorithm is the spatial layout engine inside LiteParse that converts raw PDF text boxes into ordered lines, columns, and blocks. It lives in crates/liteparse/src/projection.rs and uses a pipeline of median-based sizing, rotation normalisation, line merging, anchor extraction, and flow detection to reconstruct the visual structure of a page.
How does the grid projection algorithm handle rotated PDF text?
The algorithm normalises rotated items before grid formation through handle_rotation_reading_order in projection.rs. It detects 90°, 180°, and 270° rotations, rewrites coordinates so all text appears upright, and re-orders 180° items by X‑position to preserve reading order.
Where are the tolerances for snapping and merging configured?
Default tolerances such as FLOATING_SPACES, COLUMN_SPACES, and flow thresholds are defined in crates/liteparse/src/config.rs. The actual snapping tolerance used during line formation is derived from the median text-box size computed by compute_median_textbox_size.
How do I access the projected grid output in Python?
Call parse_path on a LiteParse instance and iterate over result.lines, which is a list of lists where each inner element is a dictionary mirroring ProjectedTextItem. You can concatenate the "text" fields of each item to reconstruct the visual line as shown in the Python example above.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →