How LiteParse's Spatial Grid Projection Algorithm Reconstructs Text Layout from PDF Coordinates

LiteParse's spatial grid projection algorithm converts raw PDF text boxes into a clean, row-by-row text grid by normalizing rotations, snapping coordinates to a median-height grid, merging adjacent characters into lines, and rendering flowing paragraphs while preserving original visual order.

LiteParse is an open-source Rust library developed by run-llama that extracts structured text from PDFs. Its spatial grid projection algorithm, implemented primarily in crates/liteparse/src/projection.rs, rebuilds the visual reading order from raw PDFium coordinates. The algorithm processes every page through a multi-stage pipeline that handles rotated text, margin noise, line formation, and paragraph flow detection before emitting plain-text lines.

Entry Point: project_pages_to_grid

After extract::extract_pages returns raw page data, the high-level parser in crates/liteparse/src/parser.rs invokes the projection engine around line 189:

let parsed_pages = projection::project_pages_to_grid(pages);

project_pages_to_grid iterates over each Page, wraps every raw TextItem in a ProjectedTextItem, and forwards the full collection to project_to_grid—the core layout routine. In projection.rs (lines 2654–2672), the wrapper preserves the original x and y coordinates and initializes snapping anchors:

pub fn project_pages_to_grid(pages: Vec<Page>) -> Vec<ParsedPage> {
    pages.into_iter().map(|page| {
        let projection_boxes = page.text_items.iter().map(|item| ProjectedTextItem {
            orig_x: item.x,
            orig_y: item.y,
            item: item.clone(),
            snap: Snap::Left,
            anchor: Anchor::Left,
            // ...
        }).collect();

        let (projected_items, text) = project_to_grid(&page, projection_boxes);
        // ...
    }).collect()
}

Normalizing Rotated Text

PDF text boxes may arrive rotated at 0°, 90°, 180°, or 270°. Before any line formation occurs, the pipeline re-orients all boxes so that downstream logic works exclusively with upright coordinates.

canonical_rotation (lines 889–911 in projection.rs) snaps any angle to the nearest cardinal rotation within a ±2° tolerance. Then handle_rotation_reading_order (lines 1133–1190) groups items by rotation, rotates 90° and 270° groups back to 0°, and merges groups that intersect non-rotated content. After this stage completes, every item carries rotation == 0.0 and a trustworthy rotated flag.

Cleaning Margins and Noise

Before the grid is built, clean_projected_items strips isolated margin artifacts such as tiny line numbers. The helper compute_median_textbox_size (lines 31–86) calculates the median width and height across all items on the page. Using that median, clean_projected_items (lines 332–384) discards small, isolated boxes located near the page center that match the line-number profile. This prevents margin noise from inserting false entries into the reading order.

Forming Visual Lines with form_lines

With clean, upright boxes, the algorithm groups them into visual lines through three operations defined in form_lines (lines 386–527):

  1. Snap Y-coordinates — snap_y rounds each box’s baseline to a shared grid using a tolerance of roughly half the median text height, never below 5 px.
  2. Sort — items are ordered by snapped Y, then by X, establishing the correct reading direction.
  3. Merge — adjacent boxes on the same baseline with gaps ≤ 0.1 px are merged into single words. During merging, merge_orig_bbox expands the original PDF bounding box so downstream OCR merging can map back to source coordinates.

Detecting and Rendering Flowing Paragraphs

PDFs often contain flowing paragraphs that should read as natural text rather than rigid coordinate dumps. LiteParse detects these blocks using anchor and width heuristics, then conditionally renders them as plain text.

extract_block_anchors (lines 1007–1044) builds left, right, and center anchor maps for each block. is_flowing_text_block (lines 1262–1305) then evaluates the number of anchors, line widths, and column gaps to decide whether a block is flowing text. When a block qualifies, render_flowing_block (lines 1349–1389) rewrites each line into a plain-text string, inserting spaces where gaps exceed a height-based threshold.

If the heuristics classify a block as non-flowing—such as a table—the algorithm preserves exact X and Y coordinates so output formatters like json.rs or text.rs can render distinct columns.

Final Text Assembly and Coordinate Preservation

project_to_grid returns two concrete outputs:

  • projected_items — the final list of ProjectedTextItem structs, each retaining the original PDF bounding box for later OCR merges.
  • text — a single string in which every line corresponds to a visual line on the page, with proper spacing and line breaks.

project_pages_to_grid packages these into a ParsedPage that stores both structured text_items for layout-aware consumers and plain text for simple downstream processing.

Rust Usage Example

You can invoke the entire projection pipeline through the LiteParse parser interface. The example below parses a PDF and inspects the grid-projected results:

use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;

#[tokio::main]
async fn main() -> Result<(), liteparse::error::LiteParseError> {
    let cfg = LiteParseConfig {
        ocr_enabled: false,
        ..Default::default()
    };
    let parser = LiteParse::new(cfg);

    let result = parser.parse("example.pdf").await?;

    println!("Full text:\n{}", result.text);

    for item in &result.pages[0].text_items {
        println!("‘{}’ at ({:.1},{:.1}) size ({:.1}×{:.1})",
                 item.text, item.x, item.y, item.width, item.height);
    }
    Ok(())
}

Internally, parser.parse triggers the extraction and project_pages_to_grid stages described above.

Summary

  • project_pages_to_grid in parser.rs orchestrates the spatial grid projection pipeline for every page.
  • Rotation normalization via canonical_rotation and handle_rotation_reading_order forces all boxes into upright coordinates before line formation.
  • Margin cleaning uses median text-box sizes to strip isolated line numbers and noise.
  • form_lines snaps Y coordinates to a median-height grid, sorts by reading order, and merges touching characters into words while preserving original bounding boxes.
  • Flowing-paragraph detection rewrites qualifying blocks into natural plain text; non-flowing blocks keep precise coordinates for table-like layouts.

Frequently Asked Questions

How does LiteParse handle text that is rotated inside a PDF?

LiteParse normalizes rotated text before line formation. The canonical_rotation helper snaps angles to the nearest 0°, 90°, 180°, or 270° with a ±2° tolerance, and handle_rotation_reading_order rotates 90° and 270° groups back to upright orientation. After this stage, every text box has rotation == 0.0 and a reliable rotated flag.

Why does the projection algorithm compute median text-box size?

The median width and height, calculated by compute_median_textbox_size, provide a stable baseline for two operations: determining the Y-axis snap tolerance in snap_y and identifying margin noise in clean_projected_items. Using the median rather than the mean prevents outlier boxes from skewing the grid.

What is the difference between flowing and non-flowing blocks in LiteParse?

Flowing blocks are standard paragraphs that LiteParse detects through anchor and width heuristics in is_flowing_text_block. These blocks are rendered by render_flowing_block into natural line-wrapped plain text. Non-flowing blocks—such as tables—retain exact X and Y coordinates so downstream formatters can preserve columnar structure.

How are original PDF coordinates preserved after merging text boxes?

During the merge step inside form_lines, the algorithm calls merge_orig_bbox to expand the original bounding box whenever two adjacent boxes are combined. This ensures that every ProjectedTextItem in the final output still carries a mapping to its source PDF coordinates, enabling accurate OCR merging later in the pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →