How LiteParse's Bounding Box System Reconstructs PDF Layouts
LiteParse's bounding box system projects raw PDF text items onto a spatial grid using viewport coordinates, normalizes rotations to 90° increments, extracts column anchors, and merges geometrically adjacent boxes to reconstruct readable document structure.
LiteParse transforms raw PDF text extraction into structured, readable layouts through a sophisticated spatial projection pipeline. At the core of this process lies the bounding box system, implemented primarily in Rust within the crates/liteparse/src/projection.rs module according to the run-llama/liteparse source code. This system operates on low-level geometric data defined in crates/liteparse/src/types.rs to convert scattered text coordinates into coherent lines and blocks.
Storing Geometric Data: The TextItem Struct
Each piece of text extracted from a PDF page is stored as a TextItem struct containing precise geometric metadata:
pub struct TextItem {
pub text: String,
/// Viewport-space coordinates (top-left origin, 72 DPI)
pub x: f32,
pub y: f32,
pub width: f32,
pub height: f32,
/// Rotation in degrees (counter-clockwise, adjusted for page rotation)
pub rotation: f32,
/* … additional font-metadata omitted for brevity … */
}
The four geometric fields (x, y, width, height) describe the bounding box of the glyph run in PDF viewport coordinates at 72 DPI with a top-left origin. The rotation field preserves text orientation because many PDFs contain vertically-oriented labels, such as pin diagrams in technical documents. The full definition can be inspected in the source file crates/liteparse/src/types.rs.
Normalizing PDF Rotations
Before any layout logic runs, LiteParse normalises each item’s rotation to the nearest multiple of 90°. The canonical_rotation helper function, located in crates/liteparse/src/projection.rs at line 89, follows a tolerance of 2° and falls back to simple rounding:
fn canonical_rotation(rotation: f32) -> i32 { … }
This guarantees that subsequent processing only sees the four canonical orientations (0°, 90°, 180°, 270°), simplifying the spatial projection logic.
Grouping by Rotation and Spatial Clusters
The handle_rotation_reading_order function in crates/liteparse/src/projection.rs (line 112) first splits items into rotation groups. For 90° and 270° groups, it further partitions items by vertical gaps to prevent unrelated labels (such as top- and bottom-side pin numbers) from being merged.
The algorithm follows three steps:
- Build
groups_by_rotationmappingrotation → Vec<item_idx> - For 90°/270° groups, sort by
yand split whenever the gap exceeds three times the maximum item height - Sort groups left-to-right based on their minimum
xcoordinate
Column Anchor Detection
After rotation is flattened, LiteParse discovers anchor points—key X-coordinates that define column structure. The system identifies three anchor types for each text item:
- Left anchor:
item.x - Right anchor:
item.x + item.width - Center anchor:
item.x + item.width / 2
These coordinates are quantized to quarter-point units via the anchor_key function and stored in three HashMap<i32, Vec<(usize, usize)>> collections (anchor_left, anchor_right, anchor_center). The extraction occurs in extract_block_anchors at line 1014 of projection.rs:
let left_key = anchor_key(bbox.item.x);
let right_key = anchor_key(bbox.item.x + bbox.item.width);
let center_key = anchor_key(bbox.item.x + bbox.item.width * 0.5);
Cleaning Anchor Noise
Raw anchors often contain noise from isolated items or spurious columns. A series of filters prune them, located in crates/liteparse/src/projection.rs around lines 1150–1190:
delta_min_filter: Removes anchors that lack a neighbouring item withinpage_height × deltaintercept_filter: Discards anchors where text from other items crosses the anchor X-position between consecutive membersmerge_nearby_anchor_groups: Merges adjacent anchor keys within a tolerance of eight quarter-points
Forming Lines and Merging Boxes
With clean anchors, LiteParse forms lines by snapping items to a Y-grid (snap_y) and merging adjacent boxes belonging to the same visual line. The form_lines function at line 886 of projection.rs executes this logic:
- Computes a median text-height to determine Y-tolerance
- Sorts items by
(snap_y(y), x) - Merges boxes using
can_mergewhen Y-difference and height-difference are within small tolerances - Generates a vector of line vectors (
Vec<Vec<ProjectedTextItem>>)
Distinguishing Tables from Flowing Text
Documents often mix structured tables with flowing paragraphs. LiteParse detects flowing blocks by analyzing anchor density, line width, and column-gap statistics via is_flowing_text_block. When identified, render_flowing_block rewrites the items into a single string with proper indentation, ensuring paragraph-style text appears correctly in the final output. This logic resides in the latter half of projection.rs (lines 1240–1500).
Retrieving Bounding Boxes via the API
After projection, processed items populate each ParsedPage as text_items. The public API entry point LiteParse::parse_path in crates/liteparse/src/parser.rs returns a ParseResult containing pages with full bounding-box metadata.
Rust Example
use liteparse::{LiteParse, LiteParseConfig};
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Create a parser with default configuration
let parser = LiteParse::new(LiteParseConfig::default())?;
// Parse a PDF file
let result = parser.parse_path("examples/demo.pdf")?;
// Iterate over pages and print each item's bounding box
for page in result.pages {
println!("--- page {} ---", page.page_number);
for item in page.text_items {
println!(
"text: \"{}\" bbox: ({:.1}, {:.1}) w={:.1} h={:.1} rot={:.0}",
item.text, item.x, item.y, item.width, item.height, item.rotation
);
}
}
Ok(())
}
Python Example
from liteparse import LiteParse
parser = LiteParse()
result = parser.parse_path("examples/demo.pdf")
for page in result.pages:
print(f"--- page {page.page_number} ---")
for item in page.text_items:
print(
f'text="{item.text}" bbox=({item.x:.1f}, {item.y:.1f}) '
f'w={item.width:.1f} h={item.height:.1f} rot={item.rotation:.0f}'
)
Both bindings expose the same TextItem fields defined in Rust, providing identical bounding-box information across language front-ends.
Summary
- LiteParse's bounding box system operates on
TextItemstructs containing viewport coordinates (72 DPI, top-left origin) and rotation data stored incrates/liteparse/src/types.rs. - The projection pipeline in
crates/liteparse/src/projection.rsnormalizes rotations to 90° increments viacanonical_rotationbefore spatial grouping. - Column structure emerges from anchor detection (left, right, center) quantized to quarter-point units, filtered for noise, and used to form lines via
form_lines. - The system distinguishes flowing text from rigid tables using density analysis, ensuring paragraphs render correctly while preserving spatial metadata.
- Access bounding boxes programmatically through
LiteParse::parse_path, which returns projectedtext_itemswith complete geometric data in both Rust and Python.
Frequently Asked Questions
What coordinate system does LiteParse use for bounding boxes?
LiteParse uses PDF viewport coordinates with a top-left origin at 72 DPI. Each TextItem stores x, y, width, and height as f32 values representing points in this coordinate space, while rotation is stored in degrees counter-clockwise adjusted for page rotation.
How does LiteParse handle rotated text in PDFs?
The system normalizes all rotations to canonical 90° increments using the canonical_rotation function with a 2° tolerance. Items are then grouped by rotation via handle_rotation_reading_order, which additionally splits 90° and 270° groups by vertical gaps to prevent merging unrelated vertical labels.
What is the purpose of anchor detection in LiteParse?
Anchor detection discovers column boundaries by identifying shared X-coordinates (left, right, and center edges) among text items. These anchors are quantized to quarter-point units and filtered to remove noise, enabling the system to reconstruct tabular layouts and align text into proper reading order.
How can I access bounding box data programmatically?
Import LiteParse and call parse_path() to receive a ParseResult containing ParsedPage objects. Each page exposes a text_items vector where every item contains x, y, width, height, and rotation fields. Both the Rust crate and Python bindings provide identical access to this geometric metadata.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →