Anchor-Based Layout Detection in LiteParse: How PDF Structure Is Reconstructed
LiteParse reconstructs document layouts by projecting text boxes into a uniform coordinate system, discovering horizontal alignment anchors (left, right, and center), and snapping content to a filtered grid that preserves the original columns, paragraphs, and tables.
LiteParse, an open-source PDF parsing library maintained by the run-llama organization, solves complex document layout reconstruction through a sophisticated anchor-based layout detection algorithm implemented in Rust. Unlike linear text extraction tools that ignore spatial structure, this engine quantizes horizontal positions into alignment anchors to identify column boundaries and reading order without external OCR dependencies. The core logic resides in crates/liteparse/src/projection.rs, where raw PDF coordinates transform into semantic document structure through a seven-stage pipeline.
The Anchor Extraction Pipeline
The layout engine processes each page through discrete stages that gradually refine raw coordinates into structured text blocks. This approach guarantees that multi-column documents, mixed-format pages, and tables retain their visual hierarchy in the final output.
Stage 1: Block Segmentation
The algorithm first isolates independent layout regions by splitting the page into logical blocks separated by double blank lines. This segmentation prevents formatting in one region from corrupting the analysis of adjacent content. The implementation uses the segment_blocks function (lines 62‑106) to identify these boundaries before anchor extraction begins.
Stage 2: Anchor Extraction and Quantization
For each detected block, the system builds three distinct maps: anchor_left, anchor_right, and anchor_center. These maps store anchor keys—quantized X-positions represented as quarter-point integers—that index every non-rotated text item by its line and box indices. The extract_block_anchors function (lines 1007‑1044) handles this extraction, utilizing anchor_key (lines 92‑94) to convert coordinates into discrete keys and anchor_to_x (lines 96‑98) to reverse the conversion when needed.
Stage 3: Merging Nearby Anchor Groups
PDF extraction and OCR artifacts often generate slightly offset coordinates for text that should align to the same column. The merge_nearby_anchor_groups function (lines 38‑76) consolidates anchor groups that lie within a configurable tolerance, creating unified column boundaries from fragmented inputs. This merging step ensures that minor positional variations do not create spurious columns in the final layout.
Stage 4: Filtering Spurious Anchors
Two specialized filters prune anchors that do not represent genuine column boundaries:
- Vertical isolation filter (
delta_min_filter, lines 48‑90): Removes anchors lacking neighboring text within a proportion of the page height, eliminating stray marks or headers that lack columnar context. - Interception filter (lines 99‑145): Discards anchors that are crossed by other text boxes between consecutive anchor members, preventing line separators or borders from being mistaken for text columns.
Stage 5: Floating-Item Alignment
Not all content belongs to a dominant anchor. Isolated symbols, small graphics, or marginal annotations may float between established columns. The try_align_floating function (lines 154‑176) optionally snaps these items to the nearest anchor on adjacent lines if they fall within a configurable margin, integrating floating content into the document flow without disrupting the primary grid.
Stage 6: Column and Flowing-Text Detection
The system distinguishes between tabular data and flowing paragraphs using width and gap heuristics:
line_max_gap(lines 28‑36) calculates the largest horizontal gap on a line to identify potential column separators.line_has_column_gap(lines 40‑58) evaluates whether a gap represents a true column boundary relative to the median character width.is_flowing_text_block(lines 62‑110) combines anchor counts and width metrics to determine if a block should render as continuous paragraph text rather than a structured table.
Stage 7: Rendering the Final Layout
With anchors finalized, the algorithm constructs the textual representation. The render_line_as_flowing_text function (lines 124‑146) builds each line string with indentation derived from the leftmost anchor, while render_flowing_block (lines 146‑166) applies this rendering across all lines in a block. This stage outputs the column-aware, indentation-preserved text that represents the original PDF layout.
Working with LiteParse in Python and Node.js
Both language bindings ultimately invoke the Rust core described above. You can parse documents and inspect the generated layout using either environment.
Parse a PDF and inspect the line-by-line layout using the Python wrapper:
from liteparse import LiteParse
parser = LiteParse()
result = parser.parse("sample.pdf", output_format="text") # → plain-text output
# The internal layout can be examined via the JSON output:
json_result = parser.parse("sample.pdf", output_format="json")
for page in json_result["pages"]:
print(f"--- Page {page['page_number']} ---")
for line in page["lines"]:
print(line["text"])
The same operation using the Node.js N-API bindings:
import { LiteParse } from "liteparse";
(async () => {
const parser = new LiteParse();
const json = await parser.parse("sample.pdf", { output: "json" });
json.pages.forEach(p => {
console.log(`--- Page ${p.page_number} ---`);
p.lines.forEach(l => console.log(l.text));
});
})();
Summary
- Anchor-based layout detection quantizes horizontal positions into left, right, and center alignment maps to reconstruct PDF structure.
- The pipeline executes in
crates/liteparse/src/projection.rsthrough seven stages: block segmentation, anchor extraction, merging, filtering, floating-item alignment, column detection, and rendering. - Quantization via
anchor_keyconverts continuous X-coordinates into discrete quarter-point integers for reliable alignment grouping. - Dual filtering (vertical isolation and interception) eliminates false column boundaries caused by borders, headers, or OCR noise.
- Language-agnostic bindings in Python and Node.js expose the Rust core engine while preserving the anchor-based layout metadata in JSON outputs.
Frequently Asked Questions
What is anchor-based layout detection in LiteParse?
Anchor-based layout detection is a spatial analysis method that identifies horizontal alignment patterns (left, right, and center anchors) within PDF text boxes to reconstruct the original document's columns and reading order. By quantizing X-coordinates into discrete anchor keys and filtering spurious alignments, LiteParse distinguishes genuine column boundaries from decorative borders or misaligned text.
How does LiteParse handle text that doesn't align to any anchor?
The engine identifies floating items—such as isolated symbols or marginal annotations—that do not belong to surviving anchors. Through the try_align_floating function, these items can be snapped to the nearest anchor on adjacent lines if they fall within a configurable margin, or they can remain as distinct elements depending on the parsing configuration.
What prevents LiteParse from detecting false columns in complex layouts?
Two filters protect against false positives: the vertical isolation filter (delta_min_filter) removes anchors lacking vertical neighbors within a page-height proportion, and the interception filter discards anchors that are crossed by other text boxes between members. Together, these ensure that lines, borders, or page headers are not mistaken for column separators.
How can I access the internal anchor structure programmatically?
While the plain-text output provides the final rendered layout, you can access the underlying structure—including line positions and anchor metadata—by specifying the JSON output format. Both the Python and Node.js APIs support output_format="json" (or { output: "json" } in Node.js), which returns the parsed pages with detailed line objects containing text content and positional data derived from the anchor detection pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →