# How AnchorMap in LiteParse Tracks Text Alignment Across Lines

> Discover how LiteParse's AnchorMap tracks text alignment across lines using a quantized hash map for consistent column layouts in PDFs. Learn about its multi-stage extraction process.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: internals
- Published: 2026-06-06

---

**The AnchorMap in LiteParse functions as a quantized hash map that groups text boxes by their horizontal alignment—left, right, or center—to maintain consistent column layouts across multi-line PDF blocks through a multi-stage extraction, cleaning, and snapping pipeline.**

LiteParse, the Rust-based PDF parsing library developed by LlamaIndex, reconstructs spatial document layouts using an **AnchorMap** system to resolve text alignment across different lines. By quantizing floating-point X-coordinates into integer keys and organizing text boxes into three distinct alignment categories, the library ensures consistent indentation and column detection even in complex documents containing rotated text or multi-column layouts. The implementation resides primarily in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) and [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs).

## Core Data Structures: Anchor and AnchorMap

The alignment tracking system rests on two fundamental definitions in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs).

The **`Anchor`** enum defines the three possible horizontal alignment directions at lines 87-91:

```rust
pub enum Anchor {
    Left,
    Right,
    Center,
}

```

The **`AnchorMap`** type, defined at lines 94-96, serves as the container for grouping text boxes that share the same quantized position:

```rust
pub type AnchorMap = HashMap<i32, Vec<(usize, usize)>>;

```

This hash map stores vectors of `(line_index, box_index)` tuples keyed by quantized X-coordinates, allowing the system to quickly locate all text boxes that align vertically across different lines.

## Coordinate Quantization Strategy

Before populating the AnchorMap, LiteParse converts floating-point coordinates to discrete integers through the **`anchor_key`** function in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) (lines 92-94). This function multiplies the X-coordinate by 4 and rounds to the nearest integer, creating a "quarter-point" key:

```rust
fn anchor_key(x: f32) -> i32 {
    (x * 4.0).round() as i32
}

```

This quantization strategy compresses minor positional variations into shared integer keys, enabling the AnchorMap to treat nearly-aligned text boxes as belonging to the same anchor column.

## Extracting Anchors from Text Blocks

The **`extract_block_anchors`** function (lines 1010-1043 of [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs)) populates three separate AnchorMaps—one for left edges, one for right edges, and one for centers. While iterating through a block of lines, the function processes every non-rotated text box and inserts three entries into the respective maps:

- **Left map**: keyed by the box's left edge (`anchor_key(x)`)
- **Right map**: keyed by the box's right edge (`anchor_key(x + width)`)
- **Center map**: keyed by the box's midpoint (`anchor_key(x + width/2.0)`)

Each insertion records the `(line_index, box_index)` tuple, creating a spatial index of alignment relationships within the current block.

## The Cleaning Pipeline: Refining Raw Anchors

Raw AnchorMaps contain noise from floating-point variations and incidental alignments. LiteParse applies three sequential cleaning operations in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) to refine them.

### Merging Nearby Groups

The **`merge_nearby_anchor_groups`** function (lines 38-44) consolidates anchor keys that differ by only a few quarter-points. This step merges anchors that likely represent the same logical column but suffered from minor PDF rendering variations.

### Vertical Isolation Filtering

The **`delta_min_filter`** (lines 48-90) removes anchors that lack neighboring text within a configurable vertical threshold. Anchors appearing on only one line with no vertical continuation are likely accidental alignments rather than true column boundaries.

### Intercept Filtering

The **`intercept_filter`** (lines 93-145) eliminates anchors whose X-position is crossed by other text in any vertical interval. If text flows across a potential anchor line, that anchor cannot represent a true column boundary and is removed from the map.

## Forward Propagation for Cross-Line Consistency

After cleaning, LiteParse stores the surviving anchors in three **`BTreeMap`** structures—`forward_left`, `forward_right`, and `forward_center`—within the `project_to_grid` function (lines 1000-1015). These forward anchor maps enable **cross-block alignment inheritance**: later text blocks can reference anchors established in earlier blocks, allowing the system to recognize that a left-aligned paragraph on page two continues the alignment pattern from page one.

This forward propagation mechanism ensures page-wide consistency for left-aligned body text, right-aligned headers, and centered titles, even when visual gaps exist between sections.

## Snapping Text Items to Anchors

During the final rendering phase, LiteParse consults the cleaned AnchorMaps to assign alignment values to each text item. Near lines 770-792 of [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs), the algorithm looks up the quantized keys for each item's left, right, and center positions. When a match exists in the corresponding AnchorMap, the system sets the `ProjectedTextItem.anchor` field to `Anchor::Left`, `Anchor::Right`, or `Anchor::Center`, and snaps the item to that column position.

Downstream renderers use this anchor field to determine indentation, spacing, and column placement, producing the final structured text output.

## Practical Example: Working with AnchorMap

The following Rust example demonstrates how to extract and inspect AnchorMaps from a fabricated page layout:

```rust
use liteparse::projection::{extract_block_anchors, anchor_key};
use liteparse::types::{ProjectedTextItem, TextItem, Snap, Anchor};

fn dummy_page_items() -> Vec<ProjectedTextItem> {
    // Three boxes on the same line: left-aligned, centre-aligned, right-aligned.
    vec![
        ProjectedTextItem {
            item: TextItem { x: 10.0, width: 30.0, ..Default::default() },
            snap: Snap::Left, anchor: Anchor::Left, rotated: false, ..Default::default()
        },
        ProjectedTextItem {
            item: TextItem { x: 200.0, width: 30.0, ..Default::default() },
            snap: Snap::Center, anchor: Anchor::Center, rotated: false, ..Default::default()
        },
        ProjectedTextItem {
            item: TextItem { x: 380.0, width: 30.0, ..Default::default() },
            snap: Snap::Right, anchor: Anchor::Right, rotated: false, ..Default::default()
        },
    ]
}

fn main() {
    // Simulate a single-line block.
    let lines = vec![dummy_page_items()];
    let block = liteparse::projection::LineRange { start: 0, end: 1 };

    // Extract the three anchor maps.
    let (left_map, right_map, center_map) = extract_block_anchors(&lines, &block);

    // Look at the quantised keys.
    println!("Left key:   {:?}", left_map.keys().collect::<Vec<_>>());
    println!("Center key: {:?}", center_map.keys().collect::<Vec<_>>());
    println!("Right key:  {:?}", right_map.keys().collect::<Vec<_>>());

    // Expected output – keys correspond to x, x+width, and midpoint.
    // (e.g. 10.0 * 4 → 40, 380.0*4 → 1520, 215.0*4 → 860)
}

```

Running this snippet prints the quantized anchor keys (40, 1520, and 860 respectively), illustrating how the AnchorMap indexes each alignment position independently.

## Summary

- **AnchorMap** is defined as `HashMap<i32, Vec<(usize, usize)>>` in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) and stores grouped text box indices by quantized X-position.
- The **`anchor_key`** function converts floating-point coordinates to quarter-point integers, collapsing minor positional variations into shared alignment keys.
- **Three parallel maps** (left, right, center) track horizontal alignment independently, populated by `extract_block_anchors` in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs).
- A **three-stage cleaning pipeline** (`merge_nearby_anchor_groups`, `delta_min_filter`, and `intercept_filter`) removes spurious anchors and consolidates related ones.
- **Forward anchor maps** allow alignment information to propagate across text blocks, maintaining consistency for paragraphs that span multiple separated regions.
- Final **snapping logic** assigns `Anchor` enum values to text items, enabling downstream renderers to produce properly indented, column-aware output.

## Frequently Asked Questions

### What data structure does AnchorMap use in LiteParse?

The AnchorMap in LiteParse uses a standard Rust `HashMap<i32, Vec<(usize, usize)>>` as defined in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) at lines 94-96. The `i32` keys represent quantized X-coordinates, while the values store vectors of `(line_index, box_index)` tuples pointing to specific text boxes in the document structure.

### How does LiteParse quantize coordinates for the AnchorMap?

LiteParse quantizes coordinates using the **`anchor_key`** function in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) (lines 92-94), which multiplies the floating-point X-coordinate by 4 and rounds to the nearest integer. This converts positions like `10.0` into integer key `40`, creating quarter-point precision that groups nearly-aligned text boxes while preserving sufficient spatial accuracy for column detection.

### What is the purpose of the intercept_filter in LiteParse?

The **`intercept_filter`** function (lines 93-145 of [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs)) removes anchor candidates whose vertical line is crossed by other text at any point. If text flows horizontally across a potential anchor position, that position cannot represent a true column boundary, so the filter eliminates it from the AnchorMap to prevent incorrect alignment assignments during the snapping phase.

### How does LiteParse maintain alignment across separate text blocks?

LiteParse maintains cross-block alignment using **forward anchor maps** (`forward_left`, `forward_right`, `forward_center`) implemented as `BTreeMap` structures in `project_to_grid` (lines 1000-1015). After processing each block, the cleaned anchors are stored in these maps, allowing subsequent blocks to inherit alignment information from earlier ones. This enables the system to recognize consistent left-aligned paragraphs or centered headers even when visual gaps or page breaks separate the content.