# How LiteParse Distinguishes Margin Line Numbers from Regular Text

> LiteParse separates margin line numbers from regular text using a boolean flag and heuristic detection. Learn how it prevents merging with paragraph content for cleaner data extraction.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-06

---

**LiteParse flags margin line numbers with a dedicated boolean field, detects them via a page-midpoint heuristic in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs), and isolates them during line formation so they never merge with regular paragraph content.**

LiteParse, an open-source PDF parser from the `run-llama/liteparse` repository, is designed to extract clean semantic text from complex academic layouts, including two-column documents with gutter digits. Understanding how LiteParse distinguishes margin line numbers from regular text content is essential for interpreting its structured JSON output and plain-text results. The entire pipeline is implemented in Rust and exposed through bindings for Python and Node.js.

## The Data Model: `is_margin_line_number`

In [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) at line 103, the `ProjectedTextItem` struct defines the boolean field `is_margin_line_number`.

This explicit flag ensures the distinction between margin metadata and body text persists through every downstream stage. When the parser projects a raw PDF text element into a `ProjectedTextItem`, the field starts as `false` and is toggled only after the layout engine confirms the element matches the margin-number heuristic.

## Heuristic Detection in `form_lines`

The detection logic lives in the `form_lines` function inside [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) (lines 395–433). While scanning projected items, the algorithm applies three filters simultaneously.

### Horizontal Center Relative to Page Midpoint

The engine first computes a narrow acceptance window around the horizontal center of the page:

```rust
let midpoint = page_width / 2.0;
let margin_left = midpoint - 5.0;
let margin_right = midpoint + 20.0;

```

For every candidate item, it measures the horizontal center:

```rust
let center = item.item.x + item.item.width / 2.0;

```

If `center` falls between `margin_left` and `margin_right`, the item passes the geometric test.

### Content Format and Width Constraints

Next, the helper `is_margin_line_number_text` (lines 401–422) validates the actual string content. It expects one to two digits, optionally followed by an `"O"` character to accommodate common OCR confusion. Finally, the item’s `width` must be less than `15` pixels.

Only items that satisfy position, content, and width criteria are marked with `is_margin_line_number = true`.

## Preventing Merges with Regular Text

Flagged margin numbers are blocked from joining normal text lines. Inside [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) (lines 35–40), the line-building routine evaluates `cur_line_has_margin` and `cur_item_has_margin`. If these two booleans differ, a `margin_mismatch` occurs and the current line is split before the new item is appended.

This rule guarantees that a margin digit will never be concatenated into the middle of a paragraph sentence. The separation preserves clean semantic boundaries between metadata and body content.

## Cleaning Orphaned Margin Numbers

After line formation, the `clean_projected_items` function ([`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs), lines 32–84) performs a final filtering pass. It removes isolated margin numbers that have no neighboring non-margin text on the same line.

Conversely, if a line contains any regular text alongside margin numbers, those numbers are retained. This preserves meaningful numbering in tables or diagrams while eliminating stray digits that would otherwise pollute the output.

## Working with Margin Numbers in Python

The Rust logic is fully exposed through LiteParse’s Python bindings. When parsing a two-column PDF, each element in the result list carries the `is_margin_line_number` field:

```python
from liteparse import LiteParse
import json

parser = LiteParse()
result = parser.parse("samples/two_column_with_margin_numbers.pdf")

print(json.dumps(result, indent=2))

```

A flagged item appears in the JSON output like this:

```json
[
  {
    "text": "1",
    "x": 8.2,
    "y": 42.0,
    "width": 6.0,
    "height": 10.0,
    "is_margin_line_number": true
  },
  {
    "text": "The quick brown fox",
    "x": 50.0,
    "y": 42.0,
    "width": 120.0,
    "height": 10.0,
    "is_margin_line_number": false
  }
]

```

Calling `parser.to_text()` omits margin numbers unless they sit on a line that also contains regular text:

```python
print(parser.to_text())

# Output: "The quick brown fox ..."

```

## Summary

- **`is_margin_line_number` flag:** Defined in [`types.rs`](https://github.com/run-llama/liteparse/blob/main/types.rs) at line 103, this boolean makes the margin-number distinction persistent across the entire pipeline.
- **Midpoint heuristic:** The `form_lines` routine in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) (lines 395–433) detects candidates by checking if their horizontal center lies within a 25-pixel window around `page_width / 2.0`, aided by `is_margin_line_number_text` (lines 401–422).
- **Anti-merge guard:** The `margin_mismatch` logic at lines 35–40 prevents margin items from being fused into normal lines.
- **Selective cleanup:** `clean_projected_items` (lines 32–84) discards orphaned margin numbers but keeps those that share a line with meaningful content.

## Frequently Asked Questions

### How does LiteParse determine whether a number is a margin line number or regular text?

LiteParse requires three simultaneous conditions: the item’s horizontal center must lie within a narrow band around the page midpoint, its text must be one to two digits (optionally ending in an OCR-artifact `"O"`), and its width must be under `15` pixels. These checks run inside [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) during the `form_lines` pass.

### Can margin line numbers remain in the final plain-text output?

Yes, but only when they share a line with non-margin text. The `clean_projected_items` stage removes isolated margin numbers yet preserves those that appear alongside legitimate content, such as in table rows or labeled diagrams.

### Where in the source code is the `is_margin_line_number` property defined?

The property is declared in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) at line 103 on the `ProjectedTextItem` struct. It is populated by the heuristic logic in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs).

### Why does LiteParse use the page midpoint to find margin numbers?

Two-column PDFs typically place line numbers in the gutter between columns, which is naturally near the horizontal center of the page. By computing `page_width / 2.0` and allowing a small tolerance window (`-5.0` to `+20.0`), LiteParse can reliably target the gutter without affecting body text, which sits far from the center in either the left or right column.