# How LiteParse Detects and Differentiates Flowing Text from Column-Based Layouts in PDFs

> Discover how LiteParse intelligently distinguishes flowing text from column layouts in PDFs using advanced anchor analysis and heuristic checks. Optimize your document parsing.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-06

---

**LiteParse detects flowing text versus column-based layouts by splitting each PDF page into blocks, extracting vertical anchors, and applying a strict heuristic that checks anchor counts, line-width ratios, and column-gap thresholds in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs).**

LiteParse parses PDF pages into a grid of anchored text boxes before deciding whether each block should render as a single flowing paragraph or as structured column-aligned content. According to the run-llama/liteparse source code, the core logic lives in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs), where functions like `is_flowing_text_block` and `detect_and_render_flowing_lines` evaluate geometric statistics to detect and differentiate between flowing text and column-based layouts in PDFs.

## Block Creation: Segmenting the Page

The pipeline begins by transforming raw PDFium boxes into lines and then grouping them into independent blocks.

In [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) at lines 6263–6295, `segment_blocks` takes the output of `form_lines` and partitions consecutive non-empty lines into `LineRange` objects:

```rust
let blocks = segment_blocks(&lines);

```

Because each `LineRange` is processed independently, LiteParse can handle pages that mix flowing paragraphs with columnar sections on the same page. A single block might become a plain paragraph, while its neighbor is snapped to multiple vertical anchors.

## Extracting Vertical Anchors from Each Block

For every block, the algorithm extracts three anchor maps that represent reliable vertical alignment columns.

At lines 6458–6478, `extract_block_anchors` returns left, right, and center anchor collections:

```rust
let (mut anchor_left, mut anchor_right, mut anchor_center) =
    extract_block_anchors(&lines, block);

```

Each bounding box contributes its x-coordinate rounded to a quarter-point grid by `anchor_key`. Rotated items are ignored because they would corrupt the column model. After extraction, the code cleans the anchor maps through `merge_nearby_anchor_groups`, `delta_min_filter`, and `intercept_filter`. The surviving anchors represent vertical lines shared by **multiple** items, making them strong signals of a structured, column-based layout.

## The Flowing-Text Heuristic

The primary decision function, `is_flowing_text_block`, spans lines 1263–1312 in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs). It evaluates a block against several hard thresholds defined near the top of the file.

### Anchor Count and Line-Width Statistics

The function applies the following criteria in order:

- **Anchor scarcity** — If the total number of distinct anchors exceeds `FLOWING_MAX_TOTAL_ANCHORS` (4), or if left anchors exceed `FLOWING_MAX_LEFT_ANCHORS` (3), the block is classified as structured.
- **Minimum line count** — A flowing block must contain at least `FLOWING_MIN_LINES` (3) non-empty lines.
- **Column-gap lines** — If **two or more** lines inside the block register a column gap via `line_has_column_gap`, the block is treated as multi-column.
- **Wide-line dominance** — `FLOWING_WIDE_LINE_RATIO` (0.5) defines a "wide" line as one whose span (`line_end - line_start`) exceeds 50 % of `page_width`. At least `FLOWING_WIDE_LINE_THRESHOLD` (0.6), or 60 %, of the lines must be wide for the block to qualify as flowing.

If all checks pass, the block is handed to `render_flowing_block` (lines 1500–1510), which concatenates the lines into a single paragraph string with indentation derived from the leftmost x-position.

### Column-Gap Detection with `line_has_column_gap`

Column gaps are large horizontal spaces that separate word clusters on the same line. The helper `line_has_column_gap` (lines 1440–1460) inspects adjacent items:

```rust
let gap = cur_start - prev_end;
if gap > median_width * 2.0 && prev_end < midpoint && cur_start > midpoint {
    return true; // column gap detected
}

```

Here, `median_width` is the median character width computed by `compute_median_textbox_size`. The gap must be **twice the median character width** and must straddle the page midpoint. This midpoint test makes the heuristic resilient to modest absolute gaps and reliably identifies two-column layouts.

## Per-Line Flow Detection Inside Structured Blocks

Even when a block is classified as structured, individual lines inside it may still be flowing text. For example, a multi-column page can contain a full-width introductory paragraph.

The second-pass function `detect_and_render_flowing_lines` (lines 1630–1690) handles this case.

### Initial Candidate Criteria

A line becomes a flowing candidate only if it satisfies **all** of the following conditions:

- It contains at least `FLOWING_MIN_LINE_ITEMS` (3) items.
- Its span exceeds `page_width * FLOWING_WIDE_LINE_RATIO` (0.5).
- Its maximum internal gap is smaller than `median_width * FLOWING_COLUMN_GAP_MULTIPLIER` (4.0).
- It does **not** contain a column gap according to `line_has_column_gap`.
- All items belong to the same snap kind; there are no mixed anchors.

### Propagation and Rendering

Once candidates are identified, forward and backward sweeps propagate flowing status to neighboring lines that share similar gaps and anchor consistency. These lines are then rendered with `render_line_as_flowing_text`, which inserts spaces when horizontal gaps exceed `FLOWING_SPACE_HEIGHT_RATIO` (0.15) and caps indentation at `FLOWING_MAX_INDENT` (8) characters.

## Parsing a PDF with the LiteParse Rust API

The projection pipeline runs automatically when you parse a document through the high-level `LiteParse` API in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs).

```rust
use liteparse::parser::{LiteParse, LiteParseConfig};

fn main() -> anyhow::Result<()> {
    let config = LiteParseConfig::default();
    let mut parser = LiteParse::new(config)?;

    let doc = parser.parse_path("samples/multi_column.pdf")?;

    for (i, page) in doc.pages.iter().enumerate() {
        println!("--- Page {} ------------------------------------------------", i + 1);
        println!("{}", page.text);
    }
    Ok(())
}

```

Under the hood, `parse_path` calls `project_pages_to_grid`, which executes the block segmentation, anchor extraction, and flowing-text heuristics described above. No manual configuration is required to trigger the flowing-versus-column detection.

## Summary

- LiteParse uses `segment_blocks` in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) to divide each PDF page into independent line ranges.
- `extract_block_anchors` builds left, right, and center anchor maps on a quarter-point grid, filtering out noise and rotated items.
- `is_flowing_text_block` decides a block’s fate by enforcing strict thresholds: no more than four total anchors, at least three lines, fewer than two column-gap lines, and at least 60 % wide lines.
- `line_has_column_gap` identifies two-column layouts by requiring a gap larger than twice the median character width that straddles the page midpoint.
- `detect_and_render_flowing_lines` performs a second pass inside structured blocks to rescue individual lines that are actually flowing paragraphs.
- The `LiteParse` API in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) exposes this behavior through a single `parse_path` call.

## Frequently Asked Questions

### What is the main Rust function that decides whether a PDF block is flowing text?

The function `is_flowing_text_block` in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) (lines 1263–1312) is the primary gatekeeper. It checks anchor counts, line widths, and column-gap presence to classify a `LineRange` as either a single flowing paragraph or a structured block that requires column snapping.

### How does LiteParse detect column gaps inside a single line?

LiteParse uses `line_has_column_gap` at lines 1440–1460 of [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs). It measures the horizontal space between adjacent items on a line; if that gap is more than twice the median character width and crosses the page midpoint, the line is flagged as containing a column gap.

### Can a multi-column PDF page still contain flowing paragraphs?

Yes. Even when a block is classified as structured, the second-pass function `detect_and_render_flowing_lines` (lines 1630–1690) scans for individual wide lines that lack mixed anchors and column gaps. These lines are extracted and rendered as flowing text before the remaining items are snapped to anchors.

### What thresholds control the flowing-text detection?

Key constants defined at the top of [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) include `FLOWING_MAX_TOTAL_ANCHORS` (4), `FLOWING_WIDE_LINE_RATIO` (0.5), `FLOWING_WIDE_LINE_THRESHOLD` (0.6), and `FLOWING_COLUMN_GAP_MULTIPLIER` (4.0). Together they enforce the rule that flowing text must have few anchors, many wide lines, and no significant column gaps.