# How the OCR Merge Algorithm Works in LiteParse (`ocr_merge.rs`)

> Explore the OCR merge algorithm in LiteParse's ocr_merge.rs. Discover its two-stage pipeline for efficient OCR processing, parallel execution, and seamless text model integration.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: internals
- Published: 2026-06-25

---

**The OCR merge algorithm in LiteParse implements a two-stage pipeline that analyzes page complexity to determine which pages need OCR, renders those pages to bitmap, executes OCR in parallel with semaphore-controlled concurrency, and merges the results back into the PDF text model while deduplicating against existing native text.**

The [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) file in the `run-llama/liteparse` repository contains the core logic for bridging OCR engines with PDF text extraction. This Rust module handles the complete lifecycle from deciding which pages require optical character recognition to integrating the recognized text into the document's `TextItem` structure.

## Deciding Which Pages Need OCR

Before any bitmap rendering occurs, the algorithm evaluates each page using `calculate_page_complexity` to determine if OCR is necessary.

### The `calculate_page_complexity` Heuristics

Located at lines 15–33 in [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs), this function applies five distinct heuristics:

- **Scanned / No-text pages**: Fewer than 20 characters and a full-page image triggers `ComplexityReason::Scanned` or `::NoText`.
- **Sparse text**: Under 2000 characters with less than 15% coverage results in `ComplexityReason::SparseText`.
- **Embedded images**: Any raster image exceeding `MIN_IMAGE_SIZE_PT` triggers `ComplexityReason::EmbeddedImages`.
- **Garbled native text**: Pages failing the `page_is_garbled` vowel-ratio test receive `ComplexityReason::Garbled`.
- **Vector-only text**: When filled-path area not covered by native text exceeds `UNCOVERED_VECTOR_AREA_THRESHOLD`, the reason becomes `ComplexityReason::VectorText`.

If any heuristic triggers, `needs_ocr` is set to `true` (lines 200–202).

## Rendering and Parallel OCR Execution

### Rendering Pages for OCR

The `render_pages_for_ocr` function (lines 46–52) iterates over all pages, re-evaluates complexity, and for pages with `needs_ocr` set to true, calls `pdfium::Page::render(dpi)` at lines 55–60. This produces a `RenderedPage` struct (lines 26–32) containing raw RGB bytes, width, and height.

### Parallel Execution with Concurrency Control

`ocr_and_merge_rendered` orchestrates the OCR workload using async concurrency:

1. **Task spawning**: One async task per rendered page (lines 100–108).
2. **Semaphore throttling**: A semaphore limits concurrent workers to `num_workers` (lines 81–86), preventing deadlocks when OCR engines use blocking I/O.
3. **Blocking offloading**: Each task acquires a permit via `sem.acquire_owned().await` and moves the actual OCR call to `spawn_blocking` (lines 111–119), ensuring the semaphore throttles CPU-blocking work correctly.

The OCR engine returns a `Vec<OcrResult>` for each processed page.

## Error Handling and Systemic Failure Detection

The algorithm implements a guard against total OCR failure. After all tasks complete, it counts `failed_tasks` and checks if any failure occurred on a sparse-text page (`failed_sparse_text_page`).

If **every** OCR task fails **and** at least one failure occurred on a sparse-text page, the function returns `LiteParseError::Ocr` (lines 186–194). This prevents silent data loss when OCR is the primary text source but allows partial failures when native text exists.

## Merging OCR Results into the Text Model

For each successful page, the merging process follows a strict protocol:

### Cleaning Native Text Items

The algorithm first clears unusable native text. If the page is garbled, all native text is removed; otherwise, only items flagged by `is_unusable_native` are dropped (lines 84–96).

### Filtering and Processing Results

1. **Confidence filtering**: OCR results with `confidence ≤ 0.1` are discarded (lines 103–106).
2. **Bounding box calculation**: When the OCR result includes a polygon, the code extracts the axis-aligned bbox and derives rotation via `polygon_rotation_deg` (lines 113–129).
3. **Overlap detection**: The OCR bbox is compared against **only native text items** using `overlaps_existing_text` (lines 139–145). If overlap exists, the OCR result is ignored to prevent duplication.
4. **Artifact cleaning**: `clean_ocr_table_artifacts` strips spurious characters like `|`, `[`, and `]` from numeric-like OCR strings (lines 150–166).
5. **Font size estimation**: For rotated text, the narrower dimension serves as the font-size heuristic (lines 167–174).

### Creating TextItem Entries

Finally, the algorithm appends a `TextItem` to the page with fields: `text`, `x`, `y`, `width`, `height`, `rotation`, `font_name = "OCR"`, `font_size`, and `confidence` (lines 176–186).

## Key Helper Utilities

The algorithm relies on several specialized utility functions:

- **`polygon_rotation_deg`** (lines 260–320): Deduces text rotation (0°, 90°, 180°, 270°) from 4-point OCR polygons, handling both reading-direction and screen-axis orderings.
- **`overlaps_existing_text`** (lines 330–350): Checks bbox overlap with configurable tolerance.
- **`clean_ocr_table_artifacts`** (lines 340–380): Removes typical OCR misreads of table borders while preserving non-numeric content.
- **`is_unusable_native`** and **`page_is_garbled`** (lines 380–420): Detect corrupted native text using Unicode-map errors and vowel-ratio heuristics.

## Practical Implementation Examples

### Triggering OCR from Application Code

```rust
use liteparse::parser::LiteParse;
use std::sync::Arc;

// Assume `doc` is an opened pdfium::Document and `pages` is a Vec<Page>.
let dpi = 300.0;

// 1. Render pages that need OCR based on complexity heuristics.
let rendered = liteparse::ocr_merge::render_pages_for_ocr(&doc, &pages, dpi)?;

// 2. Initialize an OCR engine (Tesseract example).
let ocr_engine: Arc<dyn liteparse::ocr::OcrEngine> = 
    Arc::new(liteparse::ocr::tesseract::TesseractEngine::new()?);

// 3. Execute OCR and merge results back into the pages vector.
liteparse::ocr_merge::ocr_and_merge_rendered(
    &mut pages,
    rendered,
    dpi,
    ocr_engine,
    "eng",          // language code
    4,              // parallel workers (semaphore limit)
).await?;

```

### Inspecting Merged Results

```rust
for page in pages {
    println!("--- Page {} ---", page.page_number);
    for item in &page.text_items {
        println!(
            "{} ({:.1},{:.1}) [{}×{}] rot={:.0}° conf={:?}",
            item.text,
            item.x,
            item.y,
            item.width,
            item.height,
            item.rotation,
            item.confidence
        );
    }
}

```

OCR-derived items will display `font_name = "OCR"` and include confidence scores from the recognition engine.

## Summary

- **Complexity heuristics** determine OCR necessity using character counts, coverage percentages, and image detection in `calculate_page_complexity`.
- **Parallel execution** uses semaphore-controlled async tasks to prevent resource exhaustion while maximizing throughput.
- **Failure detection** aborts parsing only when OCR is the sole text source and completely fails.
- **Deduplication logic** compares OCR bounding boxes against native text items to prevent overlapping duplicates.
- **Post-processing** includes rotation detection, table artifact cleaning, and font size estimation before creating `TextItem` entries.

## Frequently Asked Questions

### How does LiteParse decide which pages need OCR?

The algorithm evaluates five heuristics in `calculate_page_complexity` (lines 15–33): scanned/no-text pages, sparse text coverage, embedded images, garbled native text detected via vowel ratios, and vector-only content. If any condition triggers, `needs_ocr` becomes true and the page is rendered for processing.

### What happens if OCR fails on all pages?

The `ocr_and_merge_rendered` function implements a systemic-failure guard (lines 186–194). If every OCR task fails and at least one failure occurs on a sparse-text page (where OCR is the primary text source), the function returns `LiteParseError::Ocr`. This prevents silent failures when OCR is critical while allowing partial recognition to continue.

### How does the algorithm prevent duplicate text from OCR and native PDF text?

During merging, each OCR result's bounding box is checked against **only native text items** using `overlaps_existing_text` (lines 139–145). If the OCR text overlaps with existing native text, the OCR result is discarded. This ensures the final document contains either the native text or the OCR text, never both.

### Why does the algorithm use a semaphore for OCR workers?

The semaphore (lines 81–86) prevents deadlocks and resource exhaustion when OCR engines use blocking I/O. By limiting concurrent workers to `num_workers` and acquiring permits via `sem.acquire_owned().await` before calling `spawn_blocking` (lines 111–119), the algorithm ensures that CPU-intensive OCR work does not overwhelm the async runtime while maintaining parallel throughput.