# How LiteParse Combines OCR Results with Native PDF Text: A Technical Deep Dive

> Discover how LiteParse seamlessly merges OCR and native PDF text. Learn its advanced techniques for accurate document processing and deduplication.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-07

---

**LiteParse merges OCR results with native PDF text by first extracting native content, then conditionally rasterizing text-poor pages for OCR, and finally deduplicating overlapping text items using geometric overlap detection before unifying them into a single `Page` structure.**

The run-llama/liteparse repository implements a sophisticated hybrid parsing strategy that combines native PDF text extraction with optical character recognition (OCR) to handle scanned documents and text-poor layouts. When processing PDFs, the parser seamlessly integrates machine-readable text with OCR-derived content while avoiding duplicates and preserving spatial accuracy. This article examines the specific implementation details found in the Rust source code, focusing on how the library combines OCR results with native PDF text to produce unified output.

## The Two-Stage Hybrid Extraction Pipeline

The `LiteParse` engine operates in distinct phases orchestrated in [`parser.rs`](https://github.com/run-llama/liteparse/blob/main/parser.rs). First, it extracts native text directly from the PDF structure. Then, if OCR is enabled via `LiteParseConfig`, it invokes the merge logic from [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) to supplement missing content.

The pipeline follows this sequence:

1. Extract native `TextItem`s from the PDF
2. Identify pages requiring OCR based on content density
3. Rasterize selected pages to RGB bitmaps
4. Execute OCR concurrently across worker threads
5. Merge valid OCR results back into the native `Page` structures

## Detecting When Pages Require OCR

The system uses heuristic rules to avoid unnecessary OCR processing on text-rich digital documents.

### Content Density Thresholds

In [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) (lines 31-55), the `render_pages_for_ocr` function evaluates each `Page` against three criteria:

- Native text length is less than 20 bytes
- Text coverage is below 15% of the page area
- The page contains images without sufficient accompanying text

Pages meeting these conditions are flagged for rasterization into `RenderedPage` objects, while text-rich pages bypass OCR entirely.

### Garbled Text Detection

Before merging, LiteParse detects substitution-cipher-like text that indicates encoding failures. The `page_is_garbled` helper (lines 71-84) uses `is_likely_garbled` heuristics to identify pages where native text is corrupted. When detected, the parser clears existing items via `page.text_items.clear()`, allowing OCR results to replace rather than supplement the garbled content.

## Concurrent OCR Processing

### Bounded Parallelism with Tokio

The `ocr_and_merge_rendered` function (lines 91-118) manages OCR execution using a Tokio semaphore to limit concurrent operations based on `num_workers`. Each worker task calls `engine.recognize` on the bitmap, supporting both local Tesseract ([`ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/ocr/tesseract.rs)) and remote HTTP OCR services ([`ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/ocr/http_simple.rs)).

### Coordinate Transformation

OCR engines return pixel-based coordinates that must align with PDF point coordinates. The implementation applies a `scale_factor = 72.0 / dpi` transformation when constructing new `TextItem` structures, ensuring spatial consistency between native and OCR-derived content.

## Deduplication and Merge Logic

The merge process prevents visual duplicates while preserving legitimate multi-layer content.

### Geometric Overlap Detection

To avoid inserting OCR text that duplicates existing native content, the code checks `overlaps_existing_text` (lines 185-207). This function compares OCR result bounding boxes against native `TextItem` geometries using a 2-point tolerance threshold. Overlapping OCR lines are discarded to prevent redundancy.

### Artifact Cleaning

OCR outputs often misrecognize table borders as characters like `|` or `[`. The `clean_ocr_table_artifacts` function (lines 361-390) strips these common misclassifications before insertion, ensuring only meaningful text enters the final output.

### Unified Structure Insertion

Clean OCR results are inserted as new `TextItem` instances (lines 216-226) with:

- Transformed geometric coordinates
- Cleaned text content
- Synthetic `"OCR"` font name
- Confidence scores from the OCR engine

## Error Handling Strategies

The implementation distinguishes between partial and complete OCR failures. If individual pages fail OCR, they are logged and skipped. However, if all OCR tasks fail on pages that specifically required OCR (sparse-text pages), the function returns `LiteParseError::Ocr` (lines 230-246), signaling an incomplete parse to the caller rather than silently missing critical content.

## Configuration and Usage Example

To combine OCR results with native PDF text in your application, configure the `LiteParse` instance as follows:

```rust
use liteparse::LiteParse;
use liteparse::LiteParseConfig;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let config = LiteParseConfig {
        ocr_enabled: true,
        ocr_language: "eng".into(),
        ocr_workers: 4,
        // ocr_server_url: Some("http://localhost:8080/ocr".into()),
        ..Default::default()
    };

    let mut parser = LiteParse::new(config);
    let result = parser.parse_file("sample.pdf").await?;

    for page in result.pages {
        for item in page.text_items {
            println!("{} (confidence: {})", 
                item.text, 
                item.confidence.unwrap_or(0.0)
            );
        }
    }
    Ok(())
}

```

This configuration enables the full hybrid pipeline, automatically combining native extraction with OCR where needed.

## Summary

- LiteParse evaluates pages using heuristics (20-byte minimum, 15% coverage) to determine OCR necessity in [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs)
- The system detects and clears garbled native text using `is_likely_garbled` before merging to prevent corruption
- A Tokio-based worker pool bounded by `ocr_workers` executes OCR concurrently
- Geometric overlap detection with 2-point tolerance prevents duplicate text insertion
- The `scale_factor = 72/dpi` transformation aligns OCR pixel coordinates with PDF points
- `clean_ocr_table_artifacts` removes common OCR misrecognitions like border characters
- Complete OCR failure on text-poor pages raises `LiteParseError::Ocr` rather than failing silently

## Frequently Asked Questions

### How does LiteParse decide which pages need OCR?

LiteParse calculates content density in `render_pages_for_ocr` by checking if native text length falls below 20 bytes or coverage drops under 15% of page area. Pages containing images without sufficient text also trigger OCR rasterization according to the logic in lines 31-55 of [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs).

### What prevents duplicate text when OCR and native extraction overlap?

The `overlaps_existing_text` function (lines 185-207) implements geometric collision detection with a 2-point tolerance threshold. OCR results whose bounding boxes intersect with existing native `TextItem`s are discarded to avoid visual duplication in the final output.

### Can LiteParse use remote OCR services instead of local Tesseract?

Yes. The `OcrEngine` trait abstraction supports both local Tesseract via [`ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/ocr/tesseract.rs) and remote HTTP endpoints via [`ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/ocr/http_simple.rs). Configure `ocr_server_url` in `LiteParseConfig` to route OCR requests to external services while maintaining the same merge logic.

### How does LiteParse handle OCR failures?

The system differentiates between partial and catastrophic failures. Individual page failures are logged and skipped, but if all OCR tasks fail on pages that required OCR (text-poor pages), the parser returns `LiteParseError::Ocr` (lines 230-246), ensuring the caller recognizes incomplete extraction rather than receiving silently truncated results.