# How LiteParse Merges OCR Results With Native PDF Text While Preserving Confidence Scores

> Discover how LiteParse merges OCR text with native PDF content, crucially preserving confidence scores. Enhance your document processing with this advanced technique.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-06

---

**LiteParse renders text-poor PDF pages to bitmaps, runs OCR in parallel, and merges the recognized words into the native `text_items` list while preserving each OCR fragment's confidence score in the `TextItem.confidence` field.**

The `run-llama/liteparse` crate extracts native text from PDF pages using PDFium, then intelligently backfills missing or garbled content with OCR results from Tesseract or an HTTP service. By storing OCR confidence as an optional `f32` on every `TextItem`, downstream pipelines can distinguish native PDF text from OCR-supplemented text and weight it accordingly.

## Detecting Pages That Need OCR

The `render_pages_for_ocr` function in [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) evaluates every parsed `Page` and decides whether OCR is required. A page is flagged for OCR if any of four heuristics pass: the native `text_length` is less than `20`, the `text_coverage` ratio is below `0.15`, the page contains images, or the helper `page_is_garbled(page)` detects substitution-cipher corruption via `is_likely_garbled`.

```rust
let needs_ocr = text_length < 20 || text_coverage < 0.15 || has_images || page_is_garbled(page);

```

*Source: [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) lines 54–56*

When a page meets any of these conditions, LiteParse renders it to an RGB bitmap through PDFium and stores the resulting bytes in a `RenderedPage` struct for subsequent processing.

## Running OCR Concurrently

`ocr_and_merge_rendered` in [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) receives the rendered bitmaps and limits parallelism using a semaphore keyed to `num_workers`. Each spawned task invokes the OCR engine’s `recognize` method against the RGB bytes and returns a `Vec<OcrResult>`.

```rust
let engine = ocr_engine.clone();
...
rt_handle.block_on(engine.recognize(&r.rgb_bytes, r.width, r.height, &options))

```

*Source: [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) lines 99–116*

If every OCR task fails and at least one of the selected pages was text-poor, the orchestrator surfaces a `LiteParseError::Ocr` rather than silently discarding the data. This prevents pipelines from receiving empty pages when OCR was the only viable text source.

## Filtering and Cleaning OCR Results

Before any OCR result is inserted, the merger applies two hard filters. First, fragments with `r.confidence` less than or equal to `0.1` are ignored. Second, the candidate box must not overlap existing native text that was already accepted, as checked by `overlaps_existing_text` with a tolerance factor of `2.0`.

```rust
if r.confidence <= 0.1 { continue; }
...
if overlaps_existing_text(&page.text_items[..native_count], ocr_x, ocr_y, ocr_w, ocr_h, 2.0) {
    continue;
}

```

*Sources: [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) lines 191–200 and 85–94*

A utility named `clean_ocr_table_artifacts` strips common table-border characters such as `"|"`, `"["`, and `"]"` when the surrounding text appears numeric, preventing visual artifacts from polluting the final output.

## Storing Confidence Scores in TextItem

When an OCR result survives the filters, LiteParse appends a new `TextItem` to the page’s `text_items` vector. The OCR bounding box is converted from pixel coordinates to PDF coordinates using `scale_factor = 72.0 / dpi`. The confidence value is normalized and rounded to three decimal places.

```rust
page.text_items.push(TextItem {
    text: cleaned,
    x: ocr_x,
    y: ocr_y,
    width: ocr_w,
    height: ocr_h,
    font_name: Some("OCR".to_string()),
    font_size: Some(ocr_h),
    confidence: Some((r.confidence * 1000.0).round() / 1000.0),
    ..Default::default()
});

```

*Source: [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) lines 216–226*

The `TextItem` definition in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) exposes an optional `confidence` field that is `None` for native PDF text and `Some(f32)` for OCR-derived entries.

```rust
/// OCR confidence score (0.0–1.0). None for native PDF text.
pub confidence: Option<f32>,

```

*Source: [`types.rs`](https://github.com/run-llama/liteparse/blob/main/types.rs) lines 53–55*

This design makes it trivial for consumers to differentiate sources: native items lack the field, while OCR items carry the normalized score.

## End-to-End Merge Flow

The high-level parsing flow in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) coordinates the entire operation. After the initial native parse produces `Page` objects, it calls `render_pages_for_ocr` to collect bitmaps and then `ocr_and_merge_rendered` to inject the filtered OCR results back into the same pages.

```rust
let (mut pages, ocr_rendered) = if self.config.ocr_enabled {
    let rendered = ocr_merge::render_pages_for_ocr(&document, &pages, self.config.dpi)?;
    // ... later:
    ocr_merge::ocr_and_merge_rendered(&mut pages, ocr_rendered, …).await?;
};

```

*Sources: [`parser.rs`](https://github.com/run-llama/liteparse/blob/main/parser.rs) lines 102–115 and 172–176*

The end result is a unified `text_items` list per page that blends native and OCR text, with each OCR entry retaining its original quality score.

## Practical Rust Example

The following snippet enables OCR, runs the parser, and inspects the merged confidence metadata:

```rust
use liteparse::LiteParse;
use std::sync::Arc;
use liteparse::ocr::tesseract::TesseractOcrEngine;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let ocr_engine = Arc::new(TesseractOcrEngine::new()?);
    let parser = LiteParse::new()
        .with_ocr_engine(ocr_engine)
        .ocr_enabled(true)
        .ocr_language("eng".into());

    let result = parser.parse_path("sample.pdf").await?;
    for page in result.pages {
        for item in page.text_items {
            match item.confidence {
                Some(conf) => println!("OCR: '{}' (conf: {:.2})", item.text, conf),
                None => println!("Native: '{}'", item.text),
            }
        }
    }
    Ok(())
}

```

Key source files governing this behavior include:

- [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) — rendering, filtering, and merging logic
- [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) — `TextItem` definition with optional `confidence`
- [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) — orchestration of the OCR pipeline
- [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs) — built-in Tesseract engine
- [`crates/liteparse/src/ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/http_simple.rs) — optional HTTP OCR engine

## Summary

- **Page detection** in [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) uses length, coverage, image presence, and garble checks to decide when OCR is needed.
- **Concurrent execution** runs each page through `engine.recognize` under a semaphore, failing loudly with `LiteParseError::Ocr` if all tasks fail on text-poor pages.
- **Quality filtering** drops OCR results below `0.1` confidence and rejects boxes that overlap native text.
- **Confidence preservation** stores the OCR score as a rounded `f32` in `TextItem.confidence`, while native text remains `None`.
- **Unified output** means downstream consumers receive a single `text_items` list with transparent source attribution.

## Frequently Asked Questions

### How does LiteParse decide which pages need OCR?

LiteParse evaluates four heuristics in `render_pages_for_ocr`: native text length under `20` characters, text coverage below `15` percent, presence of images, and detection of garbled text via `is_likely_garbled`. If any check passes, the page is rendered to a bitmap for OCR processing.

### What happens to OCR confidence scores during the merge?

Each surviving OCR result passes its confidence value into a new `TextItem` in the `text_items` vector, rounded to three decimal places in the range `0.0` to `1.0`. Native PDF text items have `confidence: None`, creating a clear distinction between sources.

### How does LiteParse prevent OCR text from duplicating native text?

Before insertion, every OCR bounding box is checked with `overlaps_existing_text` against the existing native items. If the candidate overlaps native text within a tolerance factor of `2.0`, it is skipped entirely, ensuring no redundant duplicates appear in the final list.

### Can LiteParse use OCR engines other than Tesseract?

Yes. While the built-in option is `TesseractOcrEngine` in [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs), the library also supplies an HTTP-based engine in [`ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/ocr/http_simple.rs). Any engine implementing the `OcrEngine` trait can be passed to `LiteParse::with_ocr_engine`.