How LiteParse Merges OCR Results With Native PDF Text While Preserving Confidence Scores

LiteParse renders text-poor PDF pages to bitmaps, runs OCR in parallel, and merges the recognized words into the native text_items list while preserving each OCR fragment's confidence score in the TextItem.confidence field.

The run-llama/liteparse crate extracts native text from PDF pages using PDFium, then intelligently backfills missing or garbled content with OCR results from Tesseract or an HTTP service. By storing OCR confidence as an optional f32 on every TextItem, downstream pipelines can distinguish native PDF text from OCR-supplemented text and weight it accordingly.

Detecting Pages That Need OCR

The render_pages_for_ocr function in crates/liteparse/src/ocr_merge.rs evaluates every parsed Page and decides whether OCR is required. A page is flagged for OCR if any of four heuristics pass: the native text_length is less than 20, the text_coverage ratio is below 0.15, the page contains images, or the helper page_is_garbled(page) detects substitution-cipher corruption via is_likely_garbled.

let needs_ocr = text_length < 20 || text_coverage < 0.15 || has_images || page_is_garbled(page);

Source: ocr_merge.rs lines 54–56

When a page meets any of these conditions, LiteParse renders it to an RGB bitmap through PDFium and stores the resulting bytes in a RenderedPage struct for subsequent processing.

Running OCR Concurrently

ocr_and_merge_rendered in ocr_merge.rs receives the rendered bitmaps and limits parallelism using a semaphore keyed to num_workers. Each spawned task invokes the OCR engine’s recognize method against the RGB bytes and returns a Vec<OcrResult>.

let engine = ocr_engine.clone();
...
rt_handle.block_on(engine.recognize(&r.rgb_bytes, r.width, r.height, &options))

Source: ocr_merge.rs lines 99–116

If every OCR task fails and at least one of the selected pages was text-poor, the orchestrator surfaces a LiteParseError::Ocr rather than silently discarding the data. This prevents pipelines from receiving empty pages when OCR was the only viable text source.

Filtering and Cleaning OCR Results

Before any OCR result is inserted, the merger applies two hard filters. First, fragments with r.confidence less than or equal to 0.1 are ignored. Second, the candidate box must not overlap existing native text that was already accepted, as checked by overlaps_existing_text with a tolerance factor of 2.0.

if r.confidence <= 0.1 { continue; }
...
if overlaps_existing_text(&page.text_items[..native_count], ocr_x, ocr_y, ocr_w, ocr_h, 2.0) {
    continue;
}

Sources: ocr_merge.rs lines 191–200 and 85–94

A utility named clean_ocr_table_artifacts strips common table-border characters such as "|", "[", and "]" when the surrounding text appears numeric, preventing visual artifacts from polluting the final output.

Storing Confidence Scores in TextItem

When an OCR result survives the filters, LiteParse appends a new TextItem to the page’s text_items vector. The OCR bounding box is converted from pixel coordinates to PDF coordinates using scale_factor = 72.0 / dpi. The confidence value is normalized and rounded to three decimal places.

page.text_items.push(TextItem {
    text: cleaned,
    x: ocr_x,
    y: ocr_y,
    width: ocr_w,
    height: ocr_h,
    font_name: Some("OCR".to_string()),
    font_size: Some(ocr_h),
    confidence: Some((r.confidence * 1000.0).round() / 1000.0),
    ..Default::default()
});

Source: ocr_merge.rs lines 216–226

The TextItem definition in crates/liteparse/src/types.rs exposes an optional confidence field that is None for native PDF text and Some(f32) for OCR-derived entries.

/// OCR confidence score (0.0–1.0). None for native PDF text.
pub confidence: Option<f32>,

Source: types.rs lines 53–55

This design makes it trivial for consumers to differentiate sources: native items lack the field, while OCR items carry the normalized score.

End-to-End Merge Flow

The high-level parsing flow in crates/liteparse/src/parser.rs coordinates the entire operation. After the initial native parse produces Page objects, it calls render_pages_for_ocr to collect bitmaps and then ocr_and_merge_rendered to inject the filtered OCR results back into the same pages.

let (mut pages, ocr_rendered) = if self.config.ocr_enabled {
    let rendered = ocr_merge::render_pages_for_ocr(&document, &pages, self.config.dpi)?;
    // ... later:
    ocr_merge::ocr_and_merge_rendered(&mut pages, ocr_rendered, …).await?;
};

Sources: parser.rs lines 102–115 and 172–176

The end result is a unified text_items list per page that blends native and OCR text, with each OCR entry retaining its original quality score.

Practical Rust Example

The following snippet enables OCR, runs the parser, and inspects the merged confidence metadata:

use liteparse::LiteParse;
use std::sync::Arc;
use liteparse::ocr::tesseract::TesseractOcrEngine;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let ocr_engine = Arc::new(TesseractOcrEngine::new()?);
    let parser = LiteParse::new()
        .with_ocr_engine(ocr_engine)
        .ocr_enabled(true)
        .ocr_language("eng".into());

    let result = parser.parse_path("sample.pdf").await?;
    for page in result.pages {
        for item in page.text_items {
            match item.confidence {
                Some(conf) => println!("OCR: '{}' (conf: {:.2})", item.text, conf),
                None => println!("Native: '{}'", item.text),
            }
        }
    }
    Ok(())
}

Key source files governing this behavior include:

Summary

  • Page detection in ocr_merge.rs uses length, coverage, image presence, and garble checks to decide when OCR is needed.
  • Concurrent execution runs each page through engine.recognize under a semaphore, failing loudly with LiteParseError::Ocr if all tasks fail on text-poor pages.
  • Quality filtering drops OCR results below 0.1 confidence and rejects boxes that overlap native text.
  • Confidence preservation stores the OCR score as a rounded f32 in TextItem.confidence, while native text remains None.
  • Unified output means downstream consumers receive a single text_items list with transparent source attribution.

Frequently Asked Questions

How does LiteParse decide which pages need OCR?

LiteParse evaluates four heuristics in render_pages_for_ocr: native text length under 20 characters, text coverage below 15 percent, presence of images, and detection of garbled text via is_likely_garbled. If any check passes, the page is rendered to a bitmap for OCR processing.

What happens to OCR confidence scores during the merge?

Each surviving OCR result passes its confidence value into a new TextItem in the text_items vector, rounded to three decimal places in the range 0.0 to 1.0. Native PDF text items have confidence: None, creating a clear distinction between sources.

How does LiteParse prevent OCR text from duplicating native text?

Before insertion, every OCR bounding box is checked with overlaps_existing_text against the existing native items. If the candidate overlaps native text within a tolerance factor of 2.0, it is skipped entirely, ensuring no redundant duplicates appear in the final list.

Can LiteParse use OCR engines other than Tesseract?

Yes. While the built-in option is TesseractOcrEngine in crates/liteparse/src/ocr/tesseract.rs, the library also supplies an HTTP-based engine in ocr/http_simple.rs. Any engine implementing the OcrEngine trait can be passed to LiteParse::with_ocr_engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →