How the OCR Merge Algorithm Works in LiteParse (`ocr_merge.rs`)
The OCR merge algorithm in LiteParse implements a two-stage pipeline that analyzes page complexity to determine which pages need OCR, renders those pages to bitmap, executes OCR in parallel with semaphore-controlled concurrency, and merges the results back into the PDF text model while deduplicating against existing native text.
The ocr_merge.rs file in the run-llama/liteparse repository contains the core logic for bridging OCR engines with PDF text extraction. This Rust module handles the complete lifecycle from deciding which pages require optical character recognition to integrating the recognized text into the document's TextItem structure.
Deciding Which Pages Need OCR
Before any bitmap rendering occurs, the algorithm evaluates each page using calculate_page_complexity to determine if OCR is necessary.
The calculate_page_complexity Heuristics
Located at lines 15–33 in crates/liteparse/src/ocr_merge.rs, this function applies five distinct heuristics:
- Scanned / No-text pages: Fewer than 20 characters and a full-page image triggers
ComplexityReason::Scannedor::NoText. - Sparse text: Under 2000 characters with less than 15% coverage results in
ComplexityReason::SparseText. - Embedded images: Any raster image exceeding
MIN_IMAGE_SIZE_PTtriggersComplexityReason::EmbeddedImages. - Garbled native text: Pages failing the
page_is_garbledvowel-ratio test receiveComplexityReason::Garbled. - Vector-only text: When filled-path area not covered by native text exceeds
UNCOVERED_VECTOR_AREA_THRESHOLD, the reason becomesComplexityReason::VectorText.
If any heuristic triggers, needs_ocr is set to true (lines 200–202).
Rendering and Parallel OCR Execution
Rendering Pages for OCR
The render_pages_for_ocr function (lines 46–52) iterates over all pages, re-evaluates complexity, and for pages with needs_ocr set to true, calls pdfium::Page::render(dpi) at lines 55–60. This produces a RenderedPage struct (lines 26–32) containing raw RGB bytes, width, and height.
Parallel Execution with Concurrency Control
ocr_and_merge_rendered orchestrates the OCR workload using async concurrency:
- Task spawning: One async task per rendered page (lines 100–108).
- Semaphore throttling: A semaphore limits concurrent workers to
num_workers(lines 81–86), preventing deadlocks when OCR engines use blocking I/O. - Blocking offloading: Each task acquires a permit via
sem.acquire_owned().awaitand moves the actual OCR call tospawn_blocking(lines 111–119), ensuring the semaphore throttles CPU-blocking work correctly.
The OCR engine returns a Vec<OcrResult> for each processed page.
Error Handling and Systemic Failure Detection
The algorithm implements a guard against total OCR failure. After all tasks complete, it counts failed_tasks and checks if any failure occurred on a sparse-text page (failed_sparse_text_page).
If every OCR task fails and at least one failure occurred on a sparse-text page, the function returns LiteParseError::Ocr (lines 186–194). This prevents silent data loss when OCR is the primary text source but allows partial failures when native text exists.
Merging OCR Results into the Text Model
For each successful page, the merging process follows a strict protocol:
Cleaning Native Text Items
The algorithm first clears unusable native text. If the page is garbled, all native text is removed; otherwise, only items flagged by is_unusable_native are dropped (lines 84–96).
Filtering and Processing Results
- Confidence filtering: OCR results with
confidence ≤ 0.1are discarded (lines 103–106). - Bounding box calculation: When the OCR result includes a polygon, the code extracts the axis-aligned bbox and derives rotation via
polygon_rotation_deg(lines 113–129). - Overlap detection: The OCR bbox is compared against only native text items using
overlaps_existing_text(lines 139–145). If overlap exists, the OCR result is ignored to prevent duplication. - Artifact cleaning:
clean_ocr_table_artifactsstrips spurious characters like|,[, and]from numeric-like OCR strings (lines 150–166). - Font size estimation: For rotated text, the narrower dimension serves as the font-size heuristic (lines 167–174).
Creating TextItem Entries
Finally, the algorithm appends a TextItem to the page with fields: text, x, y, width, height, rotation, font_name = "OCR", font_size, and confidence (lines 176–186).
Key Helper Utilities
The algorithm relies on several specialized utility functions:
polygon_rotation_deg(lines 260–320): Deduces text rotation (0°, 90°, 180°, 270°) from 4-point OCR polygons, handling both reading-direction and screen-axis orderings.overlaps_existing_text(lines 330–350): Checks bbox overlap with configurable tolerance.clean_ocr_table_artifacts(lines 340–380): Removes typical OCR misreads of table borders while preserving non-numeric content.is_unusable_nativeandpage_is_garbled(lines 380–420): Detect corrupted native text using Unicode-map errors and vowel-ratio heuristics.
Practical Implementation Examples
Triggering OCR from Application Code
use liteparse::parser::LiteParse;
use std::sync::Arc;
// Assume `doc` is an opened pdfium::Document and `pages` is a Vec<Page>.
let dpi = 300.0;
// 1. Render pages that need OCR based on complexity heuristics.
let rendered = liteparse::ocr_merge::render_pages_for_ocr(&doc, &pages, dpi)?;
// 2. Initialize an OCR engine (Tesseract example).
let ocr_engine: Arc<dyn liteparse::ocr::OcrEngine> =
Arc::new(liteparse::ocr::tesseract::TesseractEngine::new()?);
// 3. Execute OCR and merge results back into the pages vector.
liteparse::ocr_merge::ocr_and_merge_rendered(
&mut pages,
rendered,
dpi,
ocr_engine,
"eng", // language code
4, // parallel workers (semaphore limit)
).await?;
Inspecting Merged Results
for page in pages {
println!("--- Page {} ---", page.page_number);
for item in &page.text_items {
println!(
"{} ({:.1},{:.1}) [{}×{}] rot={:.0}° conf={:?}",
item.text,
item.x,
item.y,
item.width,
item.height,
item.rotation,
item.confidence
);
}
}
OCR-derived items will display font_name = "OCR" and include confidence scores from the recognition engine.
Summary
- Complexity heuristics determine OCR necessity using character counts, coverage percentages, and image detection in
calculate_page_complexity. - Parallel execution uses semaphore-controlled async tasks to prevent resource exhaustion while maximizing throughput.
- Failure detection aborts parsing only when OCR is the sole text source and completely fails.
- Deduplication logic compares OCR bounding boxes against native text items to prevent overlapping duplicates.
- Post-processing includes rotation detection, table artifact cleaning, and font size estimation before creating
TextItementries.
Frequently Asked Questions
How does LiteParse decide which pages need OCR?
The algorithm evaluates five heuristics in calculate_page_complexity (lines 15–33): scanned/no-text pages, sparse text coverage, embedded images, garbled native text detected via vowel ratios, and vector-only content. If any condition triggers, needs_ocr becomes true and the page is rendered for processing.
What happens if OCR fails on all pages?
The ocr_and_merge_rendered function implements a systemic-failure guard (lines 186–194). If every OCR task fails and at least one failure occurs on a sparse-text page (where OCR is the primary text source), the function returns LiteParseError::Ocr. This prevents silent failures when OCR is critical while allowing partial recognition to continue.
How does the algorithm prevent duplicate text from OCR and native PDF text?
During merging, each OCR result's bounding box is checked against only native text items using overlaps_existing_text (lines 139–145). If the OCR text overlaps with existing native text, the OCR result is discarded. This ensures the final document contains either the native text or the OCR text, never both.
Why does the algorithm use a semaphore for OCR workers?
The semaphore (lines 81–86) prevents deadlocks and resource exhaustion when OCR engines use blocking I/O. By limiting concurrent workers to num_workers and acquiring permits via sem.acquire_owned().await before calling spawn_blocking (lines 111–119), the algorithm ensures that CPU-intensive OCR work does not overwhelm the async runtime while maintaining parallel throughput.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →