How LiteParse Combines OCR Results with Native PDF Text: A Technical Deep Dive

LiteParse merges OCR results with native PDF text by first extracting native content, then conditionally rasterizing text-poor pages for OCR, and finally deduplicating overlapping text items using geometric overlap detection before unifying them into a single Page structure.

The run-llama/liteparse repository implements a sophisticated hybrid parsing strategy that combines native PDF text extraction with optical character recognition (OCR) to handle scanned documents and text-poor layouts. When processing PDFs, the parser seamlessly integrates machine-readable text with OCR-derived content while avoiding duplicates and preserving spatial accuracy. This article examines the specific implementation details found in the Rust source code, focusing on how the library combines OCR results with native PDF text to produce unified output.

The Two-Stage Hybrid Extraction Pipeline

The LiteParse engine operates in distinct phases orchestrated in parser.rs. First, it extracts native text directly from the PDF structure. Then, if OCR is enabled via LiteParseConfig, it invokes the merge logic from ocr_merge.rs to supplement missing content.

The pipeline follows this sequence:

  1. Extract native TextItems from the PDF
  2. Identify pages requiring OCR based on content density
  3. Rasterize selected pages to RGB bitmaps
  4. Execute OCR concurrently across worker threads
  5. Merge valid OCR results back into the native Page structures

Detecting When Pages Require OCR

The system uses heuristic rules to avoid unnecessary OCR processing on text-rich digital documents.

Content Density Thresholds

In ocr_merge.rs (lines 31-55), the render_pages_for_ocr function evaluates each Page against three criteria:

  • Native text length is less than 20 bytes
  • Text coverage is below 15% of the page area
  • The page contains images without sufficient accompanying text

Pages meeting these conditions are flagged for rasterization into RenderedPage objects, while text-rich pages bypass OCR entirely.

Garbled Text Detection

Before merging, LiteParse detects substitution-cipher-like text that indicates encoding failures. The page_is_garbled helper (lines 71-84) uses is_likely_garbled heuristics to identify pages where native text is corrupted. When detected, the parser clears existing items via page.text_items.clear(), allowing OCR results to replace rather than supplement the garbled content.

Concurrent OCR Processing

Bounded Parallelism with Tokio

The ocr_and_merge_rendered function (lines 91-118) manages OCR execution using a Tokio semaphore to limit concurrent operations based on num_workers. Each worker task calls engine.recognize on the bitmap, supporting both local Tesseract (ocr/tesseract.rs) and remote HTTP OCR services (ocr/http_simple.rs).

Coordinate Transformation

OCR engines return pixel-based coordinates that must align with PDF point coordinates. The implementation applies a scale_factor = 72.0 / dpi transformation when constructing new TextItem structures, ensuring spatial consistency between native and OCR-derived content.

Deduplication and Merge Logic

The merge process prevents visual duplicates while preserving legitimate multi-layer content.

Geometric Overlap Detection

To avoid inserting OCR text that duplicates existing native content, the code checks overlaps_existing_text (lines 185-207). This function compares OCR result bounding boxes against native TextItem geometries using a 2-point tolerance threshold. Overlapping OCR lines are discarded to prevent redundancy.

Artifact Cleaning

OCR outputs often misrecognize table borders as characters like | or [. The clean_ocr_table_artifacts function (lines 361-390) strips these common misclassifications before insertion, ensuring only meaningful text enters the final output.

Unified Structure Insertion

Clean OCR results are inserted as new TextItem instances (lines 216-226) with:

  • Transformed geometric coordinates
  • Cleaned text content
  • Synthetic "OCR" font name
  • Confidence scores from the OCR engine

Error Handling Strategies

The implementation distinguishes between partial and complete OCR failures. If individual pages fail OCR, they are logged and skipped. However, if all OCR tasks fail on pages that specifically required OCR (sparse-text pages), the function returns LiteParseError::Ocr (lines 230-246), signaling an incomplete parse to the caller rather than silently missing critical content.

Configuration and Usage Example

To combine OCR results with native PDF text in your application, configure the LiteParse instance as follows:

use liteparse::LiteParse;
use liteparse::LiteParseConfig;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let config = LiteParseConfig {
        ocr_enabled: true,
        ocr_language: "eng".into(),
        ocr_workers: 4,
        // ocr_server_url: Some("http://localhost:8080/ocr".into()),
        ..Default::default()
    };

    let mut parser = LiteParse::new(config);
    let result = parser.parse_file("sample.pdf").await?;

    for page in result.pages {
        for item in page.text_items {
            println!("{} (confidence: {})", 
                item.text, 
                item.confidence.unwrap_or(0.0)
            );
        }
    }
    Ok(())
}

This configuration enables the full hybrid pipeline, automatically combining native extraction with OCR where needed.

Summary

  • LiteParse evaluates pages using heuristics (20-byte minimum, 15% coverage) to determine OCR necessity in ocr_merge.rs
  • The system detects and clears garbled native text using is_likely_garbled before merging to prevent corruption
  • A Tokio-based worker pool bounded by ocr_workers executes OCR concurrently
  • Geometric overlap detection with 2-point tolerance prevents duplicate text insertion
  • The scale_factor = 72/dpi transformation aligns OCR pixel coordinates with PDF points
  • clean_ocr_table_artifacts removes common OCR misrecognitions like border characters
  • Complete OCR failure on text-poor pages raises LiteParseError::Ocr rather than failing silently

Frequently Asked Questions

How does LiteParse decide which pages need OCR?

LiteParse calculates content density in render_pages_for_ocr by checking if native text length falls below 20 bytes or coverage drops under 15% of page area. Pages containing images without sufficient text also trigger OCR rasterization according to the logic in lines 31-55 of ocr_merge.rs.

What prevents duplicate text when OCR and native extraction overlap?

The overlaps_existing_text function (lines 185-207) implements geometric collision detection with a 2-point tolerance threshold. OCR results whose bounding boxes intersect with existing native TextItems are discarded to avoid visual duplication in the final output.

Can LiteParse use remote OCR services instead of local Tesseract?

Yes. The OcrEngine trait abstraction supports both local Tesseract via ocr/tesseract.rs and remote HTTP endpoints via ocr/http_simple.rs. Configure ocr_server_url in LiteParseConfig to route OCR requests to external services while maintaining the same merge logic.

How does LiteParse handle OCR failures?

The system differentiates between partial and catastrophic failures. Individual page failures are logged and skipped, but if all OCR tasks fail on pages that required OCR (text-poor pages), the parser returns LiteParseError::Ocr (lines 230-246), ensuring the caller recognizes incomplete extraction rather than receiving silently truncated results.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →