How LiteParse Combines OCR Results with Native PDF Text: A Technical Deep Dive
LiteParse merges OCR results with native PDF text by first extracting native content, then conditionally rasterizing text-poor pages for OCR, and finally deduplicating overlapping text items using geometric overlap detection before unifying them into a single Page structure.
The run-llama/liteparse repository implements a sophisticated hybrid parsing strategy that combines native PDF text extraction with optical character recognition (OCR) to handle scanned documents and text-poor layouts. When processing PDFs, the parser seamlessly integrates machine-readable text with OCR-derived content while avoiding duplicates and preserving spatial accuracy. This article examines the specific implementation details found in the Rust source code, focusing on how the library combines OCR results with native PDF text to produce unified output.
The Two-Stage Hybrid Extraction Pipeline
The LiteParse engine operates in distinct phases orchestrated in parser.rs. First, it extracts native text directly from the PDF structure. Then, if OCR is enabled via LiteParseConfig, it invokes the merge logic from ocr_merge.rs to supplement missing content.
The pipeline follows this sequence:
- Extract native
TextItems from the PDF - Identify pages requiring OCR based on content density
- Rasterize selected pages to RGB bitmaps
- Execute OCR concurrently across worker threads
- Merge valid OCR results back into the native
Pagestructures
Detecting When Pages Require OCR
The system uses heuristic rules to avoid unnecessary OCR processing on text-rich digital documents.
Content Density Thresholds
In ocr_merge.rs (lines 31-55), the render_pages_for_ocr function evaluates each Page against three criteria:
- Native text length is less than 20 bytes
- Text coverage is below 15% of the page area
- The page contains images without sufficient accompanying text
Pages meeting these conditions are flagged for rasterization into RenderedPage objects, while text-rich pages bypass OCR entirely.
Garbled Text Detection
Before merging, LiteParse detects substitution-cipher-like text that indicates encoding failures. The page_is_garbled helper (lines 71-84) uses is_likely_garbled heuristics to identify pages where native text is corrupted. When detected, the parser clears existing items via page.text_items.clear(), allowing OCR results to replace rather than supplement the garbled content.
Concurrent OCR Processing
Bounded Parallelism with Tokio
The ocr_and_merge_rendered function (lines 91-118) manages OCR execution using a Tokio semaphore to limit concurrent operations based on num_workers. Each worker task calls engine.recognize on the bitmap, supporting both local Tesseract (ocr/tesseract.rs) and remote HTTP OCR services (ocr/http_simple.rs).
Coordinate Transformation
OCR engines return pixel-based coordinates that must align with PDF point coordinates. The implementation applies a scale_factor = 72.0 / dpi transformation when constructing new TextItem structures, ensuring spatial consistency between native and OCR-derived content.
Deduplication and Merge Logic
The merge process prevents visual duplicates while preserving legitimate multi-layer content.
Geometric Overlap Detection
To avoid inserting OCR text that duplicates existing native content, the code checks overlaps_existing_text (lines 185-207). This function compares OCR result bounding boxes against native TextItem geometries using a 2-point tolerance threshold. Overlapping OCR lines are discarded to prevent redundancy.
Artifact Cleaning
OCR outputs often misrecognize table borders as characters like | or [. The clean_ocr_table_artifacts function (lines 361-390) strips these common misclassifications before insertion, ensuring only meaningful text enters the final output.
Unified Structure Insertion
Clean OCR results are inserted as new TextItem instances (lines 216-226) with:
- Transformed geometric coordinates
- Cleaned text content
- Synthetic
"OCR"font name - Confidence scores from the OCR engine
Error Handling Strategies
The implementation distinguishes between partial and complete OCR failures. If individual pages fail OCR, they are logged and skipped. However, if all OCR tasks fail on pages that specifically required OCR (sparse-text pages), the function returns LiteParseError::Ocr (lines 230-246), signaling an incomplete parse to the caller rather than silently missing critical content.
Configuration and Usage Example
To combine OCR results with native PDF text in your application, configure the LiteParse instance as follows:
use liteparse::LiteParse;
use liteparse::LiteParseConfig;
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let config = LiteParseConfig {
ocr_enabled: true,
ocr_language: "eng".into(),
ocr_workers: 4,
// ocr_server_url: Some("http://localhost:8080/ocr".into()),
..Default::default()
};
let mut parser = LiteParse::new(config);
let result = parser.parse_file("sample.pdf").await?;
for page in result.pages {
for item in page.text_items {
println!("{} (confidence: {})",
item.text,
item.confidence.unwrap_or(0.0)
);
}
}
Ok(())
}
This configuration enables the full hybrid pipeline, automatically combining native extraction with OCR where needed.
Summary
- LiteParse evaluates pages using heuristics (20-byte minimum, 15% coverage) to determine OCR necessity in
ocr_merge.rs - The system detects and clears garbled native text using
is_likely_garbledbefore merging to prevent corruption - A Tokio-based worker pool bounded by
ocr_workersexecutes OCR concurrently - Geometric overlap detection with 2-point tolerance prevents duplicate text insertion
- The
scale_factor = 72/dpitransformation aligns OCR pixel coordinates with PDF points clean_ocr_table_artifactsremoves common OCR misrecognitions like border characters- Complete OCR failure on text-poor pages raises
LiteParseError::Ocrrather than failing silently
Frequently Asked Questions
How does LiteParse decide which pages need OCR?
LiteParse calculates content density in render_pages_for_ocr by checking if native text length falls below 20 bytes or coverage drops under 15% of page area. Pages containing images without sufficient text also trigger OCR rasterization according to the logic in lines 31-55 of ocr_merge.rs.
What prevents duplicate text when OCR and native extraction overlap?
The overlaps_existing_text function (lines 185-207) implements geometric collision detection with a 2-point tolerance threshold. OCR results whose bounding boxes intersect with existing native TextItems are discarded to avoid visual duplication in the final output.
Can LiteParse use remote OCR services instead of local Tesseract?
Yes. The OcrEngine trait abstraction supports both local Tesseract via ocr/tesseract.rs and remote HTTP endpoints via ocr/http_simple.rs. Configure ocr_server_url in LiteParseConfig to route OCR requests to external services while maintaining the same merge logic.
How does LiteParse handle OCR failures?
The system differentiates between partial and catastrophic failures. Individual page failures are logged and skipped, but if all OCR tasks fail on pages that required OCR (text-poor pages), the parser returns LiteParseError::Ocr (lines 230-246), ensuring the caller recognizes incomplete extraction rather than receiving silently truncated results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →