Performance Implications of Enabling vs Disabling OCR in LiteParse

Enabling OCR in LiteParse makes parsing 2–5× slower for scanned documents by adding CPU-intensive rasterization, OCR inference, and result merging, while disabling it skips these steps for near-native text extraction speeds.

When working with PDF documents in the run-llama/liteparse repository, understanding the performance implications of enabling versus disabling OCR is critical for optimizing throughput. The parser's optional OCR pipeline, controlled by the ocr_enabled flag in LiteParseConfig, introduces significant computational overhead that can dramatically impact processing times depending on your document types and hardware configuration.

The Three-Stage OCR Pipeline

When ocr_enabled is set to true in /crates/liteparse/src/config.rs (lines 44-45), the parser executes three additional expensive stages for each flagged page.

1. Page Rasterization at Configured DPI

In /crates/liteparse/src/parser.rs (lines 102-115), the code triggers ocr_merge::render_pages_for_ocr(&document, &pages, self.config.dpi) to convert PDF pages into bitmap images. This step defaults to 150 DPI, but higher values generate exponentially larger bitmaps that slow both rendering and subsequent OCR processing. The rasterization is CPU-intensive and represents the first major bottleneck when OCR is active.

2. Asynchronous OCR Engine Execution

The rendered images pass to ocr_merge::ocr_and_merge_rendered in /crates/liteparse/src/ocr_merge.rs (lines 72-89), which spawns asynchronous workers (defaulting to CPU cores minus one) to process images. Depending on configuration, this invokes either:

  • Built-in Tesseract: Local CPU-bound inference
  • HTTP OCR Server: Remote processing that may offer GPU acceleration but adds network round-trip latency

3. Result Merging and Overlap Detection

Finally, the system merges OCR results back into native text items at lines 130-156 in ocr_merge.rs. The merging logic specifically discards OCR output that overlaps existing native text beyond a small tolerance (lines 140-146), preventing noise while adding computational overhead for geometric comparison and text cleaning.

Performance Factors and Bottlenecks

The exact performance penalty varies based on these configurable factors:

  • Page Complexity: Documents with many images or low native-text coverage trigger OCR on more pages, multiplying the rendering and inference workload.
  • DPI Setting: The config.dpi value directly impacts bitmap size; higher resolutions increase memory pressure and processing time linearly.
  • OCR Engine Architecture: Local Tesseract runs are strictly CPU-bound, while HTTP endpoints introduce variable network latency and bandwidth constraints.
  • Worker Parallelism: The --num-workers parameter controls concurrency. More workers speed up batch processing but increase peak CPU utilization.
  • Network Conditions: For remote OCR configurations, latency and connection stability significantly affect total throughput.

The Fast Path: OCR Disabled

When OCR is disabled via ocr_enabled: false—typically set through the --no-ocr CLI flag in /crates/liteparse/src/main.rs (lines 51-55)—the parser bypasses the entire rendering-and-OCR block. Execution flows directly from native PDF text extraction to spatial projection, eliminating rasterization, inference, and merge steps entirely. This represents the fastest possible code path with minimal memory overhead.

Selective OCR Optimization

Even when enabled, LiteParse minimizes unnecessary work through intelligent filtering. According to /crates/liteparse/src/ocr_merge.rs (lines 44-45), the engine only processes pages where native text extraction yields fewer than 20 characters or less than 15% coverage, plus any pages containing images. Additionally, the rendering step executes once per page, with the resulting bitmap reused across all OCR workers to eliminate duplicate rasterization.

Configuration Examples

Command Line Interface


# OCR enabled (default) - slower for scanned documents

liteparse parse document.pdf --output json

# Disable OCR for maximum speed on text-based PDFs

liteparse parse document.pdf --no-ocr --output json

The --no-ocr flag sets ocr_enabled: false in the LiteParseConfig struct as implemented in main.rs.

Node.js SDK

import { LiteParse } from "liteparse";

// With OCR (default) - triggers rasterization and inference
const parser = new LiteParse({
  ocrEnabled: true,
  dpi: 150,
  numWorkers: 4,
});

// Without OCR - native extraction only, minimal overhead
const fastParser = new LiteParse({ ocrEnabled: false });

The ocrEnabled option maps to config.ocr_enabled in the Rust core (/crates/liteparse-wasm/src/lib.rs lines 59-60).

Python SDK

from liteparse import LiteParse

# OCR enabled (default configuration)

lp = LiteParse(ocr_enabled=True)
result = lp.parse("document.pdf")

# Optimized for speed without OCR pipeline

lp_fast = LiteParse(ocr_enabled=False)
result_fast = lp_fast.parse("document.pdf")

The Python wrapper forwards this flag to the underlying Rust config (/crates/liteparse-python/src/lib.rs lines 196-202).

Summary

  • Enabling OCR adds three computational stages—DPI-based rasterization, OCR engine inference, and result merging—that collectively slow parsing by 2–5× for scan-heavy documents.
  • Disabling OCR via ocr_enabled: false or the --no-ocr CLI flag eliminates these bottlenecks, providing the fastest path through native text extraction only.
  • Selective processing ensures only pages with insufficient native text (<20 characters or <15% coverage) or containing images undergo OCR, minimizing unnecessary overhead.
  • Configuration tuning through DPI settings, worker counts, and engine choice (local Tesseract vs. HTTP) allows trading accuracy for speed when OCR is required.

Frequently Asked Questions

How much slower is LiteParse with OCR enabled?

Benchmarks indicate that enabling OCR makes parsing approximately 2–5× slower for documents containing scanned images or photos, while text-based PDFs with native fonts see only modest overhead since the selective logic limits OCR to pages lacking extractable text.

Does LiteParse OCR every page or only specific ones?

The engine uses selective OCR logic defined in ocr_merge.rs (lines 44-45). Only pages yielding fewer than 20 characters or less than 15% coverage from native extraction, plus pages explicitly containing images, are sent through the OCR pipeline.

Can I use a GPU-accelerated OCR engine with LiteParse?

Yes. While the built-in Tesseract implementation in /crates/liteparse/src/ocr/tesseract.rs runs CPU-bound locally, configuring an HTTP OCR server endpoint allows you to leverage GPU-accelerated remote engines. Note that this trades CPU load for network latency and bandwidth requirements.

What DPI setting provides the best balance of speed and accuracy?

The default 150 DPI offers a practical compromise for most documents. Lower values reduce rasterization time and memory usage but may miss small text, while higher values improve OCR accuracy for fine print at the cost of significantly increased processing overhead and memory consumption.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →