How to Optimize OCR Settings in LiteParse for Better Performance: 7 Proven Tuning Methods

Tune LiteParse OCR performance by adjusting dpi, num_workers, and ocr_language in LiteParseConfig, disabling OCR with --no-ocr when native text exists, or offloading recognition to an external HTTP server via --ocr-server-url.

LiteParse intelligently runs OCR only on document regions that lack native text, such as scanned images embedded in PDFs. Because the OCR pipeline is configurable through both the CLI and the LiteParseConfig struct, learning how to optimize OCR settings in LiteParse can dramatically reduce parsing time without sacrificing accuracy.

Core OCR Configuration Options

In crates/liteparse/src/config.rs, the LiteParseConfig struct groups every OCR-related runtime setting. The fields that most directly impact parsing speed include:

  • ocr_enabled (bool, default true) — Set to false or pass --no-ocr to bypass the entire OCR pipeline when source documents already contain selectable text.
  • ocr_language (String, default "eng") — Restricts Tesseract to a single language model; loading unnecessary models increases memory usage and startup latency.
  • ocr_server_url (Option<String>, default None) — Routes image bytes to a remote HTTP OCR service instead of local Tesseract, freeing the parser from CPU-bound recognition.
  • tessdata_path (Option<String>) — Overrides the TESSDATA_PREFIX environment variable to point at a lean directory containing only required .traineddata files.
  • dpi (f32, default 150.0) — Controls the resolution of rasterized page images in render.rs; lower values produce smaller images and faster OCR at the cost of fine detail.
  • num_workers (usize, default CPU cores - 1) — Sets the number of pages processed concurrently by the OCR engine.
  • preserve_very_small_text (bool, default false) — When disabled, the merge step in ocr_merge.rs discards low-resolution OCR results, speeding up layout reconstruction.

How OCR Settings Flow Through the Pipeline

The pipeline begins in src/main.rs, where CLI flags such as --ocr-language and --dpi hydrate a LiteParseConfig instance. During document processing, parser.rs reads config.ocr_enabled to decide whether to invoke the local ocr::tesseract backend or the remote ocr::http_simple client. Both backends implement the OcrEngine trait defined in crates/liteparse/src/ocr/mod.rs. The rendering step in crates/liteparse/src/render.rs rasterizes each page at config.dpi and feeds the resulting PNG bytes into the active engine. Finally, crates/liteparse/src/ocr_merge.rs blends OCR output with native PDF text, respecting preserve_very_small_text to filter out low-resolution noise.

Seven Methods to Optimize OCR Performance in LiteParse

Reduce DPI — The render.rs module rasterizes pages at the configured dpi before sending them to the engine. Dropping the default of 150.0 to 100.0–120.0 for standard text documents can reduce OCR time by roughly 30% without materially hurting accuracy.

Limit Target Pages — Use --target-pages or LiteParseConfig::target_pages to constrain parsing to specific pages. Because the parser still renders all pages if OCR is enabled, restricting the scope shrinks the total OCR workload.

Disable OCR for Native Text PDFs — Many PDFs already contain embedded text. Passing --no-ocr or setting ocr_enabled: false skips the entire OCR pipeline, making this the fastest option when scans are absent.

Select a Single Language Model — Tesseract loads .traineddata at startup. Keeping ocr_language set to "eng" for English documents avoids the CPU and memory overhead of loading multiple models.

Offload to an External OCR Server — Supplying --ocr-server-url delegates recognition to a remote service via ocr/http_simple.rs. This decouples parsing from OCR latency and allows the server to run GPU-accelerated models independently.

Tune Concurrent Workers — The num_workers field controls parallel page processing. For CPU-bound local Tesseract, matching this value to the number of physical cores typically yields optimal throughput; for HTTP-based engines, a higher count can mask network latency.

Trim the Tessdata Directory — Pointing tessdata_path to a directory that contains only the required .traineddata files minimizes file-system lookup time and reduces memory pressure during model loading.

Practical Code Examples

CLI Example: Lower DPI and Restrict Workers


# Parse only pages 1-10, use 100 dpi images, and limit to 2 concurrent OCR workers

lit parse document.pdf \
    --target-pages "1-10" \
    --dpi 100 \
    --num-workers 2 \
    -o out.json

Rust API: Custom Config for a CPU-Limited Environment

use liteparse::LiteParseConfig;
use liteparse::OutputFormat;
use liteparse::LiteParse; // assume the library entry point

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let config = LiteParseConfig {
        ocr_language: "eng".into(),
        ocr_enabled: true,
        ocr_server_url: None,               // use built-in Tesseract
        tessdata_path: Some("./tessdata".into()),
        max_pages: 500,
        target_pages: Some("1-20".into()),
        dpi: 120.0,                         // lower DPI for speed
        output_format: OutputFormat::Json,
        preserve_very_small_text: false,
        password: None,
        quiet: false,
        num_workers: 4,                     // match number of cores
    };

    let parser = LiteParse::new(config);
    let result = parser.parse_path("sample.pdf")?;
    println!("{}", serde_json::to_string_pretty(&result)?);
    Ok(())
}

Python Wrapper: Disable OCR for Text-Heavy PDFs

from liteparse import LiteParse, LiteParseConfig, OutputFormat

cfg = LiteParseConfig(
    ocr_enabled=False,          # skip OCR entirely

    output_format=OutputFormat.JSON,
)

parser = LiteParse(cfg)
out = parser.parse_path("text_only.pdf")
print(out.json())

Key Source Files for OCR Customization

Summary

  • LiteParse skips OCR on pages that already contain native text, but rasterizes images at config.dpi when OCR is required.
  • Lowering dpi from 150.0 to 100.0–120.0 is the fastest way to reduce local OCR latency.
  • Set ocr_enabled: false or use --no-ocr to eliminate the pipeline entirely for text-native PDFs.
  • Match num_workers to physical CPU cores for Tesseract, or increase it to hide latency for HTTP-based engines.
  • Point tessdata_path to a minimal directory and restrict ocr_language to a single model to cut startup overhead.
  • Use --ocr-server-url to offload recognition and scale parsing independently from OCR resources.

Frequently Asked Questions

Can I disable OCR completely in LiteParse if my PDFs already have selectable text?

Yes. Set ocr_enabled: false in LiteParseConfig or pass the --no-ocr CLI flag. This causes parser.rs to bypass both the rendering and recognition stages, returning native text immediately and dramatically improving parse speed.

What is the optimal DPI setting for fast OCR in LiteParse?

The default is 150.0, but most standard text documents parse accurately at 100.0 to 120.0 DPI. Because render.rs rasterizes every page image at this resolution, lowering DPI reduces memory pressure and can cut OCR time by approximately 30%.

How does LiteParse handle OCR for documents with mixed languages?

LiteParse uses Tesseract’s language model specified by ocr_language in config.rs. If you load multiple languages, Tesseract initializes every corresponding .traineddata file, which increases memory usage and CPU overhead. For best performance, specify only the dominant language of your document.

Is it possible to run OCR on a separate machine while using LiteParse locally?

Yes. Configure ocr_server_url with the address of an HTTP service that implements LiteParse’s JSON API. The ocr/http_simple.rs client will forward PNG bytes to the remote server, allowing the local parser to continue instantly while the dedicated machine handles recognition.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →