# How to Optimize OCR Settings in LiteParse for Better Performance: 7 Proven Tuning Methods

> Optimize LiteParse OCR performance with 7 proven tuning methods. Learn to adjust DPI, workers, language, and leverage external OCR servers for faster results.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: performance
- Published: 2026-06-07

---

**Tune LiteParse OCR performance by adjusting `dpi`, `num_workers`, and `ocr_language` in `LiteParseConfig`, disabling OCR with `--no-ocr` when native text exists, or offloading recognition to an external HTTP server via `--ocr-server-url`.**

LiteParse intelligently runs OCR only on document regions that lack native text, such as scanned images embedded in PDFs. Because the OCR pipeline is configurable through both the CLI and the `LiteParseConfig` struct, learning how to optimize OCR settings in LiteParse can dramatically reduce parsing time without sacrificing accuracy.

## Core OCR Configuration Options

In [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs), the `LiteParseConfig` struct groups every OCR-related runtime setting. The fields that most directly impact parsing speed include:

- `ocr_enabled` (`bool`, default `true`) — Set to `false` or pass `--no-ocr` to bypass the entire OCR pipeline when source documents already contain selectable text.
- `ocr_language` (`String`, default `"eng"`) — Restricts Tesseract to a single language model; loading unnecessary models increases memory usage and startup latency.
- `ocr_server_url` (`Option<String>`, default `None`) — Routes image bytes to a remote HTTP OCR service instead of local Tesseract, freeing the parser from CPU-bound recognition.
- `tessdata_path` (`Option<String>`) — Overrides the `TESSDATA_PREFIX` environment variable to point at a lean directory containing only required `.traineddata` files.
- `dpi` (`f32`, default `150.0`) — Controls the resolution of rasterized page images in [`render.rs`](https://github.com/run-llama/liteparse/blob/main/render.rs); lower values produce smaller images and faster OCR at the cost of fine detail.
- `num_workers` (`usize`, default `CPU cores - 1`) — Sets the number of pages processed concurrently by the OCR engine.
- `preserve_very_small_text` (`bool`, default `false`) — When disabled, the merge step in [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) discards low-resolution OCR results, speeding up layout reconstruction.

## How OCR Settings Flow Through the Pipeline

The pipeline begins in [`src/main.rs`](https://github.com/run-llama/liteparse/blob/main/src/main.rs), where CLI flags such as `--ocr-language` and `--dpi` hydrate a `LiteParseConfig` instance. During document processing, [`parser.rs`](https://github.com/run-llama/liteparse/blob/main/parser.rs) reads `config.ocr_enabled` to decide whether to invoke the local `ocr::tesseract` backend or the remote `ocr::http_simple` client. Both backends implement the `OcrEngine` trait defined in [`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs). The rendering step in [`crates/liteparse/src/render.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/render.rs) rasterizes each page at `config.dpi` and feeds the resulting PNG bytes into the active engine. Finally, [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) blends OCR output with native PDF text, respecting `preserve_very_small_text` to filter out low-resolution noise.

## Seven Methods to Optimize OCR Performance in LiteParse

**Reduce DPI** — The [`render.rs`](https://github.com/run-llama/liteparse/blob/main/render.rs) module rasterizes pages at the configured `dpi` before sending them to the engine. Dropping the default of `150.0` to `100.0`–`120.0` for standard text documents can reduce OCR time by roughly 30% without materially hurting accuracy.

**Limit Target Pages** — Use `--target-pages` or `LiteParseConfig::target_pages` to constrain parsing to specific pages. Because the parser still renders all pages if OCR is enabled, restricting the scope shrinks the total OCR workload.

**Disable OCR for Native Text PDFs** — Many PDFs already contain embedded text. Passing `--no-ocr` or setting `ocr_enabled: false` skips the entire OCR pipeline, making this the fastest option when scans are absent.

**Select a Single Language Model** — Tesseract loads `.traineddata` at startup. Keeping `ocr_language` set to `"eng"` for English documents avoids the CPU and memory overhead of loading multiple models.

**Offload to an External OCR Server** — Supplying `--ocr-server-url` delegates recognition to a remote service via [`ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/ocr/http_simple.rs). This decouples parsing from OCR latency and allows the server to run GPU-accelerated models independently.

**Tune Concurrent Workers** — The `num_workers` field controls parallel page processing. For CPU-bound local Tesseract, matching this value to the number of physical cores typically yields optimal throughput; for HTTP-based engines, a higher count can mask network latency.

**Trim the Tessdata Directory** — Pointing `tessdata_path` to a directory that contains only the required `.traineddata` files minimizes file-system lookup time and reduces memory pressure during model loading.

## Practical Code Examples

### CLI Example: Lower DPI and Restrict Workers

```bash

# Parse only pages 1-10, use 100 dpi images, and limit to 2 concurrent OCR workers

lit parse document.pdf \
    --target-pages "1-10" \
    --dpi 100 \
    --num-workers 2 \
    -o out.json

```

### Rust API: Custom Config for a CPU-Limited Environment

```rust
use liteparse::LiteParseConfig;
use liteparse::OutputFormat;
use liteparse::LiteParse; // assume the library entry point

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let config = LiteParseConfig {
        ocr_language: "eng".into(),
        ocr_enabled: true,
        ocr_server_url: None,               // use built-in Tesseract
        tessdata_path: Some("./tessdata".into()),
        max_pages: 500,
        target_pages: Some("1-20".into()),
        dpi: 120.0,                         // lower DPI for speed
        output_format: OutputFormat::Json,
        preserve_very_small_text: false,
        password: None,
        quiet: false,
        num_workers: 4,                     // match number of cores
    };

    let parser = LiteParse::new(config);
    let result = parser.parse_path("sample.pdf")?;
    println!("{}", serde_json::to_string_pretty(&result)?);
    Ok(())
}

```

### Python Wrapper: Disable OCR for Text-Heavy PDFs

```python
from liteparse import LiteParse, LiteParseConfig, OutputFormat

cfg = LiteParseConfig(
    ocr_enabled=False,          # skip OCR entirely

    output_format=OutputFormat.JSON,
)

parser = LiteParse(cfg)
out = parser.parse_path("text_only.pdf")
print(out.json())

```

## Key Source Files for OCR Customization

- [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs) — Holds `LiteParseConfig` and all OCR-related runtime fields.
- [`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs) — Defines the `OcrEngine` trait and shared structures such as `OcrResult` and `OcrOptions`.
- [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs) — Feature-gated Tesseract implementation that passes language and image data to the native library.
- [`crates/liteparse/src/ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/http_simple.rs) — Minimal HTTP client that forwards image bytes to a remote OCR endpoint.
- [`crates/liteparse/src/render.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/render.rs) — Generates raster images at the configured DPI for the OCR engine.
- [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) — Merges OCR results with native PDF text and applies the `preserve_very_small_text` filter.

## Summary

- LiteParse skips OCR on pages that already contain native text, but rasterizes images at `config.dpi` when OCR is required.
- Lowering `dpi` from `150.0` to `100.0`–`120.0` is the fastest way to reduce local OCR latency.
- Set `ocr_enabled: false` or use `--no-ocr` to eliminate the pipeline entirely for text-native PDFs.
- Match `num_workers` to physical CPU cores for Tesseract, or increase it to hide latency for HTTP-based engines.
- Point `tessdata_path` to a minimal directory and restrict `ocr_language` to a single model to cut startup overhead.
- Use `--ocr-server-url` to offload recognition and scale parsing independently from OCR resources.

## Frequently Asked Questions

### Can I disable OCR completely in LiteParse if my PDFs already have selectable text?

Yes. Set `ocr_enabled: false` in `LiteParseConfig` or pass the `--no-ocr` CLI flag. This causes [`parser.rs`](https://github.com/run-llama/liteparse/blob/main/parser.rs) to bypass both the rendering and recognition stages, returning native text immediately and dramatically improving parse speed.

### What is the optimal DPI setting for fast OCR in LiteParse?

The default is `150.0`, but most standard text documents parse accurately at `100.0` to `120.0` DPI. Because [`render.rs`](https://github.com/run-llama/liteparse/blob/main/render.rs) rasterizes every page image at this resolution, lowering DPI reduces memory pressure and can cut OCR time by approximately 30%.

### How does LiteParse handle OCR for documents with mixed languages?

LiteParse uses Tesseract’s language model specified by `ocr_language` in [`config.rs`](https://github.com/run-llama/liteparse/blob/main/config.rs). If you load multiple languages, Tesseract initializes every corresponding `.traineddata` file, which increases memory usage and CPU overhead. For best performance, specify only the dominant language of your document.

### Is it possible to run OCR on a separate machine while using LiteParse locally?

Yes. Configure `ocr_server_url` with the address of an HTTP service that implements LiteParse’s JSON API. The [`ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/ocr/http_simple.rs) client will forward PNG bytes to the remote server, allowing the local parser to continue instantly while the dedicated machine handles recognition.