# How to Configure Parallel OCR Workers for Faster Processing in LiteParse

> Speed up LiteParse processing by configuring parallel OCR workers. Learn how to adjust num_workers to optimize performance based on your CPU cores.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-07

---

**LiteParse controls concurrent OCR processing through the `num_workers` field in `LiteParseConfig`, using a Tokio semaphore to cap parallel tasks at a default of CPU cores minus one.**

When processing documents that contain scanned images or low-quality text, LiteParse from the `run-llama/liteparse` repository offloads optical character recognition (OCR) to a parallelized worker pool. Configuring parallel OCR workers allows you to maximize throughput on multi-core systems while managing memory and I/O pressure from CPU-intensive OCR engines.

## Understanding the Parallel OCR Architecture

The parallelization strategy relies on a semaphore-based concurrency limit defined in the configuration and enforced during the parsing pipeline.

### The Configuration Layer

In [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs), the `LiteParseConfig` struct defines the worker pool size:

```rust
pub struct LiteParseConfig {
    pub num_workers: usize,
    // ... other fields
}

```

The default value calculates available parallelism using `num_cpus::get().saturating_sub(1).max(1)`, ensuring at least one worker while reserving a CPU core for coordination overhead. You can override this default either via the CLI flag in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) (`--num_workers`) or programmatically by constructing a custom `LiteParseConfig`.

### The Concurrency Control Mechanism

The actual enforcement occurs in [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs). The `ocr_and_merge_rendered` function initializes an `Arc<tokio::sync::Semaphore>` with `num_workers` permits:

```rust
let sem = Arc::new(tokio::sync::Semaphore::new(num_workers));

```

Each OCR task spawned via `tokio::task::spawn_blocking` must acquire a permit via `sem.acquire_owned()` before invoking the OCR engine. This guarantees that no more than `num_workers` OCR operations execute simultaneously, preventing resource exhaustion when processing large batches.

## How to Configure Parallel OCR Workers

You can adjust worker concurrency through command-line arguments or Rust code depending on your integration method.

### Using the Command-Line Interface

When running the `liteparse` binary, pass the `--num_workers` flag to override the default:

```bash

# Use default worker count (CPU cores - 1)

liteparse parse document.pdf --ocr

# Explicitly allocate 8 parallel OCR workers

liteparse parse document.pdf --ocr --num_workers 8

# Batch processing with high concurrency

liteparse batch-parse ./inputs ./outputs --ocr --num_workers 16

```

The CLI parser in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) injects this value into the configuration before initializing the parser.

### Configuring Programmatically in Rust

For library integrations, construct a `LiteParseConfig` and set `num_workers` before creating the parser instance:

```rust
use liteparse::config::LiteParseConfig;
use liteparse::parser::LiteParse;

fn main() {
    // Initialize with defaults
    let mut cfg = LiteParseConfig::default();
    
    // Configure 6 parallel OCR workers
    cfg.num_workers = 6;
    
    // Instantiate parser with custom configuration
    let parser = LiteParse::new(cfg);
    // parser.parse("scanned_document.pdf") ...
}

```

The `LiteParse::new(cfg)` constructor accepts your configuration and propagates the worker count to the internal OCR pipeline defined in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs).

## How the Worker Pool Executes OCR Tasks

The parallel execution follows a four-stage pipeline implemented across the codebase:

1. **Page Selection** – The parser identifies which pages require OCR (images or garbled text) through the `render_pages_for_ocr` logic.

2. **Task Spawning** – For each selected page, `ocr_and_merge_rendered` spawns a blocking Tokio task. Before executing OCR, the task awaits a semaphore permit using `sem.acquire_owned()`.

3. **Concurrency Limiting** – The semaphore initialized with `num_workers` permits ensures only that many tasks hold permits simultaneously. Additional tasks wait in a FIFO queue until permits become available.

4. **Result Merging** – Completed OCR results merge back into the document structure. Failed tasks log errors without blocking other workers or halting the pipeline.

Because OCR engines—whether local Tesseract instances or remote HTTP services—consume significant CPU or network I/O, tuning `num_workers` lets you balance resource utilization against parsing latency. High-core machines processing batches benefit from increased worker counts, while I/O-bound remote OCR services may require conservative limits to avoid rate limiting.

## Platform-Specific Considerations

**WebAssembly Limitations:** When compiling for WebAssembly targets in [`crates/liteparse-wasm/src/lib.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse-wasm/src/lib.rs), the code forces `num_workers = 1`. Browser environments cannot safely spawn multiple native threads, forcing sequential OCR processing regardless of configuration settings.

## Summary

- **Configuration Location:** The `num_workers` field in [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs) controls the parallel OCR worker pool, defaulting to `CPU cores - 1` (minimum 1).
- **Concurrency Mechanism:** A Tokio semaphore in [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) enforces the limit, ensuring controlled resource consumption.
- **CLI Override:** Use `--num_workers` in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) to adjust concurrency without code changes.
- **Programmatic Control:** Set `cfg.num_workers` on `LiteParseConfig` before passing it to `LiteParse::new()`.
- **Platform Constraints:** WebAssembly builds automatically restrict workers to 1 due to browser threading limitations.

## Frequently Asked Questions

### What is the default number of OCR workers in LiteParse?

LiteParse defaults to `num_cpus::get().saturating_sub(1).max(1)`, which equals your CPU core count minus one, with a floor of one worker. This reserves a core for coordination while maximizing parallel OCR throughput.

### How does LiteParse handle OCR task failures in parallel mode?

Individual OCR task failures are logged but do not block other workers or halt the parsing pipeline. The semaphore permit releases back to the pool when a task completes (successfully or not), allowing pending tasks to acquire permits and continue processing.

### Can I use different OCR engines with the same worker pool configuration?

Yes. The `num_workers` setting applies to the concurrency layer regardless of the specific `OcrEngine` implementation. You can inject custom engines using `LiteParse::with_ocr_engine()` while retaining the same semaphore-controlled worker pool defined in your `LiteParseConfig`.

### Why is the worker count limited to 1 in WebAssembly builds?

The WebAssembly target in [`crates/liteparse-wasm/src/lib.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse-wasm/src/lib.rs) forces `num_workers = 1` because browser environments lack support for spawning multiple native threads through Tokio's blocking task system, making true parallel OCR impossible in WASM contexts.