How to Configure Parallel OCR Workers for Faster Processing in LiteParse
LiteParse controls concurrent OCR processing through the num_workers field in LiteParseConfig, using a Tokio semaphore to cap parallel tasks at a default of CPU cores minus one.
When processing documents that contain scanned images or low-quality text, LiteParse from the run-llama/liteparse repository offloads optical character recognition (OCR) to a parallelized worker pool. Configuring parallel OCR workers allows you to maximize throughput on multi-core systems while managing memory and I/O pressure from CPU-intensive OCR engines.
Understanding the Parallel OCR Architecture
The parallelization strategy relies on a semaphore-based concurrency limit defined in the configuration and enforced during the parsing pipeline.
The Configuration Layer
In crates/liteparse/src/config.rs, the LiteParseConfig struct defines the worker pool size:
pub struct LiteParseConfig {
pub num_workers: usize,
// ... other fields
}
The default value calculates available parallelism using num_cpus::get().saturating_sub(1).max(1), ensuring at least one worker while reserving a CPU core for coordination overhead. You can override this default either via the CLI flag in crates/liteparse/src/main.rs (--num_workers) or programmatically by constructing a custom LiteParseConfig.
The Concurrency Control Mechanism
The actual enforcement occurs in crates/liteparse/src/ocr_merge.rs. The ocr_and_merge_rendered function initializes an Arc<tokio::sync::Semaphore> with num_workers permits:
let sem = Arc::new(tokio::sync::Semaphore::new(num_workers));
Each OCR task spawned via tokio::task::spawn_blocking must acquire a permit via sem.acquire_owned() before invoking the OCR engine. This guarantees that no more than num_workers OCR operations execute simultaneously, preventing resource exhaustion when processing large batches.
How to Configure Parallel OCR Workers
You can adjust worker concurrency through command-line arguments or Rust code depending on your integration method.
Using the Command-Line Interface
When running the liteparse binary, pass the --num_workers flag to override the default:
# Use default worker count (CPU cores - 1)
liteparse parse document.pdf --ocr
# Explicitly allocate 8 parallel OCR workers
liteparse parse document.pdf --ocr --num_workers 8
# Batch processing with high concurrency
liteparse batch-parse ./inputs ./outputs --ocr --num_workers 16
The CLI parser in crates/liteparse/src/main.rs injects this value into the configuration before initializing the parser.
Configuring Programmatically in Rust
For library integrations, construct a LiteParseConfig and set num_workers before creating the parser instance:
use liteparse::config::LiteParseConfig;
use liteparse::parser::LiteParse;
fn main() {
// Initialize with defaults
let mut cfg = LiteParseConfig::default();
// Configure 6 parallel OCR workers
cfg.num_workers = 6;
// Instantiate parser with custom configuration
let parser = LiteParse::new(cfg);
// parser.parse("scanned_document.pdf") ...
}
The LiteParse::new(cfg) constructor accepts your configuration and propagates the worker count to the internal OCR pipeline defined in crates/liteparse/src/parser.rs.
How the Worker Pool Executes OCR Tasks
The parallel execution follows a four-stage pipeline implemented across the codebase:
-
Page Selection – The parser identifies which pages require OCR (images or garbled text) through the
render_pages_for_ocrlogic. -
Task Spawning – For each selected page,
ocr_and_merge_renderedspawns a blocking Tokio task. Before executing OCR, the task awaits a semaphore permit usingsem.acquire_owned(). -
Concurrency Limiting – The semaphore initialized with
num_workerspermits ensures only that many tasks hold permits simultaneously. Additional tasks wait in a FIFO queue until permits become available. -
Result Merging – Completed OCR results merge back into the document structure. Failed tasks log errors without blocking other workers or halting the pipeline.
Because OCR engines—whether local Tesseract instances or remote HTTP services—consume significant CPU or network I/O, tuning num_workers lets you balance resource utilization against parsing latency. High-core machines processing batches benefit from increased worker counts, while I/O-bound remote OCR services may require conservative limits to avoid rate limiting.
Platform-Specific Considerations
WebAssembly Limitations: When compiling for WebAssembly targets in crates/liteparse-wasm/src/lib.rs, the code forces num_workers = 1. Browser environments cannot safely spawn multiple native threads, forcing sequential OCR processing regardless of configuration settings.
Summary
- Configuration Location: The
num_workersfield incrates/liteparse/src/config.rscontrols the parallel OCR worker pool, defaulting toCPU cores - 1(minimum 1). - Concurrency Mechanism: A Tokio semaphore in
crates/liteparse/src/ocr_merge.rsenforces the limit, ensuring controlled resource consumption. - CLI Override: Use
--num_workersincrates/liteparse/src/main.rsto adjust concurrency without code changes. - Programmatic Control: Set
cfg.num_workersonLiteParseConfigbefore passing it toLiteParse::new(). - Platform Constraints: WebAssembly builds automatically restrict workers to 1 due to browser threading limitations.
Frequently Asked Questions
What is the default number of OCR workers in LiteParse?
LiteParse defaults to num_cpus::get().saturating_sub(1).max(1), which equals your CPU core count minus one, with a floor of one worker. This reserves a core for coordination while maximizing parallel OCR throughput.
How does LiteParse handle OCR task failures in parallel mode?
Individual OCR task failures are logged but do not block other workers or halt the parsing pipeline. The semaphore permit releases back to the pool when a task completes (successfully or not), allowing pending tasks to acquire permits and continue processing.
Can I use different OCR engines with the same worker pool configuration?
Yes. The num_workers setting applies to the concurrency layer regardless of the specific OcrEngine implementation. You can inject custom engines using LiteParse::with_ocr_engine() while retaining the same semaphore-controlled worker pool defined in your LiteParseConfig.
Why is the worker count limited to 1 in WebAssembly builds?
The WebAssembly target in crates/liteparse-wasm/src/lib.rs forces num_workers = 1 because browser environments lack support for spawning multiple native threads through Tokio's blocking task system, making true parallel OCR impossible in WASM contexts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →