How to Configure the Number of Concurrent OCR Workers in LiteParse

Set the num_workers field in LiteParseConfig or pass the --num-workers CLI flag to control exactly how many pages LiteParse processes simultaneously through OCR.

LiteParse is an open-source document parsing library from the run-llama organization that extracts text from PDFs using a hybrid approach of native text extraction and OCR. When processing scanned documents or pages with sparse native text, the library automatically routes content through an OCR pipeline that supports configurable parallelism to maximize throughput. Understanding how to configure the number of concurrent OCR workers in LiteParse allows you to optimize processing speed for high-core servers or reduce memory consumption in constrained environments.

Where the Configuration Lives

The concurrent OCR worker setting propagates through several layers of the LiteParse architecture, from CLI arguments down to the core Rust semaphore implementation.

Rust Core Configuration

In crates/liteparse/src/config.rs at lines 33–34, the LiteParseConfig struct defines the concurrency control:

pub struct LiteParseConfig {
    pub num_workers: usize,
    // ... other fields
}

The default value is computed by default_num_workers(), which returns CPU cores – 1 (minimum 1) by reading std::thread::available_parallelism and clamping the result.

CLI Interface

The command-line interface exposes this setting in crates/liteparse/src/main.rs at lines 97–99:

#[arg(long)]
num_workers: Option<usize>,

This generates the --num-workers flag that overrides the default when invoked.

Language Bindings

For JavaScript and Node.js users, the setting appears in crates/liteparse-napi/src/types.rs at lines 41–45 as num_workers: Option<u32>, ensuring the same concurrency control is available across all supported language bindings.

How Concurrent OCR Workers Work

LiteParse implements concurrency control using a Tokio semaphore that limits blocking OCR operations.

When LiteParse::parse_input is called (in crates/liteparse/src/parser.rs, lines 78–80), it forwards self.config.num_workers to the ocr_and_merge_rendered function in crates/liteparse/src/ocr_merge.rs. This function creates a tokio::sync::Semaphore::new(num_workers) at lines 78–80 and 94–96.

Each page requiring OCR spawns an async task that must acquire a permit from this semaphore before invoking the OCR engine (Tesseract or HTTP-based). Once the permit is released, the next waiting task proceeds. This guarantees that at most num_workers blocking threads are active simultaneously, preventing resource exhaustion and potential deadlocks.

WASM Limitation: The WebAssembly build cannot spawn OS threads, so the LiteParse constructor in crates/liteparse-wasm/src/lib.rs forces num_workers to 1 regardless of configuration.

Methods to Configure Concurrent OCR Workers

You can adjust the worker count via CLI, Python, Node.js, or native Rust APIs.

Command Line Interface

Use the --num-workers flag to override the default CPU-based calculation:

liteparse parse mydoc.pdf \
  --ocr-language eng \
  --num-workers 4 \
  --output result.json

This permits four OCR tasks to run in parallel, reducing total processing time for multi-page scanned documents.

Python API

When using the Python bindings, pass num_workers directly to LiteParseConfig:

from liteparse import LiteParse, LiteParseConfig

cfg = LiteParseConfig(
    ocr_language="eng",
    num_workers=6,               # Configure 6 concurrent workers

    ocr_enabled=True,
)

parser = LiteParse(cfg)
result = parser.parse("mydoc.pdf")
print(result.text)

Node.js API

The Node.js bindings expose the same configuration through the numWorkers property:

const { LiteParse, LiteParseConfig } = require("liteparse");

const cfg = new LiteParseConfig({
  ocrLanguage: "eng",
  numWorkers: 3,      // Set 3 concurrent OCR workers
});

const parser = new LiteParse(cfg);
parser.parse("mydoc.pdf").then(res => {
  console.log(res.text);
});

Native Rust

For programmatic use in Rust, set the field when constructing LiteParseConfig:

use liteparse::{config::LiteParseConfig, parser::LiteParse};

let cfg = LiteParseConfig {
    ocr_language: "eng".into(),
    num_workers: 5,          // Custom concurrency
    ..Default::default()
};

let parser = LiteParse::new(cfg);
let result = parser.parse("mydoc.pdf").await.unwrap();
println!("{}", result.text);

Performance Considerations

Adjusting the number of concurrent OCR workers impacts both speed and resource utilization.

  • High-CPU systems: Increasing workers (e.g., --num-workers 8 on an 8-core machine) maximizes OCR throughput for large documents, though the default CPU cores - 1 already provides near-optimal CPU utilization while leaving resources for the OS and async runtime.

  • Memory constraints: Each OCR task holds a full-resolution bitmap in memory. In containerized environments with strict memory limits, reduce num_workers to prevent out-of-memory errors during batch processing.

  • Remote OCR services: When using HTTP-based OCR servers instead of local Tesseract, the bottleneck is often network latency rather than CPU. Raising the worker count above the number of physical cores can improve throughput by keeping more requests in flight simultaneously.

Summary

  • Default behavior: LiteParse automatically sets concurrent OCR workers to CPU cores - 1 (minimum 1) via default_num_workers() in crates/liteparse/src/config.rs.
  • Configuration methods: Override via --num-workers CLI flag, num_workers in Python, numWorkers in Node.js, or the num_workers field in Rust's LiteParseConfig.
  • Implementation: The ocr_and_merge_rendered function in crates/liteparse/src/ocr_merge.rs uses a tokio::sync::Semaphore to enforce the limit, ensuring no more than the configured number of blocking OCR operations run simultaneously.
  • WASM exception: WebAssembly builds force num_workers to 1 due to thread limitations.
  • Tuning guidance: Increase workers for high-core count servers or remote OCR services; decrease for memory-constrained environments.

Frequently Asked Questions

What is the default number of concurrent OCR workers in LiteParse?

The default is calculated as CPU cores – 1 (minimum 1). This calculation occurs in crates/liteparse/src/config.rs through the default_num_workers() function, which calls std::thread::available_parallelism and subtracts one to reserve a core for the async runtime and OS operations.

Can I set the number of OCR workers when using LiteParse in a browser via WASM?

No. The WebAssembly build in crates/liteparse-wasm/src/lib.rs explicitly forces num_workers to 1 because WASM cannot spawn operating system threads. Regardless of what value you provide in the configuration, the WASM constructor overrides it to ensure single-threaded operation.

How does LiteParse prevent too many OCR tasks from running simultaneously?

LiteParse uses a Tokio semaphore (tokio::sync::Semaphore) initialized with the configured num_workers value. In crates/liteparse/src/ocr_merge.rs, each OCR task must acquire a permit from this semaphore before executing the blocking OCR call. Only when a task completes and releases its permit can another task begin, ensuring strict concurrency limits.

Should I increase num_workers beyond my CPU core count?

Only when using remote OCR services (HTTP-based OCR). For local Tesseract OCR, the default CPU cores - 1 is optimal because OCR is CPU-bound. However, with network-based OCR where latency dominates, increasing workers above the core count keeps more requests in flight and improves overall throughput without overloading local CPU.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →