How the num_workers Configuration Impacts OCR Performance in LiteParse

The num_workers setting controls concurrent OCR processing through a Tokio semaphore, defaulting to CPU cores minus one to balance speed and resource usage.

The num_workers configuration in the run-llama/liteparse repository determines how many pages the OCR engine processes simultaneously, directly impacting both throughput and system resource consumption. Understanding how this parameter interacts with the underlying concurrency mechanism allows developers to optimize document parsing for their specific hardware constraints.

Understanding the Default num_workers Calculation

LiteParse calculates the default worker count dynamically based on available hardware. In crates/liteparse/src/config.rs (lines 58-64), the default_num_workers() function determines the value by taking the number of CPU cores minus one, with a hard minimum of 1.

This default strategy ensures that on a typical multi-core machine, LiteParse spawns one worker per core while reserving a thread for coordination overhead. The result is near-optimal parallelism for CPU-bound OCR workloads without oversubscribing the processor or starving the system of computational resources.

How num_workers Controls OCR Concurrency

The configuration value acts as a concurrency limiter throughout the OCR pipeline, enforced through Rust's async runtime primitives.

The Tokio Semaphore Implementation

In crates/liteparse/src/ocr_merge.rs (lines 90-95), LiteParse creates a Tokio semaphore sized to the configured num_workers value (clamped to at least 1). Each OCR task must acquire a permit from this semaphore before invoking the OCR engine.

This mechanism guarantees that at most num_workers OCR tasks execute simultaneously, regardless of how many pages await processing. The semaphore provides backpressure, preventing memory exhaustion and CPU thrashing when processing large documents.

Runtime Flow and Task Distribution

As implemented in crates/liteparse/src/parser.rs (lines 71-79), the configured worker count propagates from LiteParseConfig to the OCR merging function. The parser passes this value to the dispatcher, which forwards it to the OCR pipeline where the semaphore governs actual execution.

Performance Trade-offs and Tuning Strategies

Adjusting num_workers creates a direct trade-off between processing speed and resource utilization.

  • Higher num_workers increases parallelism, reducing total OCR time for CPU-bound workloads by processing multiple page bitmaps concurrently.
  • Too many workers (exceeding physical cores) introduces thread contention, increased context switching, and higher memory usage, potentially slowing overall parse time.
  • Lower num_workers (e.g., 1) minimizes CPU load and memory pressure but extends processing time; this is mandatory for constrained environments such as WASM builds where the value is forced to 1.

The default provides an optimal balance for most desktop and server deployments, but adjustment becomes necessary when running alongside other CPU-intensive applications or when operating on single-core containers.

Configuring num_workers Across Interfaces

LiteParse exposes num_workers through multiple APIs, defined in crates/liteparse/src/config.rs (lines 5-30) within the LiteParseConfig struct.

Rust API

Configure programmatically before instantiating the parser:

use liteparse::{LiteParse, LiteParseConfig};

#[tokio::main]
async fn main() {
    let mut cfg = LiteParseConfig::default();
    cfg.num_workers = 4;  // Explicitly set 4 parallel OCR workers

    let parser = LiteParse::new(cfg);
    let result = parser.parse("document.pdf").await.unwrap();
    println!("Parsed {} pages", result.pages.len());
}

Command Line Interface

Set via the --num-workers flag parsed in crates/liteparse/src/main.rs (lines 89-94):

liteparse parse document.pdf --num-workers 8

Python Bindings

Accessed through the configuration bridge in liteparse-python/src/lib.rs:

from liteparse import LiteParse, LiteParseConfig

cfg = LiteParseConfig()
cfg.num_workers = 6  # Use six concurrent OCR workers

parser = LiteParse(cfg)

result = parser.parse("document.pdf")
print("Pages:", len(result.pages))

Node.js Bindings

Exposed via the N-API types defined in crates/liteparse-napi/src/types.rs (lines 37-80):

const { LiteParse, LiteParseConfig } = require('liteparse');

const cfg = new LiteParseConfig();
cfg.num_workers = 2;  // Two parallel OCR workers
const parser = new LiteParse(cfg);

parser.parse('document.pdf')
  .then(res => console.log(`Parsed ${res.pages.length} pages`));

Summary

  • num_workers controls OCR concurrency via a Tokio semaphore in ocr_merge.rs, capping simultaneous page processing.
  • The default value uses CPU cores minus one (minimum 1) to maximize throughput without system overload.
  • Increasing the value speeds up OCR on multi-core machines, while decreasing it conserves resources for constrained environments.
  • All LiteParse interfaces (Rust, CLI, Python, Node.js) expose this configuration through LiteParseConfig or equivalent flags.

Frequently Asked Questions

What is the optimal num_workers value for OCR performance?

The optimal value typically equals your physical CPU core count minus one, which is the default calculation in config.rs. For a machine with 8 cores, setting num_workers to 7 or 8 usually provides the best throughput. However, if running LiteParse alongside other CPU-intensive applications, reducing this value prevents resource contention and may improve overall system responsiveness.

Can I set num_workers higher than my CPU core count?

While technically possible, setting num_workers higher than your physical core count typically degrades performance. According to the semaphore implementation in ocr_merge.rs, excess workers cause context switching overhead and memory pressure without increasing actual parallelism, potentially slowing OCR processing due to thread contention.

Why does LiteParse force num_workers to 1 in WASM environments?

WebAssembly environments execute within single-threaded or highly constrained JavaScript runtimes that lack true POSIX thread support. LiteParse forces num_workers to 1 in WASM builds to prevent deadlocks and ensure compatibility, as the Tokio runtime cannot spawn native threads in these sandboxed contexts.

How does num_workers affect memory usage during OCR processing?

Each concurrent OCR worker allocates memory for page bitmaps and recognition buffers. Higher num_workers values increase peak memory consumption linearly, as the semaphore in ocr_merge.rs allows more simultaneous allocations. If processing large documents or high-resolution scans, reducing num_workers prevents out-of-memory errors on systems with limited RAM.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →