How to Optimize OCR Settings in LiteParse for Better Performance: 7 Proven Tuning Methods
Tune LiteParse OCR performance by adjusting dpi, num_workers, and ocr_language in LiteParseConfig, disabling OCR with --no-ocr when native text exists, or offloading recognition to an external HTTP server via --ocr-server-url.
LiteParse intelligently runs OCR only on document regions that lack native text, such as scanned images embedded in PDFs. Because the OCR pipeline is configurable through both the CLI and the LiteParseConfig struct, learning how to optimize OCR settings in LiteParse can dramatically reduce parsing time without sacrificing accuracy.
Core OCR Configuration Options
In crates/liteparse/src/config.rs, the LiteParseConfig struct groups every OCR-related runtime setting. The fields that most directly impact parsing speed include:
ocr_enabled(bool, defaulttrue) — Set tofalseor pass--no-ocrto bypass the entire OCR pipeline when source documents already contain selectable text.ocr_language(String, default"eng") — Restricts Tesseract to a single language model; loading unnecessary models increases memory usage and startup latency.ocr_server_url(Option<String>, defaultNone) — Routes image bytes to a remote HTTP OCR service instead of local Tesseract, freeing the parser from CPU-bound recognition.tessdata_path(Option<String>) — Overrides theTESSDATA_PREFIXenvironment variable to point at a lean directory containing only required.traineddatafiles.dpi(f32, default150.0) — Controls the resolution of rasterized page images inrender.rs; lower values produce smaller images and faster OCR at the cost of fine detail.num_workers(usize, defaultCPU cores - 1) — Sets the number of pages processed concurrently by the OCR engine.preserve_very_small_text(bool, defaultfalse) — When disabled, the merge step inocr_merge.rsdiscards low-resolution OCR results, speeding up layout reconstruction.
How OCR Settings Flow Through the Pipeline
The pipeline begins in src/main.rs, where CLI flags such as --ocr-language and --dpi hydrate a LiteParseConfig instance. During document processing, parser.rs reads config.ocr_enabled to decide whether to invoke the local ocr::tesseract backend or the remote ocr::http_simple client. Both backends implement the OcrEngine trait defined in crates/liteparse/src/ocr/mod.rs. The rendering step in crates/liteparse/src/render.rs rasterizes each page at config.dpi and feeds the resulting PNG bytes into the active engine. Finally, crates/liteparse/src/ocr_merge.rs blends OCR output with native PDF text, respecting preserve_very_small_text to filter out low-resolution noise.
Seven Methods to Optimize OCR Performance in LiteParse
Reduce DPI — The render.rs module rasterizes pages at the configured dpi before sending them to the engine. Dropping the default of 150.0 to 100.0–120.0 for standard text documents can reduce OCR time by roughly 30% without materially hurting accuracy.
Limit Target Pages — Use --target-pages or LiteParseConfig::target_pages to constrain parsing to specific pages. Because the parser still renders all pages if OCR is enabled, restricting the scope shrinks the total OCR workload.
Disable OCR for Native Text PDFs — Many PDFs already contain embedded text. Passing --no-ocr or setting ocr_enabled: false skips the entire OCR pipeline, making this the fastest option when scans are absent.
Select a Single Language Model — Tesseract loads .traineddata at startup. Keeping ocr_language set to "eng" for English documents avoids the CPU and memory overhead of loading multiple models.
Offload to an External OCR Server — Supplying --ocr-server-url delegates recognition to a remote service via ocr/http_simple.rs. This decouples parsing from OCR latency and allows the server to run GPU-accelerated models independently.
Tune Concurrent Workers — The num_workers field controls parallel page processing. For CPU-bound local Tesseract, matching this value to the number of physical cores typically yields optimal throughput; for HTTP-based engines, a higher count can mask network latency.
Trim the Tessdata Directory — Pointing tessdata_path to a directory that contains only the required .traineddata files minimizes file-system lookup time and reduces memory pressure during model loading.
Practical Code Examples
CLI Example: Lower DPI and Restrict Workers
# Parse only pages 1-10, use 100 dpi images, and limit to 2 concurrent OCR workers
lit parse document.pdf \
--target-pages "1-10" \
--dpi 100 \
--num-workers 2 \
-o out.json
Rust API: Custom Config for a CPU-Limited Environment
use liteparse::LiteParseConfig;
use liteparse::OutputFormat;
use liteparse::LiteParse; // assume the library entry point
fn main() -> Result<(), Box<dyn std::error::Error>> {
let config = LiteParseConfig {
ocr_language: "eng".into(),
ocr_enabled: true,
ocr_server_url: None, // use built-in Tesseract
tessdata_path: Some("./tessdata".into()),
max_pages: 500,
target_pages: Some("1-20".into()),
dpi: 120.0, // lower DPI for speed
output_format: OutputFormat::Json,
preserve_very_small_text: false,
password: None,
quiet: false,
num_workers: 4, // match number of cores
};
let parser = LiteParse::new(config);
let result = parser.parse_path("sample.pdf")?;
println!("{}", serde_json::to_string_pretty(&result)?);
Ok(())
}
Python Wrapper: Disable OCR for Text-Heavy PDFs
from liteparse import LiteParse, LiteParseConfig, OutputFormat
cfg = LiteParseConfig(
ocr_enabled=False, # skip OCR entirely
output_format=OutputFormat.JSON,
)
parser = LiteParse(cfg)
out = parser.parse_path("text_only.pdf")
print(out.json())
Key Source Files for OCR Customization
crates/liteparse/src/config.rs— HoldsLiteParseConfigand all OCR-related runtime fields.crates/liteparse/src/ocr/mod.rs— Defines theOcrEnginetrait and shared structures such asOcrResultandOcrOptions.crates/liteparse/src/ocr/tesseract.rs— Feature-gated Tesseract implementation that passes language and image data to the native library.crates/liteparse/src/ocr/http_simple.rs— Minimal HTTP client that forwards image bytes to a remote OCR endpoint.crates/liteparse/src/render.rs— Generates raster images at the configured DPI for the OCR engine.crates/liteparse/src/ocr_merge.rs— Merges OCR results with native PDF text and applies thepreserve_very_small_textfilter.
Summary
- LiteParse skips OCR on pages that already contain native text, but rasterizes images at
config.dpiwhen OCR is required. - Lowering
dpifrom150.0to100.0–120.0is the fastest way to reduce local OCR latency. - Set
ocr_enabled: falseor use--no-ocrto eliminate the pipeline entirely for text-native PDFs. - Match
num_workersto physical CPU cores for Tesseract, or increase it to hide latency for HTTP-based engines. - Point
tessdata_pathto a minimal directory and restrictocr_languageto a single model to cut startup overhead. - Use
--ocr-server-urlto offload recognition and scale parsing independently from OCR resources.
Frequently Asked Questions
Can I disable OCR completely in LiteParse if my PDFs already have selectable text?
Yes. Set ocr_enabled: false in LiteParseConfig or pass the --no-ocr CLI flag. This causes parser.rs to bypass both the rendering and recognition stages, returning native text immediately and dramatically improving parse speed.
What is the optimal DPI setting for fast OCR in LiteParse?
The default is 150.0, but most standard text documents parse accurately at 100.0 to 120.0 DPI. Because render.rs rasterizes every page image at this resolution, lowering DPI reduces memory pressure and can cut OCR time by approximately 30%.
How does LiteParse handle OCR for documents with mixed languages?
LiteParse uses Tesseract’s language model specified by ocr_language in config.rs. If you load multiple languages, Tesseract initializes every corresponding .traineddata file, which increases memory usage and CPU overhead. For best performance, specify only the dominant language of your document.
Is it possible to run OCR on a separate machine while using LiteParse locally?
Yes. Configure ocr_server_url with the address of an HTTP service that implements LiteParse’s JSON API. The ocr/http_simple.rs client will forward PNG bytes to the remote server, allowing the local parser to continue instantly while the dedicated machine handles recognition.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →