How to Implement a Custom OCR Engine with the LiteParse OCR API
LiteParse exposes a trait-based OCR API that lets you inject any backend—whether a local library, HTTP service, or JavaScript callback—by implementing the OcrEngine trait and registering it via with_ocr_engine.
The LiteParse library provides a modular parsing pipeline that supports pluggable OCR backends for PDF text extraction. By implementing the OcrEngine trait defined in crates/liteparse/src/ocr/mod.rs, you can integrate custom recognition logic without modifying the core extraction code, whether you are building native Rust applications or WebAssembly browser bundles.
Understanding the OCR API Structure
At the heart of the LiteParse OCR system are three core components that define the contract between the parser and your custom engine.
The OcrEngine Trait
The OcrEngine trait in crates/liteparse/src/ocr/mod.rs (lines 28-43) specifies the interface every engine must satisfy. It requires two methods:
name()– Returns a static string identifier for the enginerecognize()– An async method that receives raw RGB bytes, image dimensions, and OCR options, returning a future that resolves to a vector ofOcrResultobjects
On native targets, the trait requires Send + Sync bounds to support multi-threaded execution, while the WebAssembly build relaxes these constraints to accommodate JavaScript's single-threaded nature.
Data Structures
The API uses two primary data structures also defined in crates/liteparse/src/ocr/mod.rs:
OcrResult(lines 8-16) – Contains the extractedtext, bounding box coordinates[x1, y1, x2, y2], and aconfidencescore between 0.0 and 1.0OcrOptions(lines 18-21) – Currently carries a singlelanguagestring field, which your engine can use to select appropriate models or remote endpoints
Implementing a Custom Engine
To implement a custom OCR engine with the LiteParse OCR API, you must define a struct that implements the OcrEngine trait and then inject it into the parser using LiteParse::with_ocr_engine.
Step-by-Step Guide
- Create your engine file under
crates/liteparse/src/ocr/(e.g.,my_engine.rs) - Implement the trait by defining
name()andrecognize()methods that returnPin<Box<dyn Future<...>>> - Normalize outputs to ensure bounding boxes use floating-point coordinates and confidence values fall within the 0.0 to 1.0 range
- Export the module by adding
pub mod my_engine;tocrates/liteparse/src/ocr/mod.rs - Inject the engine when constructing the parser using
Arc<dyn OcrEngine>
The parser's downstream pipeline (ocr_merge::ocr_and_merge_rendered) automatically calls your engine's recognize method for each rendered page requiring OCR, merging the results with native PDF text before spatial grid projection.
Code Examples
Example 1: Minimal Dummy Engine
This pass-through implementation demonstrates the trait structure and is useful for testing the integration without external dependencies.
// File: crates/liteparse/src/ocr/dummy.rs
use super::{OcrEngine, OcrOptions, OcrResult};
use std::pin::Pin;
use futures::future::BoxFuture;
/// A trivial engine that returns a single word containing the language name.
pub struct DummyEngine;
impl DummyEngine {
pub fn new() -> Self { Self }
}
impl OcrEngine for DummyEngine {
fn name(&self) -> &str { "dummy" }
fn recognize<'a, 'b: 'a, 'c: 'a>(
&'a self,
_image_data: &'c [u8],
_width: u32,
_height: u32,
options: &'b OcrOptions,
) -> Pin<BoxFuture<'a, Result<Vec<OcrResult>, Box<dyn std::error::Error + Send + Sync>>>> {
Box::pin(async move {
Ok(vec![OcrResult {
text: format!("lang={}", options.language),
bbox: [0.0, 0.0, 10.0, 10.0],
confidence: 0.9,
}])
})
}
}
Usage:
use liteparse::parser::LiteParse;
use liteparse::ocr::dummy::DummyEngine;
use std::sync::Arc;
let cfg = liteparse::config::LiteParseConfig::default();
let parser = LiteParse::new(cfg)
.with_ocr_engine(Arc::new(DummyEngine::new()));
let pdf_bytes = std::fs::read("sample.pdf")?;
let result = parser.parse_input(liteparse::types::PdfInput::Bytes(pdf_bytes)).await?;
Example 2: HTTP-Based OCR Service
This implementation converts RGB bytes to PNG and sends them to a remote OCR endpoint, parsing the JSON response into OcrResult objects.
// File: crates/liteparse/src/ocr/my_http.rs
use super::{OcrEngine, OcrOptions, OcrResult};
use std::{pin::Pin, future::Future};
use reqwest::{Client, multipart::{Form, Part}};
use image::{ImageBuffer, ImageFormat, RgbImage};
use std::io::Cursor;
pub struct MyHttpEngine {
name: String,
server_url: String,
client: Client,
api_key: Option<String>,
}
impl MyHttpEngine {
pub fn new(url: impl Into<String>, api_key: Option<String>) -> Self {
Self {
name: "my-http-ocr".into(),
server_url: url.into(),
client: Client::new(),
api_key,
}
}
}
impl OcrEngine for MyHttpEngine {
fn name(&self) -> &str { &self.name }
fn recognize<'a, 'b: 'a, 'c: 'a>(
&'a self,
image_data: &'c [u8],
width: u32,
height: u32,
options: &'b OcrOptions,
) -> Pin<Box<dyn Future<Output = Result<Vec<OcrResult>, Box<dyn std::error::Error + Send + Sync>>> + Send + 'a>> {
Box::pin(async move {
let img: RgbImage = ImageBuffer::from_raw(width, height, image_data.to_vec())
.ok_or("failed to create image buffer")?;
let mut png = Vec::new();
img.write_to(&mut Cursor::new(&mut png), ImageFormat::Png)?;
let mut form = Form::new()
.part("file", Part::bytes(png).file_name("page.png").mime_str("image/png")?)
.text("language", options.language.clone());
if let Some(ref key) = self.api_key {
form = form.text("api_key", key.clone());
}
let resp = self.client
.post(&self.server_url)
.multipart(form)
.send()
.await?
.json::<serde_json::Value>()
.await?;
let results = resp["results"]
.as_array()
.ok_or("malformed OCR response")?
.iter()
.map(|item| OcrResult {
text: item["text"].as_str().unwrap_or("").to_string(),
bbox: [
item["bbox"][0].as_f64().unwrap_or(0.0) as f32,
item["bbox"][1].as_f64().unwrap_or(0.0) as f32,
item["bbox"][2].as_f64().unwrap_or(0.0) as f32,
item["bbox"][3].as_f64().unwrap_or(0.0) as f32,
],
confidence: item["confidence"].as_f64().unwrap_or(0.0) as f32,
})
.collect();
Ok(results)
})
}
}
Wiring the custom HTTP engine:
use liteparse::parser::LiteParse;
use liteparse::ocr::my_http::MyHttpEngine;
use std::sync::Arc;
let cfg = liteparse::config::LiteParseConfig {
ocr_enabled: true,
..Default::default()
};
let my_engine = MyHttpEngine::new(
"https://my-ocr-service.example.com/ocr".to_string(),
Some("my-secret-api-key".to_string()),
);
let parser = LiteParse::new(cfg).with_ocr_engine(Arc::new(my_engine));
Example 3: JavaScript Bridge for WebAssembly
When consuming LiteParse in the browser via liteparse-wasm, you can supply a JavaScript object implementing an async recognize method. The JsOcrEngine wrapper in crates/liteparse-wasm/src/lib.rs (lines 63-71) automatically adapts this object to the Rust trait.
import { LiteParse } from "liteparse-wasm";
const ocrEngine = {
async recognize(imageData, width, height, language) {
// imageData is a Uint8Array of raw RGB bytes
const response = await fetch("https://my-ocr-service/api/ocr", {
method: "POST",
body: makeFormData(imageData, width, height, language),
});
const json = await response.json();
return json.results; // Array of { text, bbox, confidence }
},
};
const parser = await LiteParse.new({
ocrEnabled: true,
ocrEngine, // Injected JavaScript callback
ocrLanguage: "eng",
});
const pdfBytes = await fetch("/files/report.pdf").then(r => r.arrayBuffer());
const result = await parser.parse(new Uint8Array(pdfBytes));
console.log(result.text);
Engine Registration and Configuration
The LiteParse struct in crates/liteparse/src/parser.rs (lines 52-58) provides the with_ocr_engine method for overriding the default engine selection logic. This method accepts an Arc<dyn OcrEngine> and stores it for use during the parsing pipeline.
When the parser encounters pages requiring OCR, it passes the OcrOptions struct—populated from configuration values like --ocr-language—to your engine's recognize method. This allows your implementation to dynamically adjust behavior based on user-specified language codes or other extended options.
Summary
- Implement
OcrEngine– Definename()andrecognize()in a new module undercrates/liteparse/src/ocr/ - Handle async correctly – Return
Pin<Box<dyn Future>>usingBox::pin(async move { ... }) - Normalize outputs – Ensure confidence values are 0.0-1.0 and bounding boxes use
[x1, y1, x2, y2]format - Thread safety – Native engines must be
Send + Sync; WASM engines can use non-Send JavaScript values - Register with parser – Use
LiteParse::new(cfg).with_ocr_engine(Arc::new(your_engine)) - Reference implementations – Study
http_simple.rsandtesseract.rsin theocrmodule for production patterns
Frequently Asked Questions
What is the minimum required implementation for the OcrEngine trait?
You must implement two methods: name() returning a &str identifier, and recognize() accepting image bytes, dimensions, and options, returning a future that resolves to Vec<OcrResult>. The recognize method must be async and return Pin<Box<dyn Future>> to satisfy the trait bounds required by the parsing pipeline.
How do I pass configuration options to my custom OCR engine?
The OcrOptions struct passed to recognize contains a language field (and can be extended with additional fields). Access options.language inside your implementation to determine which model to load or which API endpoint to query. For additional static configuration, add fields to your engine's struct during initialization.
Can I use a JavaScript-only OCR service with LiteParse?
Yes. When using the liteparse-wasm package, provide a JavaScript object with an async recognize(imageData, width, height, language) method to the LiteParse.new() constructor. The JsOcrEngine wrapper in crates/liteparse-wasm/src/lib.rs automatically bridges this JavaScript callback to the Rust OcrEngine trait without requiring you to write Rust code.
Are there thread-safety requirements for custom engines?
On native targets (non-WASM), your engine must be Send + Sync because the parser may process pages concurrently. The future returned by recognize must also be Send. On WebAssembly, these bounds are relaxed because JavaScript values are not Send, allowing you to safely use browser APIs and WebWorkers within your JS callback.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →