How to Implement a Custom OCR Engine with the LiteParse OCR API

LiteParse exposes a trait-based OCR API that lets you inject any backend—whether a local library, HTTP service, or JavaScript callback—by implementing the OcrEngine trait and registering it via with_ocr_engine.

The LiteParse library provides a modular parsing pipeline that supports pluggable OCR backends for PDF text extraction. By implementing the OcrEngine trait defined in crates/liteparse/src/ocr/mod.rs, you can integrate custom recognition logic without modifying the core extraction code, whether you are building native Rust applications or WebAssembly browser bundles.

Understanding the OCR API Structure

At the heart of the LiteParse OCR system are three core components that define the contract between the parser and your custom engine.

The OcrEngine Trait

The OcrEngine trait in crates/liteparse/src/ocr/mod.rs (lines 28-43) specifies the interface every engine must satisfy. It requires two methods:

  • name() – Returns a static string identifier for the engine
  • recognize() – An async method that receives raw RGB bytes, image dimensions, and OCR options, returning a future that resolves to a vector of OcrResult objects

On native targets, the trait requires Send + Sync bounds to support multi-threaded execution, while the WebAssembly build relaxes these constraints to accommodate JavaScript's single-threaded nature.

Data Structures

The API uses two primary data structures also defined in crates/liteparse/src/ocr/mod.rs:

  • OcrResult (lines 8-16) – Contains the extracted text, bounding box coordinates [x1, y1, x2, y2], and a confidence score between 0.0 and 1.0
  • OcrOptions (lines 18-21) – Currently carries a single language string field, which your engine can use to select appropriate models or remote endpoints

Implementing a Custom Engine

To implement a custom OCR engine with the LiteParse OCR API, you must define a struct that implements the OcrEngine trait and then inject it into the parser using LiteParse::with_ocr_engine.

Step-by-Step Guide

  1. Create your engine file under crates/liteparse/src/ocr/ (e.g., my_engine.rs)
  2. Implement the trait by defining name() and recognize() methods that return Pin<Box<dyn Future<...>>>
  3. Normalize outputs to ensure bounding boxes use floating-point coordinates and confidence values fall within the 0.0 to 1.0 range
  4. Export the module by adding pub mod my_engine; to crates/liteparse/src/ocr/mod.rs
  5. Inject the engine when constructing the parser using Arc<dyn OcrEngine>

The parser's downstream pipeline (ocr_merge::ocr_and_merge_rendered) automatically calls your engine's recognize method for each rendered page requiring OCR, merging the results with native PDF text before spatial grid projection.

Code Examples

Example 1: Minimal Dummy Engine

This pass-through implementation demonstrates the trait structure and is useful for testing the integration without external dependencies.

// File: crates/liteparse/src/ocr/dummy.rs
use super::{OcrEngine, OcrOptions, OcrResult};
use std::pin::Pin;
use futures::future::BoxFuture;

/// A trivial engine that returns a single word containing the language name.
pub struct DummyEngine;

impl DummyEngine {
    pub fn new() -> Self { Self }
}

impl OcrEngine for DummyEngine {
    fn name(&self) -> &str { "dummy" }

    fn recognize<'a, 'b: 'a, 'c: 'a>(
        &'a self,
        _image_data: &'c [u8],
        _width: u32,
        _height: u32,
        options: &'b OcrOptions,
    ) -> Pin<BoxFuture<'a, Result<Vec<OcrResult>, Box<dyn std::error::Error + Send + Sync>>>> {
        Box::pin(async move {
            Ok(vec![OcrResult {
                text: format!("lang={}", options.language),
                bbox: [0.0, 0.0, 10.0, 10.0],
                confidence: 0.9,
            }])
        })
    }
}

Usage:

use liteparse::parser::LiteParse;
use liteparse::ocr::dummy::DummyEngine;
use std::sync::Arc;

let cfg = liteparse::config::LiteParseConfig::default();
let parser = LiteParse::new(cfg)
    .with_ocr_engine(Arc::new(DummyEngine::new()));

let pdf_bytes = std::fs::read("sample.pdf")?;
let result = parser.parse_input(liteparse::types::PdfInput::Bytes(pdf_bytes)).await?;

Example 2: HTTP-Based OCR Service

This implementation converts RGB bytes to PNG and sends them to a remote OCR endpoint, parsing the JSON response into OcrResult objects.

// File: crates/liteparse/src/ocr/my_http.rs
use super::{OcrEngine, OcrOptions, OcrResult};
use std::{pin::Pin, future::Future};
use reqwest::{Client, multipart::{Form, Part}};
use image::{ImageBuffer, ImageFormat, RgbImage};
use std::io::Cursor;

pub struct MyHttpEngine {
    name: String,
    server_url: String,
    client: Client,
    api_key: Option<String>,
}

impl MyHttpEngine {
    pub fn new(url: impl Into<String>, api_key: Option<String>) -> Self {
        Self {
            name: "my-http-ocr".into(),
            server_url: url.into(),
            client: Client::new(),
            api_key,
        }
    }
}

impl OcrEngine for MyHttpEngine {
    fn name(&self) -> &str { &self.name }

    fn recognize<'a, 'b: 'a, 'c: 'a>(
        &'a self,
        image_data: &'c [u8],
        width: u32,
        height: u32,
        options: &'b OcrOptions,
    ) -> Pin<Box<dyn Future<Output = Result<Vec<OcrResult>, Box<dyn std::error::Error + Send + Sync>>> + Send + 'a>> {
        Box::pin(async move {
            let img: RgbImage = ImageBuffer::from_raw(width, height, image_data.to_vec())
                .ok_or("failed to create image buffer")?;
            let mut png = Vec::new();
            img.write_to(&mut Cursor::new(&mut png), ImageFormat::Png)?;

            let mut form = Form::new()
                .part("file", Part::bytes(png).file_name("page.png").mime_str("image/png")?)
                .text("language", options.language.clone());

            if let Some(ref key) = self.api_key {
                form = form.text("api_key", key.clone());
            }

            let resp = self.client
                .post(&self.server_url)
                .multipart(form)
                .send()
                .await?
                .json::<serde_json::Value>()
                .await?;

            let results = resp["results"]
                .as_array()
                .ok_or("malformed OCR response")?
                .iter()
                .map(|item| OcrResult {
                    text: item["text"].as_str().unwrap_or("").to_string(),
                    bbox: [
                        item["bbox"][0].as_f64().unwrap_or(0.0) as f32,
                        item["bbox"][1].as_f64().unwrap_or(0.0) as f32,
                        item["bbox"][2].as_f64().unwrap_or(0.0) as f32,
                        item["bbox"][3].as_f64().unwrap_or(0.0) as f32,
                    ],
                    confidence: item["confidence"].as_f64().unwrap_or(0.0) as f32,
                })
                .collect();

            Ok(results)
        })
    }
}

Wiring the custom HTTP engine:

use liteparse::parser::LiteParse;
use liteparse::ocr::my_http::MyHttpEngine;
use std::sync::Arc;

let cfg = liteparse::config::LiteParseConfig {
    ocr_enabled: true,
    ..Default::default()
};

let my_engine = MyHttpEngine::new(
    "https://my-ocr-service.example.com/ocr".to_string(),
    Some("my-secret-api-key".to_string()),
);

let parser = LiteParse::new(cfg).with_ocr_engine(Arc::new(my_engine));

Example 3: JavaScript Bridge for WebAssembly

When consuming LiteParse in the browser via liteparse-wasm, you can supply a JavaScript object implementing an async recognize method. The JsOcrEngine wrapper in crates/liteparse-wasm/src/lib.rs (lines 63-71) automatically adapts this object to the Rust trait.

import { LiteParse } from "liteparse-wasm";

const ocrEngine = {
  async recognize(imageData, width, height, language) {
    // imageData is a Uint8Array of raw RGB bytes
    const response = await fetch("https://my-ocr-service/api/ocr", {
      method: "POST",
      body: makeFormData(imageData, width, height, language),
    });
    const json = await response.json();
    return json.results; // Array of { text, bbox, confidence }
  },
};

const parser = await LiteParse.new({
  ocrEnabled: true,
  ocrEngine,          // Injected JavaScript callback
  ocrLanguage: "eng",
});

const pdfBytes = await fetch("/files/report.pdf").then(r => r.arrayBuffer());
const result = await parser.parse(new Uint8Array(pdfBytes));
console.log(result.text);

Engine Registration and Configuration

The LiteParse struct in crates/liteparse/src/parser.rs (lines 52-58) provides the with_ocr_engine method for overriding the default engine selection logic. This method accepts an Arc<dyn OcrEngine> and stores it for use during the parsing pipeline.

When the parser encounters pages requiring OCR, it passes the OcrOptions struct—populated from configuration values like --ocr-language—to your engine's recognize method. This allows your implementation to dynamically adjust behavior based on user-specified language codes or other extended options.

Summary

  • Implement OcrEngine – Define name() and recognize() in a new module under crates/liteparse/src/ocr/
  • Handle async correctly – Return Pin<Box<dyn Future>> using Box::pin(async move { ... })
  • Normalize outputs – Ensure confidence values are 0.0-1.0 and bounding boxes use [x1, y1, x2, y2] format
  • Thread safety – Native engines must be Send + Sync; WASM engines can use non-Send JavaScript values
  • Register with parser – Use LiteParse::new(cfg).with_ocr_engine(Arc::new(your_engine))
  • Reference implementations – Study http_simple.rs and tesseract.rs in the ocr module for production patterns

Frequently Asked Questions

What is the minimum required implementation for the OcrEngine trait?

You must implement two methods: name() returning a &str identifier, and recognize() accepting image bytes, dimensions, and options, returning a future that resolves to Vec<OcrResult>. The recognize method must be async and return Pin<Box<dyn Future>> to satisfy the trait bounds required by the parsing pipeline.

How do I pass configuration options to my custom OCR engine?

The OcrOptions struct passed to recognize contains a language field (and can be extended with additional fields). Access options.language inside your implementation to determine which model to load or which API endpoint to query. For additional static configuration, add fields to your engine's struct during initialization.

Can I use a JavaScript-only OCR service with LiteParse?

Yes. When using the liteparse-wasm package, provide a JavaScript object with an async recognize(imageData, width, height, language) method to the LiteParse.new() constructor. The JsOcrEngine wrapper in crates/liteparse-wasm/src/lib.rs automatically bridges this JavaScript callback to the Rust OcrEngine trait without requiring you to write Rust code.

Are there thread-safety requirements for custom engines?

On native targets (non-WASM), your engine must be Send + Sync because the parser may process pages concurrently. The future returned by recognize must also be Send. On WebAssembly, these bounds are relaxed because JavaScript values are not Send, allowing you to safely use browser APIs and WebWorkers within your JS callback.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →