# How to Implement a Custom OCR Engine with the LiteParse OCR API

> Learn to implement a custom OCR engine with the LiteParse OCR API. Inject any backend by implementing the OcrEngine trait and boost your application's OCR capabilities.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-07

---

**LiteParse exposes a trait-based OCR API that lets you inject any backend—whether a local library, HTTP service, or JavaScript callback—by implementing the `OcrEngine` trait and registering it via `with_ocr_engine`.**

The LiteParse library provides a modular parsing pipeline that supports pluggable OCR backends for PDF text extraction. By implementing the `OcrEngine` trait defined in [`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs), you can integrate custom recognition logic without modifying the core extraction code, whether you are building native Rust applications or WebAssembly browser bundles.

## Understanding the OCR API Structure

At the heart of the LiteParse OCR system are three core components that define the contract between the parser and your custom engine.

### The OcrEngine Trait

The `OcrEngine` trait in [`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs) (lines 28-43) specifies the interface every engine must satisfy. It requires two methods:

- **`name()`** – Returns a static string identifier for the engine
- **`recognize()`** – An async method that receives raw RGB bytes, image dimensions, and OCR options, returning a future that resolves to a vector of `OcrResult` objects

On native targets, the trait requires `Send + Sync` bounds to support multi-threaded execution, while the WebAssembly build relaxes these constraints to accommodate JavaScript's single-threaded nature.

### Data Structures

The API uses two primary data structures also defined in [`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs):

- **`OcrResult`** (lines 8-16) – Contains the extracted `text`, bounding box coordinates `[x1, y1, x2, y2]`, and a `confidence` score between 0.0 and 1.0
- **`OcrOptions`** (lines 18-21) – Currently carries a single `language` string field, which your engine can use to select appropriate models or remote endpoints

## Implementing a Custom Engine

To implement a custom OCR engine with the LiteParse OCR API, you must define a struct that implements the `OcrEngine` trait and then inject it into the parser using `LiteParse::with_ocr_engine`.

### Step-by-Step Guide

1. **Create your engine file** under `crates/liteparse/src/ocr/` (e.g., [`my_engine.rs`](https://github.com/run-llama/liteparse/blob/main/my_engine.rs))
2. **Implement the trait** by defining `name()` and `recognize()` methods that return `Pin<Box<dyn Future<...>>>`
3. **Normalize outputs** to ensure bounding boxes use floating-point coordinates and confidence values fall within the 0.0 to 1.0 range
4. **Export the module** by adding `pub mod my_engine;` to [`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs)
5. **Inject the engine** when constructing the parser using `Arc<dyn OcrEngine>`

The parser's downstream pipeline (`ocr_merge::ocr_and_merge_rendered`) automatically calls your engine's `recognize` method for each rendered page requiring OCR, merging the results with native PDF text before spatial grid projection.

## Code Examples

### Example 1: Minimal Dummy Engine

This pass-through implementation demonstrates the trait structure and is useful for testing the integration without external dependencies.

```rust
// File: crates/liteparse/src/ocr/dummy.rs
use super::{OcrEngine, OcrOptions, OcrResult};
use std::pin::Pin;
use futures::future::BoxFuture;

/// A trivial engine that returns a single word containing the language name.
pub struct DummyEngine;

impl DummyEngine {
    pub fn new() -> Self { Self }
}

impl OcrEngine for DummyEngine {
    fn name(&self) -> &str { "dummy" }

    fn recognize<'a, 'b: 'a, 'c: 'a>(
        &'a self,
        _image_data: &'c [u8],
        _width: u32,
        _height: u32,
        options: &'b OcrOptions,
    ) -> Pin<BoxFuture<'a, Result<Vec<OcrResult>, Box<dyn std::error::Error + Send + Sync>>>> {
        Box::pin(async move {
            Ok(vec![OcrResult {
                text: format!("lang={}", options.language),
                bbox: [0.0, 0.0, 10.0, 10.0],
                confidence: 0.9,
            }])
        })
    }
}

```

**Usage:**

```rust
use liteparse::parser::LiteParse;
use liteparse::ocr::dummy::DummyEngine;
use std::sync::Arc;

let cfg = liteparse::config::LiteParseConfig::default();
let parser = LiteParse::new(cfg)
    .with_ocr_engine(Arc::new(DummyEngine::new()));

let pdf_bytes = std::fs::read("sample.pdf")?;
let result = parser.parse_input(liteparse::types::PdfInput::Bytes(pdf_bytes)).await?;

```

### Example 2: HTTP-Based OCR Service

This implementation converts RGB bytes to PNG and sends them to a remote OCR endpoint, parsing the JSON response into `OcrResult` objects.

```rust
// File: crates/liteparse/src/ocr/my_http.rs
use super::{OcrEngine, OcrOptions, OcrResult};
use std::{pin::Pin, future::Future};
use reqwest::{Client, multipart::{Form, Part}};
use image::{ImageBuffer, ImageFormat, RgbImage};
use std::io::Cursor;

pub struct MyHttpEngine {
    name: String,
    server_url: String,
    client: Client,
    api_key: Option<String>,
}

impl MyHttpEngine {
    pub fn new(url: impl Into<String>, api_key: Option<String>) -> Self {
        Self {
            name: "my-http-ocr".into(),
            server_url: url.into(),
            client: Client::new(),
            api_key,
        }
    }
}

impl OcrEngine for MyHttpEngine {
    fn name(&self) -> &str { &self.name }

    fn recognize<'a, 'b: 'a, 'c: 'a>(
        &'a self,
        image_data: &'c [u8],
        width: u32,
        height: u32,
        options: &'b OcrOptions,
    ) -> Pin<Box<dyn Future<Output = Result<Vec<OcrResult>, Box<dyn std::error::Error + Send + Sync>>> + Send + 'a>> {
        Box::pin(async move {
            let img: RgbImage = ImageBuffer::from_raw(width, height, image_data.to_vec())
                .ok_or("failed to create image buffer")?;
            let mut png = Vec::new();
            img.write_to(&mut Cursor::new(&mut png), ImageFormat::Png)?;

            let mut form = Form::new()
                .part("file", Part::bytes(png).file_name("page.png").mime_str("image/png")?)
                .text("language", options.language.clone());

            if let Some(ref key) = self.api_key {
                form = form.text("api_key", key.clone());
            }

            let resp = self.client
                .post(&self.server_url)
                .multipart(form)
                .send()
                .await?
                .json::<serde_json::Value>()
                .await?;

            let results = resp["results"]
                .as_array()
                .ok_or("malformed OCR response")?
                .iter()
                .map(|item| OcrResult {
                    text: item["text"].as_str().unwrap_or("").to_string(),
                    bbox: [
                        item["bbox"][0].as_f64().unwrap_or(0.0) as f32,
                        item["bbox"][1].as_f64().unwrap_or(0.0) as f32,
                        item["bbox"][2].as_f64().unwrap_or(0.0) as f32,
                        item["bbox"][3].as_f64().unwrap_or(0.0) as f32,
                    ],
                    confidence: item["confidence"].as_f64().unwrap_or(0.0) as f32,
                })
                .collect();

            Ok(results)
        })
    }
}

```

**Wiring the custom HTTP engine:**

```rust
use liteparse::parser::LiteParse;
use liteparse::ocr::my_http::MyHttpEngine;
use std::sync::Arc;

let cfg = liteparse::config::LiteParseConfig {
    ocr_enabled: true,
    ..Default::default()
};

let my_engine = MyHttpEngine::new(
    "https://my-ocr-service.example.com/ocr".to_string(),
    Some("my-secret-api-key".to_string()),
);

let parser = LiteParse::new(cfg).with_ocr_engine(Arc::new(my_engine));

```

### Example 3: JavaScript Bridge for WebAssembly

When consuming LiteParse in the browser via `liteparse-wasm`, you can supply a JavaScript object implementing an async `recognize` method. The `JsOcrEngine` wrapper in [`crates/liteparse-wasm/src/lib.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse-wasm/src/lib.rs) (lines 63-71) automatically adapts this object to the Rust trait.

```javascript
import { LiteParse } from "liteparse-wasm";

const ocrEngine = {
  async recognize(imageData, width, height, language) {
    // imageData is a Uint8Array of raw RGB bytes
    const response = await fetch("https://my-ocr-service/api/ocr", {
      method: "POST",
      body: makeFormData(imageData, width, height, language),
    });
    const json = await response.json();
    return json.results; // Array of { text, bbox, confidence }
  },
};

const parser = await LiteParse.new({
  ocrEnabled: true,
  ocrEngine,          // Injected JavaScript callback
  ocrLanguage: "eng",
});

const pdfBytes = await fetch("/files/report.pdf").then(r => r.arrayBuffer());
const result = await parser.parse(new Uint8Array(pdfBytes));
console.log(result.text);

```

## Engine Registration and Configuration

The `LiteParse` struct in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) (lines 52-58) provides the `with_ocr_engine` method for overriding the default engine selection logic. This method accepts an `Arc<dyn OcrEngine>` and stores it for use during the parsing pipeline.

When the parser encounters pages requiring OCR, it passes the `OcrOptions` struct—populated from configuration values like `--ocr-language`—to your engine's `recognize` method. This allows your implementation to dynamically adjust behavior based on user-specified language codes or other extended options.

## Summary

- **Implement `OcrEngine`** – Define `name()` and `recognize()` in a new module under `crates/liteparse/src/ocr/`
- **Handle async correctly** – Return `Pin<Box<dyn Future>>` using `Box::pin(async move { ... })`
- **Normalize outputs** – Ensure confidence values are 0.0-1.0 and bounding boxes use `[x1, y1, x2, y2]` format
- **Thread safety** – Native engines must be `Send + Sync`; WASM engines can use non-Send JavaScript values
- **Register with parser** – Use `LiteParse::new(cfg).with_ocr_engine(Arc::new(your_engine))`
- **Reference implementations** – Study [`http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/http_simple.rs) and [`tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/tesseract.rs) in the `ocr` module for production patterns

## Frequently Asked Questions

### What is the minimum required implementation for the OcrEngine trait?

You must implement two methods: `name()` returning a `&str` identifier, and `recognize()` accepting image bytes, dimensions, and options, returning a future that resolves to `Vec<OcrResult>`. The `recognize` method must be async and return `Pin<Box<dyn Future>>` to satisfy the trait bounds required by the parsing pipeline.

### How do I pass configuration options to my custom OCR engine?

The `OcrOptions` struct passed to `recognize` contains a `language` field (and can be extended with additional fields). Access `options.language` inside your implementation to determine which model to load or which API endpoint to query. For additional static configuration, add fields to your engine's struct during initialization.

### Can I use a JavaScript-only OCR service with LiteParse?

Yes. When using the `liteparse-wasm` package, provide a JavaScript object with an async `recognize(imageData, width, height, language)` method to the `LiteParse.new()` constructor. The `JsOcrEngine` wrapper in [`crates/liteparse-wasm/src/lib.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse-wasm/src/lib.rs) automatically bridges this JavaScript callback to the Rust `OcrEngine` trait without requiring you to write Rust code.

### Are there thread-safety requirements for custom engines?

On native targets (non-WASM), your engine must be `Send + Sync` because the parser may process pages concurrently. The future returned by `recognize` must also be `Send`. On WebAssembly, these bounds are relaxed because JavaScript values are not `Send`, allowing you to safely use browser APIs and WebWorkers within your JS callback.