Can LiteParse Use External HTTP OCR Servers Instead of Tesseract?

Yes, LiteParse can use external HTTP OCR servers instead of Tesseract by configuring the ocr_server_url option, which switches the backend from the local Tesseract binary to a remote HTTP service.

LiteParse, the Rust-based document parsing library from the run-llama organization, provides a pluggable OCR architecture that decouples text recognition from the core parsing pipeline. While it ships with a default Tesseract integration for local processing, the codebase is designed around the OcrEngine trait, allowing seamless substitution with any HTTP-based OCR service that conforms to the expected JSON API.

How LiteParse Configures External OCR Servers

The switch between local and remote OCR happens at initialization time through a simple configuration field.

Configuration via LiteParseConfig

All OCR-related settings are centralized in LiteParseConfig, defined in [crates/liteparse/src/config.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs). The critical field for external services is ocr_server_url, an optional string that, when present, triggers the HTTP engine instantiation.

let cfg = LiteParseConfig {
    ocr_enabled: true,
    ocr_server_url: Some("http://my-ocr-service.example.com/ocr".into()),
    ..Default::default()
};

Engine Selection Logic

During parser construction in [crates/liteparse/src/parser.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs), LiteParse evaluates config.ocr_server_url to determine which engine implementation to load:

Because both engines implement the same OcrEngine trait defined in [ocr/mod.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs), downstream components—such as projection handling, OCR merging, and output formatting—remain agnostic to the backend choice.

Technical Implementation of HTTP OCR

The HTTP OCR engine follows a predictable request-response pattern designed for stateless external services.

The HttpOcrEngine Structure

HttpOcrEngine is a lightweight wrapper around an HTTP client that implements the OcrEngine trait. Unlike the Tesseract backend, which requires linking to native libraries, this engine communicates via standard HTTP POST requests, making it compatible with containerized microservices or cloud OCR APIs.

Request/Response Flow

When recognize is called on HttpOcrEngine, the following sequence occurs:

  1. Image Encoding: The raw RGB image buffer is encoded as a PNG.

  2. Multipart POST: The engine submits a POST request to the configured URL with:

    • A file field containing the PNG bytes
    • A language field specifying the OCR language code
  3. JSON Deserialization: The server must return a JSON object with the following structure:

    {
      "results": [
        {
          "text": "extracted text",
          "bbox": [x1, y1, x2, y2],
          "confidence": 0.95
        }
      ]
    }
  4. Result Mapping: The response is parsed into Vec<OcrResult>, matching the exact return type used by the Tesseract engine.

Practical Usage Examples

You can leverage external OCR servers via command-line interface or programmatic Rust code.

CLI Usage

Run LiteParse from the terminal with the --ocr-server-url flag to route all recognition tasks to your HTTP endpoint:

liteparse input.pdf \
  --ocr-enabled \
  --ocr-server-url http://my-ocr-service.example.com/ocr \
  --output-format json

Programmatic Configuration

Integrate external OCR into your Rust application by constructing a LiteParse instance with a custom configuration:

use liteparse::parser::LiteParse;
use liteparse::config::LiteParseConfig;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let cfg = LiteParseConfig {
        ocr_enabled: true,
        ocr_server_url: Some("http://my-ocr-service.example.com/ocr".into()),
        ..Default::default()
    };

    let parser = LiteParse::new(cfg)?;
    let result = parser.parse_path("sample.pdf")?;
    println!("{}", serde_json::to_string_pretty(&result)?);
    Ok(())
}

Minimal HTTP OCR Server Implementation

To build a compatible OCR server, implement the endpoint contract described in the repository's OCR_API_SPEC.md. Below is a minimal Python Flask example using Tesseract as the backend processor:

from flask import Flask, request, jsonify
from PIL import Image
import pytesseract

app = Flask(__name__)

@app.post("/ocr")
def ocr():
    img = Image.open(request.files["file"])
    lang = request.form.get("language", "eng")
    data = pytesseract.image_to_data(img, lang=lang, output_type=pytesseract.Output.DICT)
    results = [
        {
            "text": txt,
            "bbox": [float(x), float(y), float(x + w), float(y + h)],
            "confidence": float(conf) / 100.0,
        }
        for txt, x, y, w, h, conf in zip(
            data["text"], data["left"], data["top"], data["width"], data["height"], data["conf"]
        )
        if txt.strip()
    ]
    return jsonify({"results": results})

if __name__ == "__main__":
    app.run(host="0.0.0.0", port=8000)

Once deployed, point LiteParse to http://localhost:8000/ocr to process documents through this custom service.

Summary

  • Trait-based architecture: LiteParse uses the OcrEngine trait in ocr/mod.rs to abstract OCR implementations, allowing hot-swapping between backends.
  • Simple configuration: Set ocr_server_url in LiteParseConfig to automatically instantiate HttpOcrEngine instead of the default Tesseract engine.
  • Standard HTTP protocol: External servers receive PNG images via multipart POST and return JSON containing text, bounding boxes, and confidence scores.
  • Zero downstream impact: Parser components remain unchanged when switching between local Tesseract and remote HTTP OCR servers.

Frequently Asked Questions

What JSON format must external HTTP OCR servers return?

External servers must return a JSON object with a results array, where each element contains text (string), bbox (array of four floats), and confidence (float between 0 and 1). This format is identical to the internal representation used by LiteParse's Tesseract integration, ensuring consistent behavior across engines.

Can I use multiple OCR servers simultaneously with LiteParse?

No. The current implementation in parser.rs selects a single OcrEngine implementation during initialization based on the ocr_server_url configuration. To distribute load across multiple servers, you must implement a reverse proxy or load balancer in front of your OCR services, then point LiteParse to that aggregated endpoint.

Is Tesseract required to be installed when using external HTTP OCR servers?

No. When ocr_server_url is configured, LiteParse exclusively uses HttpOcrEngine and does not invoke local Tesseract binaries. You can compile or run LiteParse without the tesseract feature flag enabled, reducing binary size and eliminating the native dependency entirely.

How does LiteParse handle HTTP OCR failures?

The HttpOcrEngine::recognize implementation propagates HTTP errors as standard Rust Result types. If the external server returns a non-2xx status code or malformed JSON, the error bubbles up through the parser's parse_path or equivalent methods, allowing your application to handle network timeouts or service unavailability through standard error handling patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →