How to Implement a Custom HTTP OCR Server Compatible with LiteParse
You can implement a custom HTTP OCR server compatible with LiteParse by exposing a POST /ocr endpoint that accepts multipart/form-data image uploads and returns JSON containing recognized text, bounding boxes, and confidence scores, then configuring LiteParse to use your server URL.
LiteParse is a Rust-based PDF parsing library from the run-llama/liteparse repository that supports pluggable OCR backends. When you require more control over text recognition than the built-in Tesseract engine provides, implementing a custom HTTP OCR server allows you to integrate specialized engines like EasyOCR, PaddleOCR, or proprietary solutions while maintaining seamless compatibility with LiteParse's document processing pipeline.
Understanding the LiteParse HTTP OCR Architecture
LiteParse delegates OCR operations to an external service when you configure an HTTP endpoint. Understanding how the library selects and communicates with this engine is essential for building a compatible server.
Configuration and Engine Selection
The OCR backend is determined by the ocr_server_url field in LiteParseConfig, defined in crates/liteparse/src/config.rs. When this URL is provided, LiteParse instantiates an HttpOcrEngine instead of the default local Tesseract engine.
In crates/liteparse/src/parser.rs, the LiteParse::parse_input method contains the selection logic:
if let Some(ref url) = self.config.ocr_server_url {
std::sync::Arc::new(HttpOcrEngine::new(url.clone()))
} else {
// fallback to Tesseract if the `tesseract` feature is enabled
}
This conditional check ensures that setting ocr_server_url automatically routes all image recognition tasks through the HTTP pathway.
The HTTP OCR Engine Implementation
The HttpOcrEngine struct, located in crates/liteparse/src/ocr/http_simple.rs, implements the OcrEngine trait defined in crates/liteparse/src/ocr/mod.rs. Its recognize method performs three critical operations:
- Converts the raw RGB page image into PNG format
- Constructs a multipart/form-data POST request to the configured server URL
- Deserializes the JSON response into
OcrResultobjects containing text, bounding boxes, and confidence values
This implementation expects your server to handle binary image data and respond with a specific JSON schema.
The LiteParse OCR API Specification
Any custom HTTP OCR server must conform to the LiteParse OCR API Specification documented in OCR_API_SPEC.md at the repository root. This contract defines the exact interface between LiteParse and your service.
Required Endpoint and Parameters
Your server must expose:
- Endpoint:
POST /ocr - Content-Type:
multipart/form-data - Parameters:
file: Binary image data (PNG or JPEG)language(optional): ISO language code (e.g., "en", "zh", "eng")
Response Format
The endpoint must return HTTP 200 with a JSON body structured as follows:
{
"results": [
{
"text": "recognized text",
"bbox": [x1, y1, x2, y2],
"confidence": 0.97
}
]
}
Each result object requires:
text: The recognized stringbbox: An array of four floats representing the bounding box coordinates[left, top, right, bottom]confidence: A float between 0.0 and 1.0 indicating recognition certainty
Implementing Your Custom HTTP OCR Server
Building a compliant server requires handling image decoding, processing through your chosen OCR engine, and formatting results according to the specification.
Minimal FastAPI Implementation
Below is a complete, minimal implementation using Python and FastAPI that satisfies LiteParse requirements:
from fastapi import FastAPI, File, Form, UploadFile
from pydantic import BaseModel
import your_ocr_engine # Replace with actual OCR library
app = FastAPI()
class OcrResponse(BaseModel):
results: list[dict]
@app.post("/ocr")
async def ocr_endpoint(
file: UploadFile = File(...),
language: str = Form(default="en")
) -> OcrResponse:
# Read image bytes from the request
image_bytes = await file.read()
# Process with your OCR engine
# This is pseudocode - implement according to your engine's API
raw_results = your_ocr_engine.process(image_bytes, lang=language)
# Format results to match LiteParse specification
formatted_results = [
{
"text": item.text,
"bbox": [item.x1, item.y1, item.x2, item.y2],
"confidence": item.confidence
}
for item in raw_results
]
return OcrResponse(results=formatted_results)
Run this server with uvicorn main:app --host 0.0.0.0 --port 9000.
Handling Image Processing
The HttpOcrEngine sends PNG-encoded images via multipart/form-data. Your server should:
- Accept the
filefield as binary data - Support common image formats (PNG is guaranteed from LiteParse, but JPEG handling adds flexibility)
- Respect the optional
languageparameter for multilingual documents - Return an empty
resultsarray for pages containing no text rather than HTTP error codes
Integrating Your Server with LiteParse
Once your server is running, configure LiteParse to use it instead of the local Tesseract engine.
Command Line Interface
Use the --ocr-server-url flag when parsing documents:
lit parse document.pdf --ocr-server-url http://localhost:9000/ocr
For multilingual documents, specify the language code:
lit parse document.pdf --ocr-server-url http://localhost:9000/ocr --ocr-language zh
Programmatic Integration (Rust)
In Rust applications using liteparse directly, set the ocr_server_url field in LiteParseConfig:
use liteparse::{LiteParse, LiteParseConfig};
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let cfg = LiteParseConfig {
ocr_enabled: true,
ocr_server_url: Some("http://localhost:9000/ocr".into()),
ocr_language: Some("en".into()),
..Default::default()
};
let parser = LiteParse::new(cfg);
let result = parser.parse_input(
liteparse::PdfInput::Path("document.pdf".into())
).await?;
println!("{}", result.text);
Ok(())
}
Programmatic Integration (Node.js/TypeScript)
When using the JavaScript bindings, pass the URL in the configuration object:
import { LiteParse } from 'liteparse';
const parser = new LiteParse({
ocrServerUrl: 'http://localhost:9000/ocr',
ocrLanguage: 'en',
});
const result = await parser.parse('document.pdf');
console.log(result.text);
Reference Implementations
The run-llama/liteparse repository provides two production-ready reference servers that demonstrate best practices for implementing the specification.
EasyOCR Server
Located at ocr/easyocr/server.py, this FastAPI application wraps the EasyOCR library:
cd ocr/easyocr
uv run server.py # Starts on http://0.0.0.0:8828
This implementation handles automatic GPU detection and supports all languages available in the EasyOCR model zoo.
PaddleOCR Server
Found at ocr/paddleocr/server.py, this server leverages PaddleOCR for high-performance text recognition:
cd ocr/paddleocr
uv run server.py # Starts on http://0.0.0.0:8829
You can enable GPU acceleration by modifying the use_gpu parameter in the PaddleOCRServer class initialization.
Both servers expose GET /health endpoints for load balancer health checks and implement detailed request logging compatible with the LiteParse HTTP OCR engine expectations.
Summary
Implementing a custom HTTP OCR server compatible with LiteParse requires adherence to a specific API contract while providing flexibility for your chosen OCR technology:
- Configuration: Set
ocr_server_urlinLiteParseConfigor use the--ocr-server-urlCLI flag to route OCR requests to your server - API Contract: Implement
POST /ocraccepting multipart/form-data withfileand optionallanguagefields, returning JSON withresultscontainingtext,bbox, andconfidence - Engine Selection: LiteParse automatically instantiates
HttpOcrEnginefromcrates/liteparse/src/ocr/http_simple.rswhen a URL is configured, converting PDF pages to PNG and sending them via HTTP - Reference Code: Study the EasyOCR (
ocr/easyocr/server.py) and PaddleOCR (ocr/paddleocr/server.py) implementations for production-ready patterns
Frequently Asked Questions
What image format does LiteParse send to the HTTP OCR server?
LiteParse converts PDF pages to PNG format before transmission. The HttpOcrEngine::recognize method in crates/liteparse/src/ocr/http_simple.rs handles the RGB to PNG encoding automatically, ensuring lossless quality for optimal text recognition accuracy.
Can I use GPU acceleration with a custom HTTP OCR server?
Yes. GPU acceleration is independent of LiteParse and depends entirely on your server implementation. The reference PaddleOCR server at ocr/paddleocr/server.py demonstrates GPU support by setting use_gpu=True in the PaddleOCR constructor, while the client configuration in LiteParse remains unchanged.
How does LiteParse handle OCR server failures or timeouts?
The HttpOcrEngine in crates/liteparse/src/ocr/http_simple.rs propagates HTTP errors as OCR failures. If your server returns non-200 status codes or timeouts occur, the parsing operation will fail with the corresponding error. Implement proper health checks using the GET /health endpoint pattern shown in the reference servers to ensure high availability.
Is the language parameter required in the OCR API specification?
No, the language parameter is optional. According to the specification in OCR_API_SPEC.md, your server should accept a language form field but provide sensible defaults (typically "en" or "eng") when the parameter is omitted. LiteParse sends this value based on the ocr_language configuration field.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →