# How to Implement a Custom HTTP OCR Server Compatible with LiteParse

> Implement a custom HTTP OCR server for LiteParse. Learn to create a POST /ocr endpoint for image uploads and configure LiteParse to seamlessly integrate your OCR solution.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-06

---

**You can implement a custom HTTP OCR server compatible with LiteParse by exposing a `POST /ocr` endpoint that accepts multipart/form-data image uploads and returns JSON containing recognized text, bounding boxes, and confidence scores, then configuring LiteParse to use your server URL.**

LiteParse is a Rust-based PDF parsing library from the `run-llama/liteparse` repository that supports pluggable OCR backends. When you require more control over text recognition than the built-in Tesseract engine provides, implementing a custom HTTP OCR server allows you to integrate specialized engines like EasyOCR, PaddleOCR, or proprietary solutions while maintaining seamless compatibility with LiteParse's document processing pipeline.

## Understanding the LiteParse HTTP OCR Architecture

LiteParse delegates OCR operations to an external service when you configure an HTTP endpoint. Understanding how the library selects and communicates with this engine is essential for building a compatible server.

### Configuration and Engine Selection

The OCR backend is determined by the `ocr_server_url` field in `LiteParseConfig`, defined in [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs). When this URL is provided, LiteParse instantiates an `HttpOcrEngine` instead of the default local Tesseract engine.

In [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs), the `LiteParse::parse_input` method contains the selection logic:

```rust
if let Some(ref url) = self.config.ocr_server_url {
    std::sync::Arc::new(HttpOcrEngine::new(url.clone()))
} else {
    // fallback to Tesseract if the `tesseract` feature is enabled
}

```

This conditional check ensures that setting `ocr_server_url` automatically routes all image recognition tasks through the HTTP pathway.

### The HTTP OCR Engine Implementation

The `HttpOcrEngine` struct, located in [`crates/liteparse/src/ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/http_simple.rs), implements the `OcrEngine` trait defined in [`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs). Its `recognize` method performs three critical operations:

1. Converts the raw RGB page image into PNG format
2. Constructs a multipart/form-data POST request to the configured server URL
3. Deserializes the JSON response into `OcrResult` objects containing text, bounding boxes, and confidence values

This implementation expects your server to handle binary image data and respond with a specific JSON schema.

## The LiteParse OCR API Specification

Any custom HTTP OCR server must conform to the **LiteParse OCR API Specification** documented in [`OCR_API_SPEC.md`](https://github.com/run-llama/liteparse/blob/main/OCR_API_SPEC.md) at the repository root. This contract defines the exact interface between LiteParse and your service.

### Required Endpoint and Parameters

Your server must expose:

- **Endpoint**: `POST /ocr`
- **Content-Type**: `multipart/form-data`
- **Parameters**:
  - `file`: Binary image data (PNG or JPEG)
  - `language` (optional): ISO language code (e.g., "en", "zh", "eng")

### Response Format

The endpoint must return HTTP 200 with a JSON body structured as follows:

```json
{
  "results": [
    {
      "text": "recognized text",
      "bbox": [x1, y1, x2, y2],
      "confidence": 0.97
    }
  ]
}

```

Each result object requires:

- `text`: The recognized string
- `bbox`: An array of four floats representing the bounding box coordinates `[left, top, right, bottom]`
- `confidence`: A float between 0.0 and 1.0 indicating recognition certainty

## Implementing Your Custom HTTP OCR Server

Building a compliant server requires handling image decoding, processing through your chosen OCR engine, and formatting results according to the specification.

### Minimal FastAPI Implementation

Below is a complete, minimal implementation using Python and FastAPI that satisfies LiteParse requirements:

```python
from fastapi import FastAPI, File, Form, UploadFile
from pydantic import BaseModel
import your_ocr_engine  # Replace with actual OCR library

app = FastAPI()

class OcrResponse(BaseModel):
    results: list[dict]

@app.post("/ocr")
async def ocr_endpoint(
    file: UploadFile = File(...), 
    language: str = Form(default="en")
) -> OcrResponse:
    # Read image bytes from the request

    image_bytes = await file.read()
    
    # Process with your OCR engine

    # This is pseudocode - implement according to your engine's API

    raw_results = your_ocr_engine.process(image_bytes, lang=language)
    
    # Format results to match LiteParse specification

    formatted_results = [
        {
            "text": item.text,
            "bbox": [item.x1, item.y1, item.x2, item.y2],
            "confidence": item.confidence
        }
        for item in raw_results
    ]
    
    return OcrResponse(results=formatted_results)

```

Run this server with `uvicorn main:app --host 0.0.0.0 --port 9000`.

### Handling Image Processing

The `HttpOcrEngine` sends PNG-encoded images via multipart/form-data. Your server should:

- Accept the `file` field as binary data
- Support common image formats (PNG is guaranteed from LiteParse, but JPEG handling adds flexibility)
- Respect the optional `language` parameter for multilingual documents
- Return an empty `results` array for pages containing no text rather than HTTP error codes

## Integrating Your Server with LiteParse

Once your server is running, configure LiteParse to use it instead of the local Tesseract engine.

### Command Line Interface

Use the `--ocr-server-url` flag when parsing documents:

```bash
lit parse document.pdf --ocr-server-url http://localhost:9000/ocr

```

For multilingual documents, specify the language code:

```bash
lit parse document.pdf --ocr-server-url http://localhost:9000/ocr --ocr-language zh

```

### Programmatic Integration (Rust)

In Rust applications using `liteparse` directly, set the `ocr_server_url` field in `LiteParseConfig`:

```rust
use liteparse::{LiteParse, LiteParseConfig};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let cfg = LiteParseConfig {
        ocr_enabled: true,
        ocr_server_url: Some("http://localhost:9000/ocr".into()),
        ocr_language: Some("en".into()),
        ..Default::default()
    };
    
    let parser = LiteParse::new(cfg);
    let result = parser.parse_input(
        liteparse::PdfInput::Path("document.pdf".into())
    ).await?;
    
    println!("{}", result.text);
    Ok(())
}

```

### Programmatic Integration (Node.js/TypeScript)

When using the JavaScript bindings, pass the URL in the configuration object:

```typescript
import { LiteParse } from 'liteparse';

const parser = new LiteParse({
  ocrServerUrl: 'http://localhost:9000/ocr',
  ocrLanguage: 'en',
});

const result = await parser.parse('document.pdf');
console.log(result.text);

```

## Reference Implementations

The `run-llama/liteparse` repository provides two production-ready reference servers that demonstrate best practices for implementing the specification.

### EasyOCR Server

Located at [`ocr/easyocr/server.py`](https://github.com/run-llama/liteparse/blob/main/ocr/easyocr/server.py), this FastAPI application wraps the EasyOCR library:

```bash
cd ocr/easyocr
uv run server.py  # Starts on http://0.0.0.0:8828

```

This implementation handles automatic GPU detection and supports all languages available in the EasyOCR model zoo.

### PaddleOCR Server

Found at [`ocr/paddleocr/server.py`](https://github.com/run-llama/liteparse/blob/main/ocr/paddleocr/server.py), this server leverages PaddleOCR for high-performance text recognition:

```bash
cd ocr/paddleocr
uv run server.py  # Starts on http://0.0.0.0:8829

```

You can enable GPU acceleration by modifying the `use_gpu` parameter in the `PaddleOCRServer` class initialization.

Both servers expose `GET /health` endpoints for load balancer health checks and implement detailed request logging compatible with the LiteParse HTTP OCR engine expectations.

## Summary

Implementing a custom HTTP OCR server compatible with LiteParse requires adherence to a specific API contract while providing flexibility for your chosen OCR technology:

- **Configuration**: Set `ocr_server_url` in `LiteParseConfig` or use the `--ocr-server-url` CLI flag to route OCR requests to your server
- **API Contract**: Implement `POST /ocr` accepting multipart/form-data with `file` and optional `language` fields, returning JSON with `results` containing `text`, `bbox`, and `confidence`
- **Engine Selection**: LiteParse automatically instantiates `HttpOcrEngine` from [`crates/liteparse/src/ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/http_simple.rs) when a URL is configured, converting PDF pages to PNG and sending them via HTTP
- **Reference Code**: Study the EasyOCR ([`ocr/easyocr/server.py`](https://github.com/run-llama/liteparse/blob/main/ocr/easyocr/server.py)) and PaddleOCR ([`ocr/paddleocr/server.py`](https://github.com/run-llama/liteparse/blob/main/ocr/paddleocr/server.py)) implementations for production-ready patterns

## Frequently Asked Questions

### What image format does LiteParse send to the HTTP OCR server?

LiteParse converts PDF pages to **PNG format** before transmission. The `HttpOcrEngine::recognize` method in [`crates/liteparse/src/ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/http_simple.rs) handles the RGB to PNG encoding automatically, ensuring lossless quality for optimal text recognition accuracy.

### Can I use GPU acceleration with a custom HTTP OCR server?

Yes. GPU acceleration is independent of LiteParse and depends entirely on your server implementation. The reference **PaddleOCR server** at [`ocr/paddleocr/server.py`](https://github.com/run-llama/liteparse/blob/main/ocr/paddleocr/server.py) demonstrates GPU support by setting `use_gpu=True` in the PaddleOCR constructor, while the client configuration in LiteParse remains unchanged.

### How does LiteParse handle OCR server failures or timeouts?

The `HttpOcrEngine` in [`crates/liteparse/src/ocr/http_simple.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/http_simple.rs) propagates HTTP errors as OCR failures. If your server returns non-200 status codes or timeouts occur, the parsing operation will fail with the corresponding error. Implement proper health checks using the `GET /health` endpoint pattern shown in the reference servers to ensure high availability.

### Is the language parameter required in the OCR API specification?

No, the `language` parameter is optional. According to the specification in [`OCR_API_SPEC.md`](https://github.com/run-llama/liteparse/blob/main/OCR_API_SPEC.md), your server should accept a `language` form field but provide sensible defaults (typically "en" or "eng") when the parameter is omitted. LiteParse sends this value based on the `ocr_language` configuration field.