How to Integrate LiteParse with EasyOCR and PaddleOCR HTTP Servers

LiteParse delegates optical character recognition to external HTTP servers via a simple REST API, allowing seamless integration with EasyOCR and PaddleOCR wrappers provided in the ocr/ directory of the repository.

LiteParse ships with a flexible OCR subsystem that can fall back to a built-in Tesseract engine or delegate OCR to an external HTTP server. The repository provides ready-to-use Flask wrappers for both EasyOCR and PaddleOCR that implement the LiteParse OCR API specification, enabling high-quality text extraction from scanned PDFs and images without bundling heavy ML models into the core library.

Architecture Overview

When you integrate LiteParse with external OCR HTTP servers, the parsing pipeline follows this flow:

  1. Detection: The core engine in crates/liteparse/src/ocr/http_simple.rs identifies pages lacking native text (scanned images or embedded pictures).
  2. Delegation: If --ocr-server-url is configured, LiteParse instantiates an HttpOcrEngine that sends rendered images to your specified endpoint.
  3. Processing: The HTTP server receives a multipart/form-data request containing a file field (PNG/JPEG) and optional language parameter, then runs its OCR model.
  4. Response: The server returns JSON containing text, bbox, and confidence fields as defined in OCR_API_SPEC.md.
  5. Merging: Results are merged with native text via ocr_merge.rs, preserving spatial coordinates and confidence scores.
  6. Output: The final document is emitted in your chosen format (JSON, Markdown, or plain text).

Setting Up the EasyOCR HTTP Server

The EasyOCR wrapper is located at ocr/easyocr/ and defaults to port 8828.

First, clone the repository and start the server:

git clone https://github.com/run-llama/liteparse.git
cd liteparse/ocr/easyocr
pip install -r requirements.txt
uv run server.py

The server exposes POST /ocr at http://localhost:8828/ocr and implements the LiteParse OCR API specification.

Setting Up the PaddleOCR HTTP Server

The PaddleOCR wrapper resides in ocr/paddleocr/ and listens on port 8829 by default.

Start the server with:

cd liteparse/ocr/paddleocr
pip install -r requirements.txt
uv run server.py

This wrapper is optimized for multilingual documents, particularly Chinese text recognition, and accepts the same multipart request format as the EasyOCR server.

Connecting LiteParse to OCR Servers

Once your HTTP server is running, configure LiteParse to route image processing through it using CLI flags or language binding options.

Command Line Interface

Add --ocr-server-url and optionally --ocr-language to any lit parse command:

lit parse my_scanned.pdf \
    --ocr-server-url http://localhost:8828/ocr \
    --ocr-language en \
    --format markdown -o out.md

For PaddleOCR with Chinese documents:

lit parse my_chinese.pdf \
    --ocr-server-url http://localhost:8829/ocr \
    --ocr-language zh \
    --format json -o out.json

Node.js / TypeScript

Pass ocrServerUrl and ocrLanguage to the LiteParse constructor as shown in ocr/easyocr/README.md:

import { LiteParse } from 'liteparse';

const parser = new LiteParse({
  ocrServerUrl: 'http://localhost:8828/ocr',
  ocrLanguage: 'en',
});

const result = await parser.parse('my_scanned.pdf');
console.log(result.markdown);

Python

The Python bindings in packages/python/liteparse/parser.py mirror the same constructor options:

from liteparse import LiteParse

parser = LiteParse(
    ocr_server_url="http://localhost:8829/ocr",
    ocr_language="zh"
)
result = parser.parse("document.pdf")

Building Custom OCR HTTP Endpoints

You can integrate any OCR engine that conforms to the LiteParse HTTP contract. Your server must accept multipart/form-data POST requests to /ocr with:

  • file: The image file (PNG or JPEG)
  • language (optional): ISO language code (e.g., en, zh)

The response must match the schema in OCR_API_SPEC.md:

{
  "text": "extracted text content",
  "bbox": [x1, y1, x2, y2],
  "confidence": 0.95
}

Configure LiteParse to point to your custom endpoint:

const parser = new LiteParse({
  ocrServerUrl: 'http://my-custom-ocr.com/ocr',
  ocrLanguage: 'fr',
});

Summary

  • LiteParse supports external OCR via HTTP through the HttpOcrEngine implementation in crates/liteparse/src/ocr/http_simple.rs.
  • EasyOCR and PaddleOCR wrappers are provided in ocr/easyocr/ and ocr/paddleocr/ respectively, defaulting to ports 8828 and 8829.
  • The OCR API requires a POST /ocr endpoint accepting multipart form data and returning JSON with text, bbox, and confidence.
  • Configure integration via --ocr-server-url CLI flag or ocrServerUrl constructor parameter in Node.js and Python.
  • Results are automatically merged with native PDF text using ocr_merge.rs to preserve document structure.

Frequently Asked Questions

What HTTP API specification must OCR servers follow to work with LiteParse?

Your server must implement the contract defined in OCR_API_SPEC.md at the repository root. This requires accepting multipart/form-data POST requests with a file field containing the image, and returning a JSON object with text (string), bbox (array of four numbers), and confidence (float between 0 and 1).

Can I use a custom OCR engine instead of EasyOCR or PaddleOCR?

Yes. Any OCR service that exposes an HTTP endpoint matching the LiteParse specification can be used. Simply start your custom server and point LiteParse to it using --ocr-server-url or the equivalent constructor option in your language binding. The core parsing pipeline remains unchanged.

How does LiteParse handle mixed documents containing both native text and scanned images?

The engine in crates/liteparse/src/ocr/http_simple.rs automatically detects pages requiring OCR and routes only those images to the HTTP server. Native text layers are extracted directly from the PDF. The ocr_merge.rs module then combines both streams, maintaining correct reading order and spatial alignment using the bounding box coordinates returned by the OCR server.

What are the default ports for the EasyOCR and PaddleOCR wrappers?

The EasyOCR Flask server defaults to port 8828, while the PaddleOCR server defaults to port 8829. These are documented in ocr/easyocr/README.md and ocr/paddleocr/README.md respectively. You can modify these by editing the server.py files or using environment variables before starting the services.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →