How to Integrate LiteParse with LlamaIndex for Document Processing

You can integrate LiteParse with LlamaIndex by using its Rust core engine to parse documents into structured JSON, then feeding that output into LlamaIndex's Document objects for vector indexing and retrieval.

LiteParse is a high-performance document parsing engine that converts PDFs and other formats into structured data using a Rust core with Python and Node.js bindings. When you integrate LiteParse with LlamaIndex, you gain spatially-aware text extraction with optional OCR capabilities that feed directly into LlamaIndex's ingestion pipeline. This combination enables accurate document chunking and metadata-rich indexing for retrieval-augmented generation (RAG) applications.

Understanding the LiteParse Architecture

LiteParse follows a three-layer architecture that separates high-performance parsing logic from user-facing APIs.

Core Rust Engine

The parsing logic resides in the Rust core, specifically in [crates/liteparse/src/parser.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs). This file orchestrates format conversion, PDFium text extraction, selective OCR, and spatial grid projection to maintain reading order. The engine supports pluggable OCR backends through the OcrEngine trait defined in [crates/liteparse/src/ocr/mod.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs).

Language Bindings

LiteParse exposes its functionality through thin, language-specific wrappers. The Python binding in [packages/python/liteparse/parser.py](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py) provides the LiteParse class, while the Node.js binding in [packages/node/src/lib.ts](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) offers TypeScript support with async/await patterns. These bindings transfer the structured JSON output from Rust into native data structures for LlamaIndex consumption.

Step-by-Step Integration Workflow

Follow this workflow to connect LiteParse output to LlamaIndex indexing:

  1. Install LiteParse using your language's package manager (e.g., pip install liteparse for Python or npm install @llamaindex/liteparse for Node.js).
  2. Parse documents using the LiteParse class to generate JSON containing text, bounding boxes, and page metadata.
  3. Transform the JSON into LlamaIndex Document objects, mapping LiteParse fields (like text, page_number, and bbox) to document text and metadata.
  4. Chunk and index using LlamaIndex's SimpleNodeParser or custom node parsers, then feed nodes into your vector store index.

Python Integration Example

The Python binding enables direct integration with LlamaIndex's Document and SimpleNodeParser classes:

from liteparse import LiteParse
from llama_index import Document, SimpleNodeParser, GPTVectorStoreIndex
import json

# Initialize parser

lp = LiteParse()

# Parse PDF to structured JSON with spatial data

result = lp.parse("contract.pdf", format="json")
pages = json.loads(result)

# Convert to LlamaIndex Documents with metadata

documents = []
for page in pages:
    doc = Document(
        text=page["text"],
        metadata={
            "page_number": page["page_number"],
            "source": "contract.pdf",
            "bbox": page.get("bbox")
        }
    )
    documents.append(doc)

# Chunk and create index

parser = SimpleNodeParser()
nodes = parser.get_nodes_from_documents(documents)
index = GPTVectorStoreIndex(nodes)

Node.js Integration Example

For TypeScript or JavaScript applications, use the Node.js binding to achieve the same pipeline:

import { LiteParse } from "@llamaindex/liteparse";
import { Document, SimpleNodeParser, GPTVectorStoreIndex } from "llamaindex";

const lp = new LiteParse();

// Parse PDF → JSON
const jsonStr = await lp.parse("contract.pdf", { format: "json" });
const pages = JSON.parse(jsonStr);

// Map to LlamaIndex Documents
const docs = pages.map((p: any) => new Document({
  text: p.text,
  metadata: {
    page_number: p.page_number,
    source: "contract.pdf",
    bbox: p.bbox,
  },
}));

// Create index
const parser = new SimpleNodeParser();
const nodes = await parser.getNodesFromDocuments(docs);
const index = new GPTVectorStoreIndex(nodes);

Key Source Files for Reference

Component File Path Purpose
Core Parser [crates/liteparse/src/parser.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) Orchestrates text extraction, OCR, and spatial projection
OCR Interface [crates/liteparse/src/ocr/mod.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs) Defines OcrEngine trait for Tesseract or HTTP OCR servers
Python Binding [packages/python/liteparse/parser.py](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py) Exposes LiteParse class to Python
Node.js Binding [packages/node/src/lib.ts](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) TypeScript wrapper for the Rust core
CLI Tool [crates/liteparse/src/main.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) Implements the lit command-line interface
URL Parsing Guide [docs/src/content/docs/liteparse/guides/parsing-urls.md](https://github.com/run-llama/liteparse/blob/main/docs/src/content/docs/liteparse/guides/parsing-urls.md) Documentation for remote document ingestion

Summary

Frequently Asked Questions

What output format does LiteParse produce for LlamaIndex ingestion?

LiteParse outputs structured JSON containing page objects with text, page_number, and optional bbox fields. This format maps directly to LlamaIndex's Document constructor, where the text becomes the document content and the remaining fields populate the metadata dictionary for filtering or retrieval strategies.

Can I use LiteParse with LlamaIndex in a browser environment?

Yes. LiteParse compiles to WebAssembly (WASM) for browser use. You can initialize the WASM module, parse files client-side, and feed the resulting JSON into LlamaIndex's browser-compatible indexing solutions (such as LocalVectorStore) without requiring a backend server.

How does LiteParse handle OCR for scanned PDFs?

The engine implements selective OCR through the OcrEngine trait in [crates/liteparse/src/ocr/mod.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs). It can use built-in Tesseract or external HTTP OCR services to extract text from scanned images, then merges those results with native PDF text extraction in [parser.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) to produce a unified spatial reading order.

Is there a CLI option for quick testing before integration?

Yes. The lit command-line tool defined in [crates/liteparse/src/main.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) allows you to parse documents and preview JSON output locally. This is useful for validating extraction quality before writing integration code that connects LiteParse to your LlamaIndex pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →