How to Integrate LiteParse with LlamaIndex for Document Processing
You can integrate LiteParse with LlamaIndex by using its Rust core engine to parse documents into structured JSON, then feeding that output into LlamaIndex's Document objects for vector indexing and retrieval.
LiteParse is a high-performance document parsing engine that converts PDFs and other formats into structured data using a Rust core with Python and Node.js bindings. When you integrate LiteParse with LlamaIndex, you gain spatially-aware text extraction with optional OCR capabilities that feed directly into LlamaIndex's ingestion pipeline. This combination enables accurate document chunking and metadata-rich indexing for retrieval-augmented generation (RAG) applications.
Understanding the LiteParse Architecture
LiteParse follows a three-layer architecture that separates high-performance parsing logic from user-facing APIs.
Core Rust Engine
The parsing logic resides in the Rust core, specifically in [crates/liteparse/src/parser.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs). This file orchestrates format conversion, PDFium text extraction, selective OCR, and spatial grid projection to maintain reading order. The engine supports pluggable OCR backends through the OcrEngine trait defined in [crates/liteparse/src/ocr/mod.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs).
Language Bindings
LiteParse exposes its functionality through thin, language-specific wrappers. The Python binding in [packages/python/liteparse/parser.py](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py) provides the LiteParse class, while the Node.js binding in [packages/node/src/lib.ts](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) offers TypeScript support with async/await patterns. These bindings transfer the structured JSON output from Rust into native data structures for LlamaIndex consumption.
Step-by-Step Integration Workflow
Follow this workflow to connect LiteParse output to LlamaIndex indexing:
- Install LiteParse using your language's package manager (e.g.,
pip install liteparsefor Python ornpm install @llamaindex/liteparsefor Node.js). - Parse documents using the
LiteParseclass to generate JSON containing text, bounding boxes, and page metadata. - Transform the JSON into LlamaIndex
Documentobjects, mapping LiteParse fields (liketext,page_number, andbbox) to document text and metadata. - Chunk and index using LlamaIndex's
SimpleNodeParseror custom node parsers, then feed nodes into your vector store index.
Python Integration Example
The Python binding enables direct integration with LlamaIndex's Document and SimpleNodeParser classes:
from liteparse import LiteParse
from llama_index import Document, SimpleNodeParser, GPTVectorStoreIndex
import json
# Initialize parser
lp = LiteParse()
# Parse PDF to structured JSON with spatial data
result = lp.parse("contract.pdf", format="json")
pages = json.loads(result)
# Convert to LlamaIndex Documents with metadata
documents = []
for page in pages:
doc = Document(
text=page["text"],
metadata={
"page_number": page["page_number"],
"source": "contract.pdf",
"bbox": page.get("bbox")
}
)
documents.append(doc)
# Chunk and create index
parser = SimpleNodeParser()
nodes = parser.get_nodes_from_documents(documents)
index = GPTVectorStoreIndex(nodes)
Node.js Integration Example
For TypeScript or JavaScript applications, use the Node.js binding to achieve the same pipeline:
import { LiteParse } from "@llamaindex/liteparse";
import { Document, SimpleNodeParser, GPTVectorStoreIndex } from "llamaindex";
const lp = new LiteParse();
// Parse PDF → JSON
const jsonStr = await lp.parse("contract.pdf", { format: "json" });
const pages = JSON.parse(jsonStr);
// Map to LlamaIndex Documents
const docs = pages.map((p: any) => new Document({
text: p.text,
metadata: {
page_number: p.page_number,
source: "contract.pdf",
bbox: p.bbox,
},
}));
// Create index
const parser = new SimpleNodeParser();
const nodes = await parser.getNodesFromDocuments(docs);
const index = new GPTVectorStoreIndex(nodes);
Key Source Files for Reference
Summary
- LiteParse provides a Rust-based parsing core with JSON output containing spatial text data and optional OCR results.
- Language bindings in [
packages/python/liteparse/parser.py](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py) and [packages/node/src/lib.ts](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) bridge the Rust engine to LlamaIndex's Python and TypeScript ecosystems. - Integration requires parsing documents to JSON, converting pages to LlamaIndex
Documentobjects with metadata, then using standard LlamaIndex node parsers for chunking. - Spatial metadata (bounding boxes, page numbers) preserved from [
parser.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) enables advanced retrieval strategies like visual grounding or section-aware chunking.
Frequently Asked Questions
What output format does LiteParse produce for LlamaIndex ingestion?
LiteParse outputs structured JSON containing page objects with text, page_number, and optional bbox fields. This format maps directly to LlamaIndex's Document constructor, where the text becomes the document content and the remaining fields populate the metadata dictionary for filtering or retrieval strategies.
Can I use LiteParse with LlamaIndex in a browser environment?
Yes. LiteParse compiles to WebAssembly (WASM) for browser use. You can initialize the WASM module, parse files client-side, and feed the resulting JSON into LlamaIndex's browser-compatible indexing solutions (such as LocalVectorStore) without requiring a backend server.
How does LiteParse handle OCR for scanned PDFs?
The engine implements selective OCR through the OcrEngine trait in [crates/liteparse/src/ocr/mod.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs). It can use built-in Tesseract or external HTTP OCR services to extract text from scanned images, then merges those results with native PDF text extraction in [parser.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) to produce a unified spatial reading order.
Is there a CLI option for quick testing before integration?
Yes. The lit command-line tool defined in [crates/liteparse/src/main.rs](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) allows you to parse documents and preview JSON output locally. This is useful for validating extraction quality before writing integration code that connects LiteParse to your LlamaIndex pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →