# How to Integrate LiteParse with LlamaIndex for Document Processing

> Integrate LiteParse with LlamaIndex to efficiently parse documents into structured JSON for advanced vector indexing and retrieval with LlamaIndex's Document objects.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-07

---

**You can integrate LiteParse with LlamaIndex by using its Rust core engine to parse documents into structured JSON, then feeding that output into LlamaIndex's Document objects for vector indexing and retrieval.**

LiteParse is a high-performance document parsing engine that converts PDFs and other formats into structured data using a Rust core with Python and Node.js bindings. When you integrate LiteParse with LlamaIndex, you gain spatially-aware text extraction with optional OCR capabilities that feed directly into LlamaIndex's ingestion pipeline. This combination enables accurate document chunking and metadata-rich indexing for retrieval-augmented generation (RAG) applications.

## Understanding the LiteParse Architecture

LiteParse follows a three-layer architecture that separates high-performance parsing logic from user-facing APIs.

### Core Rust Engine

The parsing logic resides in the Rust core, specifically in [[`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs). This file orchestrates format conversion, PDFium text extraction, selective OCR, and spatial grid projection to maintain reading order. The engine supports pluggable OCR backends through the `OcrEngine` trait defined in [[`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs).

### Language Bindings

LiteParse exposes its functionality through thin, language-specific wrappers. The Python binding in [[`packages/python/liteparse/parser.py`](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py)](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py) provides the `LiteParse` class, while the Node.js binding in [[`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts)](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) offers TypeScript support with async/await patterns. These bindings transfer the structured JSON output from Rust into native data structures for LlamaIndex consumption.

## Step-by-Step Integration Workflow

Follow this workflow to connect LiteParse output to LlamaIndex indexing:

1. **Install LiteParse** using your language's package manager (e.g., `pip install liteparse` for Python or `npm install @llamaindex/liteparse` for Node.js).
2. **Parse documents** using the `LiteParse` class to generate JSON containing text, bounding boxes, and page metadata.
3. **Transform the JSON** into LlamaIndex `Document` objects, mapping LiteParse fields (like `text`, `page_number`, and `bbox`) to document text and metadata.
4. **Chunk and index** using LlamaIndex's `SimpleNodeParser` or custom node parsers, then feed nodes into your vector store index.

## Python Integration Example

The Python binding enables direct integration with LlamaIndex's `Document` and `SimpleNodeParser` classes:

```python
from liteparse import LiteParse
from llama_index import Document, SimpleNodeParser, GPTVectorStoreIndex
import json

# Initialize parser

lp = LiteParse()

# Parse PDF to structured JSON with spatial data

result = lp.parse("contract.pdf", format="json")
pages = json.loads(result)

# Convert to LlamaIndex Documents with metadata

documents = []
for page in pages:
    doc = Document(
        text=page["text"],
        metadata={
            "page_number": page["page_number"],
            "source": "contract.pdf",
            "bbox": page.get("bbox")
        }
    )
    documents.append(doc)

# Chunk and create index

parser = SimpleNodeParser()
nodes = parser.get_nodes_from_documents(documents)
index = GPTVectorStoreIndex(nodes)

```

## Node.js Integration Example

For TypeScript or JavaScript applications, use the Node.js binding to achieve the same pipeline:

```typescript
import { LiteParse } from "@llamaindex/liteparse";
import { Document, SimpleNodeParser, GPTVectorStoreIndex } from "llamaindex";

const lp = new LiteParse();

// Parse PDF → JSON
const jsonStr = await lp.parse("contract.pdf", { format: "json" });
const pages = JSON.parse(jsonStr);

// Map to LlamaIndex Documents
const docs = pages.map((p: any) => new Document({
  text: p.text,
  metadata: {
    page_number: p.page_number,
    source: "contract.pdf",
    bbox: p.bbox,
  },
}));

// Create index
const parser = new SimpleNodeParser();
const nodes = await parser.getNodesFromDocuments(docs);
const index = new GPTVectorStoreIndex(nodes);

```

## Key Source Files for Reference

| Component | File Path | Purpose |
|-----------|-----------|---------|
| **Core Parser** | [[`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) | Orchestrates text extraction, OCR, and spatial projection |
| **OCR Interface** | [[`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs) | Defines `OcrEngine` trait for Tesseract or HTTP OCR servers |
| **Python Binding** | [[`packages/python/liteparse/parser.py`](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py)](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py) | Exposes `LiteParse` class to Python |
| **Node.js Binding** | [[`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts)](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) | TypeScript wrapper for the Rust core |
| **CLI Tool** | [[`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) | Implements the `lit` command-line interface |
| **URL Parsing Guide** | [[`docs/src/content/docs/liteparse/guides/parsing-urls.md`](https://github.com/run-llama/liteparse/blob/main/docs/src/content/docs/liteparse/guides/parsing-urls.md)](https://github.com/run-llama/liteparse/blob/main/docs/src/content/docs/liteparse/guides/parsing-urls.md) | Documentation for remote document ingestion |

## Summary

- **LiteParse** provides a Rust-based parsing core with JSON output containing spatial text data and optional OCR results.
- **Language bindings** in [[`packages/python/liteparse/parser.py`](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py)](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py) and [[`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts)](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) bridge the Rust engine to LlamaIndex's Python and TypeScript ecosystems.
- **Integration** requires parsing documents to JSON, converting pages to LlamaIndex `Document` objects with metadata, then using standard LlamaIndex node parsers for chunking.
- **Spatial metadata** (bounding boxes, page numbers) preserved from [[`parser.rs`](https://github.com/run-llama/liteparse/blob/main/parser.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) enables advanced retrieval strategies like visual grounding or section-aware chunking.

## Frequently Asked Questions

### What output format does LiteParse produce for LlamaIndex ingestion?

LiteParse outputs structured **JSON** containing page objects with `text`, `page_number`, and optional `bbox` fields. This format maps directly to LlamaIndex's `Document` constructor, where the text becomes the document content and the remaining fields populate the metadata dictionary for filtering or retrieval strategies.

### Can I use LiteParse with LlamaIndex in a browser environment?

Yes. LiteParse compiles to **WebAssembly (WASM)** for browser use. You can initialize the WASM module, parse files client-side, and feed the resulting JSON into LlamaIndex's browser-compatible indexing solutions (such as `LocalVectorStore`) without requiring a backend server.

### How does LiteParse handle OCR for scanned PDFs?

The engine implements selective OCR through the `OcrEngine` trait in [[`crates/liteparse/src/ocr/mod.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/mod.rs). It can use built-in Tesseract or external HTTP OCR services to extract text from scanned images, then merges those results with native PDF text extraction in [[`parser.rs`](https://github.com/run-llama/liteparse/blob/main/parser.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) to produce a unified spatial reading order.

### Is there a CLI option for quick testing before integration?

Yes. The `lit` command-line tool defined in [[`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs)](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) allows you to parse documents and preview JSON output locally. This is useful for validating extraction quality before writing integration code that connects LiteParse to your LlamaIndex pipeline.