How to Use LiteParse in Browser Environments with WASM
LiteParse ships a dedicated WebAssembly (WASM) build via the npm package @llamaindex/liteparse-wasm that enables complete PDF parsing—including spatial text extraction and optional OCR—entirely within the browser without server dependencies.
The run-llama/liteparse repository provides a Rust-based PDF parsing pipeline that compiles to WebAssembly for browser deployment. By leveraging the @llamaindex/liteparse-wasm package, you can execute the full parsing logic client-side, processing PDF bytes directly in the browser using a JavaScript-friendly API that mirrors the native Node.js bindings.
Installing the WASM Package
The WASM distribution is published as @llamaindex/liteparse-wasm and contains the pre-compiled binary along with TypeScript definitions. Install it via npm or your preferred package manager:
npm install @llamaindex/liteparse-wasm
Loading and Initializing the Module
The entry point exports an init function that must be called once to load the .wasm binary. In crates/liteparse-wasm/src/lib.rs, the glue code handles instantiation of the Rust core compiled to WASM and sets up the JavaScript bindings.
import init, { LiteParse } from "@llamaindex/liteparse-wasm";
// Initialize the WASM module; this loads the bundled .wasm file
await init();
After initialization, the LiteParse class becomes available for instantiation.
Parsing PDFs in the Browser
The LiteParse class exposed in crates/liteparse-wasm/src/lib.rs accepts a plain JavaScript object for configuration. The glue code converts camelCased fields (e.g., ocrEnabled, outputFormat) into the internal LiteParseConfig struct using the JsLiteParseConfig::into_core method.
const parser = new LiteParse({
ocrEnabled: false,
outputFormat: "json", // or "text"
maxPages: 100,
});
// Convert a File or Blob into Uint8Array
const file: File = /* from <input> or drag-drop */;
const bytes = new Uint8Array(await file.arrayBuffer());
// Parse the PDF
const result = await parser.parse(bytes);
Understanding the Result Structure
The parse method returns a JSON-serializable object built from JsParsedPage and JsTextItem structs. Access result.text for the full document text, or iterate through result.pages[i].textItems to obtain per-item bounding boxes, font information, and confidence scores.
console.log("Full text:", result.text);
console.log("First page items:", result.pages[0].textItems);
Implementing Browser-Based OCR
Because native Tesseract or HTTP OCR backends cannot run in the browser, the WASM crate provides the JsOcrEngine wrapper. The ocrEngine configuration field accepts any JavaScript object implementing an async recognize(imageData, width, height, language) method, which the Rust code invokes via the bridge in lib.rs.
const parser = new LiteParse({
ocrEnabled: true,
ocrLanguage: "eng",
ocrEngine: {
async recognize(imageData: Uint8Array, width: number, height: number, language: string) {
// imageData is a PNG bytes buffer produced by LiteParse
const { data } = await Tesseract.recognize(
new Uint8Array(imageData),
language,
{ rectangle: { left: 0, top: 0, width, height } }
);
// Return format expected by JsOcrEngine
return data.words.map(w => ({
text: w.text,
bbox: [w.bbox.x0, w.bbox.y0, w.bbox.x1, w.bbox.y1],
confidence: w.confidence / 100,
}));
},
},
});
WASM Architecture and Limitations
The WASM build runs single-threaded because browsers currently expose only a single thread to WebAssembly. According to the source in crates/liteparse/src/config.rs, the configuration explicitly sets cfg.num_workers = 1. All heavy lifting—PDF rendering, text extraction, and spatial projection—is performed by the Rust core in crates/liteparse/src/parser.rs compiled to WASM, delivering performance comparable to the native CLI.
For a complete working implementation, reference the demo page in wasm-demo-site/index.html, which demonstrates CDN loading, drag-and-drop UI integration, and status handling.
Summary
- @llamaindex/liteparse-wasm provides the official browser distribution of LiteParse, located in
crates/liteparse-wasm. - Initialize the module with
init()before constructing theLiteParseclass using camelCased config options likeocrEnabledandoutputFormat. - Parse PDFs by passing a
Uint8Arrayto the asyncparse()method, which returnsJsParsedPageandJsTextItemdata structures. - Enable OCR in the browser by supplying a JavaScript engine (e.g., tesseract.js) to the
ocrEnginefield, which must implement therecognize(imageData, width, height, language)method. - The WASM execution is limited to a single thread (
cfg.num_workers = 1), though the Rust core handles all parsing operations efficiently within that constraint.
Frequently Asked Questions
How do I load the LiteParse WASM module from a CDN?
You can load the module directly from a CDN or from node_modules/@llamaindex/liteparse-wasm/pkg. The wasm-demo-site/index.html file demonstrates the pattern: import the package, then call the default export (the init function) with the path to the .wasm binary. This initializes the internal WASM memory and Rust runtime before you instantiate the LiteParse class.
Why is the WASM build limited to single-threaded execution?
Browsers currently only expose a single thread to WebAssembly, so the LiteParse configuration explicitly sets cfg.num_workers = 1 when running in WASM. Despite this limitation, the Rust core compiled from crates/liteparse/src/parser.rs performs all PDF rendering, text extraction, and spatial projection within that single thread, maintaining performance comparable to the native CLI for most documents.
How do I integrate OCR when using LiteParse in the browser?
You must provide a JavaScript OCR engine via the ocrEngine configuration field because native Tesseract backends cannot run in the browser. The engine must expose an async recognize(imageData, width, height, language) method that accepts PNG byte buffers from LiteParse and returns text items with bounding boxes. The JsOcrEngine wrapper in crates/liteparse-wasm/src/lib.rs forwards these calls between the Rust parser and your JavaScript implementation.
What is the difference between the "json" and "text" output formats?
When outputFormat is set to "json", the parse() method returns a structured object containing pages and textItems with full metadata including bounding boxes and fonts. When set to "text", the parser returns a simplified object where result.text contains the plain document text without per-item spatial data. Both formats are serialized from JsParsedPage structures in the WASM glue code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →