# How to Use LiteParse Node.js Bindings (@llamaindex/liteparse): A Complete Guide

> Master LiteParse Node.js bindings with this guide. Learn to parse PDFs, perform OCR, and generate Markdown using type-safe TypeScript and Rust native binary integration.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-25

---

**LiteParse Node.js bindings provide a type-safe TypeScript wrapper around the Rust core, exposing a minimal `LiteParse` class that handles PDF parsing, OCR, and Markdown generation through native binary integration.**

The `@llamaindex/liteparse` package delivers high-performance document parsing for Node.js applications by loading a compiled native binary (`liteparse-native`) and exposing a clean TypeScript API. This guide demonstrates how to integrate LiteParse into your projects using the official source code from the `run-llama/liteparse` repository.

## Installation and Setup

Install the package via npm to add the native bindings and TypeScript definitions to your project.

```bash
npm install @llamaindex/liteparse

```

The package automatically resolves the correct platform-specific binary through a shim loader located at [`packages/node/src/native.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/native.ts). No manual configuration of the Rust library is required.

## Core Architecture and Configuration

The Node.js wrapper implemented in [`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) performs three critical functions to bridge JavaScript and Rust.

### Configuration Translation

The constructor builds a `LiteParseNativeConfig` object from user-provided partial options and passes it to the native side (`native.LiteParse`). Default values are read back from the native instance, ensuring the JavaScript view matches Rust defaults defined in [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs).

All configuration options map one-to-one with the Rust `LiteParseConfig`, including:
- **OCR toggles** and DPI settings
- **Output format** (`json`, `text`, `markdown`)
- **Image handling mode** and worker pool size

### Input Handling

All public methods accept either a **file path** (`string`) or an **in-memory buffer** (`Buffer`/`Uint8Array`). Bytes are converted to a `Buffer` before being handed to the native library, allowing the binary to work with a single data type regardless of input source.

### Result Mapping

After native execution, the wrapper maps raw Rust structs (`NativeParsedPage`, `NativeExtractedImage`) into friendly TypeScript interfaces (`ParsedPage`, `ExtractedImage`, `ParseResult`). Helper functions like `toPage`, `toImage`, and `toTextItem` preserve spatial metadata including coordinates, font information, OCR confidence, and rotation data.

## Parsing Documents with LiteParse

The API surface is deliberately minimal, offering four primary methods through the `LiteParse` class.

### Basic PDF Parsing

The `parse()` method performs full document parsing including OCR, layout reconstruction, and optional Markdown rendering. It accepts file paths or buffers and returns a `ParseResult` containing pages, concatenated text, and optional embedded images.

```typescript
import { LiteParse } from '@llamaindex/liteparse';

const parser = new LiteParse();
const result = await parser.parse('sample.pdf');
console.log(result.text);

```

### Markdown Output and Image Extraction

Configure the parser to output Markdown format and extract image placeholders for downstream processing.

```typescript
import { LiteParse } from '@llamaindex/liteparse';

const parser = new LiteParse({
  outputFormat: 'markdown',
  imageMode: 'placeholder',
  extractLinks: true,
});
const { text, images } = await parser.parse('invoice.pdf');
console.log(text);
console.log(images.map(i => i.id));

```

The Markdown generation logic resides in [`crates/liteparse/src/output/markdown.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/markdown.rs), producing formatted output with `![](image_p1_0.png)` style placeholders when configured.

### Complexity Checking for OCR Optimization

Use `isComplex()` to run a cheap, text-only scan that flags pages requiring OCR before committing to full processing.

```typescript
import { LiteParse } from '@llamaindex/liteparse';

const parser = new LiteParse({ ocrEnabled: false });
const pageStats = await parser.isComplex('scanned.pdf');

if (pageStats.some(p => p.needsOcr)) {
  const fullResult = await parser.parse('scanned.pdf');
  console.log(fullResult.text);
}

```

This method returns `PageComplexityStats[]` containing per-page verdicts and reasons, allowing conditional OCR activation to optimize performance.

### Generating Page Screenshots

The `screenshot()` method renders selected pages to PNG buffers without requiring external dependencies.

```typescript
import { LiteParse } from '@llamaindex/liteparse';
import { writeFile } from 'fs/promises';

const parser = new LiteParse();
const screenshots = await parser.screenshot('presentation.pdf', [1, 3]);

for (const snap of screenshots) {
  await writeFile(`page-${snap.pageNum}.png`, snap.imageBuffer);
}

```

This returns `ScreenshotResult[]` containing page numbers and binary image data.

## Working with Parse Results

### Searching Text Items

The package exports utility functions to search within parsed content using spatial and text criteria.

```typescript
import { LiteParse, searchItems } from '@llamaindex/liteparse';

const parser = new LiteParse();
const result = await parser.parse('report.pdf');

const hits = searchItems(result.pages[0].textItems, {
  phrase: 'total revenue',
  caseSensitive: false,
});
console.log('Found on page 1 at positions:', hits.map(h => ({x: h.x, y: h.y})));

```

## Advanced Configuration Options

The `LiteParse` constructor accepts a partial configuration object that inherits defaults from [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs). Key parameters include:

- **`ocrEnabled`**: Boolean to toggle Tesseract OCR processing (implemented in `crates/liteparse/src/ocr/`)
- **`outputFormat`**: Specifies whether to return `json`, `text`, or `markdown`
- **`imageMode`**: Controls how images are handled (`embed`, `placeholder`, or `ignore`)
- **`workerPoolSize`**: Thread pool size for parallel processing

Retrieve resolved configuration using `getConfig()`:

```typescript
const parser = new LiteParse({ dpi: 300 });
const config = parser.getConfig();
console.log(config.dpi); // 300

```

## Summary

- **LiteParse Node.js bindings** wrap the Rust core in a type-safe TypeScript API located in [`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts).
- The wrapper handles **configuration translation**, **input normalization** (strings or Buffers), and **result mapping** to TypeScript interfaces.
- Four primary methods provide document processing: `parse()`, `parsePages()`, `isComplex()`, and `screenshot()`.
- Configuration options mirror the Rust `LiteParseConfig` struct exactly, ensuring consistent behavior across languages.
- The package includes utilities like `searchItems()` for working with spatial text data extracted from PDFs.

## Frequently Asked Questions

### How do I enable OCR for scanned documents in LiteParse Node.js?

Set `ocrEnabled: true` in the constructor options. The bindings interface with the OCR engine abstraction layer in `crates/liteparse/src/ocr/`, which supports both built-in Tesseract and HTTP-based OCR services. For performance optimization, first run `isComplex()` to identify which pages actually require OCR before processing the full document.

### Can LiteParse handle PDFs stored as buffers instead of file paths?

Yes. The `parse()` method accepts either a file path string or a `Buffer`/`Uint8Array`. The wrapper in [`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) converts all inputs to Node.js `Buffer` objects before passing them to the native binary, ensuring consistent handling for in-memory PDFs received from HTTP requests or databases.

### What is the difference between `parse()` and `parsePages()` methods?

`parse()` performs full document processing including PDF-level text extraction, layout analysis, and OCR. `parsePages()` skips PDF-level text extraction and directly projects pre-extracted items, which is useful when you have already processed the document through an external OCR pipeline and want to leverage LiteParse's layout reconstruction and formatting capabilities.

### How are images handled when parsing documents?

The `imageMode` configuration option controls image behavior: `embed` includes base64-encoded image data in results, `placeholder` inserts Markdown-style references (processed in [`crates/liteparse/src/output/markdown.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/markdown.rs)), and `ignore` excludes images entirely. Extracted images are returned as `ExtractedImage` objects with metadata including dimensions and page coordinates.