# How to Perform Batch Parsing of Multiple Documents with LiteParse

> Effortlessly batch parse multiple documents with LiteParse. Use the lit batch-parse command or the LiteParse class to process entire directories offline and deterministically.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-07

---

**Use the `lit batch-parse` command or programmatically iterate over files using the `LiteParse` class with a shared configuration to process entire directories offline and deterministically.**

Batch parsing multiple documents with LiteParse leverages the Rust-powered `lit` CLI and core library to convert file directories into structured text or JSON without network dependencies. The run-llama/liteparse repository separates file orchestration (Node.js/TypeScript) from heavy-duty parsing (Rust), enabling efficient multi-file processing across CLI, Node.js, and Python environments.

## Understanding LiteParse Batch Architecture

The batch implementation follows a thin-wrapper pattern: the Node.js CLI handles file discovery and I/O orchestration, while the Rust crate `liteparse` performs all document processing.

### File Discovery and Orchestration

In [`packages/node/src/cli.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/cli.ts), the `collectFiles` function traverses input directories respecting `--recursive` and `--extension` filters. This utility gathers matching file paths before any parsing begins, allowing the system to batch-process documents sequentially or with concurrency controls.

### Core Parsing Pipeline

For each discovered file, the system instantiates a `LiteParse` struct (defined in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs)) and invokes the `parse_input` method. This Rust-native pipeline:

1. Loads PDFs directly or converts non-PDF inputs to PDF format via [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs)
2. Extracts native text layers and optionally runs OCR through [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs)
3. Projects spatial layouts into a structured grid via [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs)
4. Serializes results according to `LiteParseConfig::output_format` settings

All processing occurs locally unless you explicitly configure an external OCR server, making batch parsing fully **offline** and **deterministic**.

## Batch Parsing with the CLI

The simplest approach uses the `lit` command-line interface installed via the Node.js or Python packages.

Parse an entire directory to plain text:

```bash
lit batch-parse ./my-docs ./parsed-output

```

Generate JSON outputs with specific filtering options:

```bash
lit batch-parse ./my-docs ./json-output \
    --format json \
    --no-ocr \
    --recursive \
    --extension .pdf

```

This command maps directly to the implementation at lines 59-74 of [`packages/node/src/cli.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/cli.ts), which walks the directory, filters extensions, and preserves the original hierarchy in the output folder.

## Programmatic Batch Parsing in Node.js

For custom workflows, import the `LiteParse` class directly and implement recursive file walking. Reuse a single parser instance across files to minimize initialization overhead.

```typescript
import { LiteParse, LiteParseConfig } from '@llamaindex/liteparse';
import { readdirSync, statSync, join, mkdir, writeFile } from 'fs';
import { parse, dirname } from 'path';
import { promisify } from 'util';

const mkdirAsync = promisify(mkdir);
const writeFileAsync = promisify(writeFile);

// Recursively collect files by extension
function walk(dir: string, ext?: string, files: string[] = []): string[] {
  for (const entry of readdirSync(dir, { withFileTypes: true })) {
    const full = join(dir, entry.name);
    if (entry.isDirectory()) {
      walk(full, ext, files);
    } else if (!ext || entry.name.toLowerCase().endsWith(ext)) {
      files.push(full);
    }
  }
  return files;
}

// Configure once and reuse
const config: Partial<LiteParseConfig> = {
  outputFormat: 'json',
  ocrEnabled: true,
  ocrLanguage: 'eng',
  numWorkers: 4,
};
const parser = new LiteParse(config);

async function batchProcess(inputDir: string, outDir: string) {
  const files = walk(inputDir, '.pdf');
  
  for (const filePath of files) {
    const result = await parser.parse(filePath);
    const rel = dirname(filePath).replace(inputDir, '');
    const outPath = join(outDir, rel, `${parse(filePath).name}.json`);
    
    await mkdirAsync(join(outDir, rel), { recursive: true });
    await writeFileAsync(outPath, JSON.stringify(result, null, 2));
    console.log(`Processed: ${filePath} → ${outPath}`);
  }
}

batchProcess('./my-docs', './parsed-output');

```

This pattern mirrors the CLI's internal logic while allowing custom preprocessing or postprocessing hooks.

## Batch Parsing from Python

Python users access the same Rust core through the shared `lit` binary shipped with the package. Use `subprocess` to invoke batch operations while maintaining the full feature set.

```python
import subprocess
import pathlib

input_dir = pathlib.Path("./my-docs")
output_dir = pathlib.Path("./parsed-output")

subprocess.run([
    "lit", "batch-parse",
    str(input_dir),
    str(output_dir),
    "--format", "json",
    "--no-ocr",
    "--recursive"
], check=True)

```

The Python package bundles the Node.js CLI implementation, ensuring identical behavior across languages.

## Configuration and Performance Options

### Output Formats and OCR Settings

Control batch behavior through `LiteParseConfig` parameters defined in [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs):

- **`outputFormat`**: Choose `text` for plain extraction or `json` for structured output with spatial metadata
- **`ocrEnabled`**: Toggle OCR processing for scanned documents
- **`ocrLanguage`**: Specify language packs (e.g., `eng`, `deu`) for Tesseract-based recognition

### Worker Concurrency

When processing large directories, set `numWorkers` to parallelize Rust operations. This configures internal thread pools during `LiteParse` instantiation without blocking the Node.js event loop.

## Summary

- **Use `lit batch-parse`** for command-line directory processing with automatic hierarchy preservation
- **Call `collectFiles`** from [`packages/node/src/cli.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/cli.ts) when implementing custom file discovery logic
- **Reuse `LiteParse` instances** across multiple files to avoid重复 initialization overhead in programmatic batch processing
- **Configure via `LiteParseConfig`** in [`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs) to control OCR, output format, and concurrency
- **Process offline** by default—all document handling in `parse_input` occurs locally in the Rust crate unless external OCR servers are explicitly enabled

## Frequently Asked Questions

### Does LiteParse require an internet connection for batch parsing?

No. According to the run-llama/liteparse source code, batch parsing is fully offline and deterministic. All document conversion, text extraction, OCR (via Tesseract), and spatial projection occur within the local Rust crate `liteparse` unless you explicitly configure an external OCR server endpoint.

### Can I parse non-PDF documents in batch mode?

Yes. The `parse_input` method in [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs) automatically converts supported non-PDF inputs to PDF format before processing. Simply include files like Word documents or images in your input directory, or use the `--extension` filter to target specific formats during the batch operation.

### How do I preserve directory structure when parsing output?

Both the CLI and programmatic approaches preserve hierarchy. The `lit batch-parse` command writes outputs to mirror the input directory structure automatically. When implementing custom Node.js solutions, use `path.relative()` or similar methods to calculate subdirectory paths when writing results, as demonstrated in the [`cli.ts`](https://github.com/run-llama/liteparse/blob/main/cli.ts) reference implementation.

### Is there a limit to the number of files I can process in one batch?

There is no hardcoded limit in the LiteParse core. Performance depends on available system memory and the `numWorkers` configuration setting. For massive directories (thousands of files), consider processing in chunks or ensuring adequate disk space for temporary conversion files during the extraction phase managed by [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs).