How to Perform Batch Parsing of Multiple Documents with LiteParse

Use the lit batch-parse command or programmatically iterate over files using the LiteParse class with a shared configuration to process entire directories offline and deterministically.

Batch parsing multiple documents with LiteParse leverages the Rust-powered lit CLI and core library to convert file directories into structured text or JSON without network dependencies. The run-llama/liteparse repository separates file orchestration (Node.js/TypeScript) from heavy-duty parsing (Rust), enabling efficient multi-file processing across CLI, Node.js, and Python environments.

Understanding LiteParse Batch Architecture

The batch implementation follows a thin-wrapper pattern: the Node.js CLI handles file discovery and I/O orchestration, while the Rust crate liteparse performs all document processing.

File Discovery and Orchestration

In packages/node/src/cli.ts, the collectFiles function traverses input directories respecting --recursive and --extension filters. This utility gathers matching file paths before any parsing begins, allowing the system to batch-process documents sequentially or with concurrency controls.

Core Parsing Pipeline

For each discovered file, the system instantiates a LiteParse struct (defined in crates/liteparse/src/parser.rs) and invokes the parse_input method. This Rust-native pipeline:

  1. Loads PDFs directly or converts non-PDF inputs to PDF format via crates/liteparse/src/extract.rs
  2. Extracts native text layers and optionally runs OCR through crates/liteparse/src/ocr_merge.rs
  3. Projects spatial layouts into a structured grid via crates/liteparse/src/projection.rs
  4. Serializes results according to LiteParseConfig::output_format settings

All processing occurs locally unless you explicitly configure an external OCR server, making batch parsing fully offline and deterministic.

Batch Parsing with the CLI

The simplest approach uses the lit command-line interface installed via the Node.js or Python packages.

Parse an entire directory to plain text:

lit batch-parse ./my-docs ./parsed-output

Generate JSON outputs with specific filtering options:

lit batch-parse ./my-docs ./json-output \
    --format json \
    --no-ocr \
    --recursive \
    --extension .pdf

This command maps directly to the implementation at lines 59-74 of packages/node/src/cli.ts, which walks the directory, filters extensions, and preserves the original hierarchy in the output folder.

Programmatic Batch Parsing in Node.js

For custom workflows, import the LiteParse class directly and implement recursive file walking. Reuse a single parser instance across files to minimize initialization overhead.

import { LiteParse, LiteParseConfig } from '@llamaindex/liteparse';
import { readdirSync, statSync, join, mkdir, writeFile } from 'fs';
import { parse, dirname } from 'path';
import { promisify } from 'util';

const mkdirAsync = promisify(mkdir);
const writeFileAsync = promisify(writeFile);

// Recursively collect files by extension
function walk(dir: string, ext?: string, files: string[] = []): string[] {
  for (const entry of readdirSync(dir, { withFileTypes: true })) {
    const full = join(dir, entry.name);
    if (entry.isDirectory()) {
      walk(full, ext, files);
    } else if (!ext || entry.name.toLowerCase().endsWith(ext)) {
      files.push(full);
    }
  }
  return files;
}

// Configure once and reuse
const config: Partial<LiteParseConfig> = {
  outputFormat: 'json',
  ocrEnabled: true,
  ocrLanguage: 'eng',
  numWorkers: 4,
};
const parser = new LiteParse(config);

async function batchProcess(inputDir: string, outDir: string) {
  const files = walk(inputDir, '.pdf');
  
  for (const filePath of files) {
    const result = await parser.parse(filePath);
    const rel = dirname(filePath).replace(inputDir, '');
    const outPath = join(outDir, rel, `${parse(filePath).name}.json`);
    
    await mkdirAsync(join(outDir, rel), { recursive: true });
    await writeFileAsync(outPath, JSON.stringify(result, null, 2));
    console.log(`Processed: ${filePath} → ${outPath}`);
  }
}

batchProcess('./my-docs', './parsed-output');

This pattern mirrors the CLI's internal logic while allowing custom preprocessing or postprocessing hooks.

Batch Parsing from Python

Python users access the same Rust core through the shared lit binary shipped with the package. Use subprocess to invoke batch operations while maintaining the full feature set.

import subprocess
import pathlib

input_dir = pathlib.Path("./my-docs")
output_dir = pathlib.Path("./parsed-output")

subprocess.run([
    "lit", "batch-parse",
    str(input_dir),
    str(output_dir),
    "--format", "json",
    "--no-ocr",
    "--recursive"
], check=True)

The Python package bundles the Node.js CLI implementation, ensuring identical behavior across languages.

Configuration and Performance Options

Output Formats and OCR Settings

Control batch behavior through LiteParseConfig parameters defined in crates/liteparse/src/config.rs:

  • outputFormat: Choose text for plain extraction or json for structured output with spatial metadata
  • ocrEnabled: Toggle OCR processing for scanned documents
  • ocrLanguage: Specify language packs (e.g., eng, deu) for Tesseract-based recognition

Worker Concurrency

When processing large directories, set numWorkers to parallelize Rust operations. This configures internal thread pools during LiteParse instantiation without blocking the Node.js event loop.

Summary

  • Use lit batch-parse for command-line directory processing with automatic hierarchy preservation
  • Call collectFiles from packages/node/src/cli.ts when implementing custom file discovery logic
  • Reuse LiteParse instances across multiple files to avoid重复 initialization overhead in programmatic batch processing
  • Configure via LiteParseConfig in crates/liteparse/src/config.rs to control OCR, output format, and concurrency
  • Process offline by default—all document handling in parse_input occurs locally in the Rust crate unless external OCR servers are explicitly enabled

Frequently Asked Questions

Does LiteParse require an internet connection for batch parsing?

No. According to the run-llama/liteparse source code, batch parsing is fully offline and deterministic. All document conversion, text extraction, OCR (via Tesseract), and spatial projection occur within the local Rust crate liteparse unless you explicitly configure an external OCR server endpoint.

Can I parse non-PDF documents in batch mode?

Yes. The parse_input method in crates/liteparse/src/parser.rs automatically converts supported non-PDF inputs to PDF format before processing. Simply include files like Word documents or images in your input directory, or use the --extension filter to target specific formats during the batch operation.

How do I preserve directory structure when parsing output?

Both the CLI and programmatic approaches preserve hierarchy. The lit batch-parse command writes outputs to mirror the input directory structure automatically. When implementing custom Node.js solutions, use path.relative() or similar methods to calculate subdirectory paths when writing results, as demonstrated in the cli.ts reference implementation.

Is there a limit to the number of files I can process in one batch?

There is no hardcoded limit in the LiteParse core. Performance depends on available system memory and the numWorkers configuration setting. For massive directories (thousands of files), consider processing in chunks or ensuring adequate disk space for temporary conversion files during the extraction phase managed by crates/liteparse/src/extract.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →