LiteParse Batch-Parse: Processing Multiple Documents in One Command

LiteParse's batch-parse functionality allows you to process entire directories of documents (PDFs, DOCXs, images) through a single CLI command or API call, reusing a single LiteParse instance to minimize resource overhead while preserving directory structure in the output.

The run-llama/liteparse repository provides a dedicated batch-parse command that transforms document processing from a file-by-file operation into an efficient batch workflow. This functionality processes mixed document types automatically and outputs structured results without requiring manual orchestration of individual parse calls.

How Batch-Parse Works in LiteParse

The implementation lives in crates/liteparse/src/main.rs, where the batch-parse subcommand orchestrates the pipeline. The process follows six distinct stages that optimize for throughput and error resilience.

CLI Entry Point and Configuration

The BatchParseCommand struct defines all CLI flags including input directory, output directory, format selection, OCR toggles, DPI settings, and recursion options. This configuration is collected upfront in crates/liteparse/src/main.rs to avoid re-parsing arguments during the document loop.

File Discovery and Filtering

The helper function collect_files walks the input directory recursively when requested, returning a sorted Vec<String> of file paths. An optional extension filter applied during collection ensures only relevant documents enter the processing queue.

Output Path Resolution

For each input file, the batch_output_path function computes the corresponding output location by stripping the input directory prefix, preserving nested subfolder structures, and appending the appropriate extension based on the selected output format.

The Parsing Loop and Resource Management

A single LiteParse instance is instantiated with the batch-specific configuration before entering the loop. This approach reuses the OCR worker pool and other heavy resources across all documents, significantly reducing per-file overhead. The CLI iterates over collected files, invoking lp.parse(file_path).await for each entry, then formats the ParseResult according to the requested output format.

Command-Line Usage Examples

Use the lit batch-parse command to process documents without writing code.

Basic recursive parsing:

lit batch-parse \
    --input-dir ./documents \
    --output-dir ./out \
    --format markdown \
    --recursive \
    --extension .pdf

This command scans ./documents recursively, extracts all PDFs, parses each with OCR enabled by default, and writes .md files mirroring the input directory structure under ./out.

Programmatic Batch Processing

Beyond the CLI, LiteParse exposes the same batch functionality through Rust and Node.js APIs.

Rust Implementation

When building custom pipelines, instantiate LiteParse once and reuse it across multiple files:

use liteparse::{LiteParse, config::LiteParseConfig, parser::PdfInput};
use std::path::Path;

let config = LiteParseConfig {
    ocr_enabled: true,
    ocr_language: "eng".into(),
    output_format: liteparse::config::OutputFormat::Json,
    max_pages: 1000,
    ..Default::default()
};

let parser = LiteParse::new(config);

for entry in walkdir::WalkDir::new("./batch-input")
    .into_iter()
    .filter_map(Result::ok)
    .filter(|e| e.path().extension() == Some(std::ffi::OsStr::new("pdf")))
{
    let input = PdfInput::Path(entry.path().to_string_lossy().into_owned());
    let result = parser.parse(&input).await?;
    let json = liteparse::output::json::format_json(&result.pages)?;
    let out_path = entry.path()
        .strip_prefix("./batch-input")?
        .with_extension("json");
    std::fs::create_dir_all(out_path.parent().unwrap())?;
    std::fs::write(out_path, json)?;
}

This approach mirrors the CLI logic while providing full control over file discovery and error handling.

Node.js Wrapper

The Node.js API provides a high-level batchParse method:

import { LiteParse } from "liteparse";

const lp = new LiteParse({ outputFormat: "json", ocrEnabled: false });

await lp.batchParse({
  inputDir: "src/pdfs",
  outputDir: "out/json",
  recursive: true,
  extension: ".pdf",
});

The wrapper forwards options to the native binary, using the same underlying Rust implementation described in crates/liteparse/src/main.rs.

Key Source Files and Architecture

Understanding the codebase structure helps when extending batch functionality:

Summary

  • LiteParse's batch-parse functionality processes entire directories through a single command or API call.
  • The BatchParseCommand struct in crates/liteparse/src/main.rs handles CLI option parsing for input directories, output formats, and filtering.
  • The collect_files helper discovers documents recursively while batch_output_path preserves directory structure in outputs.
  • A single LiteParse instance is reused across all files, keeping OCR workers and heavy resources allocated once for efficiency.
  • Output formats include JSON, plain text, and Markdown, implemented in dedicated modules under crates/liteparse/src/output/.
  • The command exits with non-zero status if any file fails, making it suitable for CI/CD pipelines and automated scripts.

Frequently Asked Questions

What file types does LiteParse batch-parse support?

LiteParse supports PDFs, DOCXs, and various image formats in batch mode. The collect_files helper can filter by extension using the --extension flag, allowing you to target specific document types within mixed directories while ignoring unsupported files.

How does LiteParse handle errors during batch processing?

The CLI tracks successes and failures throughout the parsing loop. If any document fails to parse, the command prints the error, continues processing remaining files, and exits with a non-zero status code. This behavior makes the command suitable for shell scripts and CI/CD pipelines that need to detect processing failures.

Can I use OCR in batch mode?

Yes. OCR is controlled via the LiteParseConfig settings passed to the batch command. When enabled, the same OCR worker pool is initialized once and reused across all documents, minimizing the performance penalty of repeated initialization. Set ocr_enabled: true in Rust or ocrEnabled: true in Node.js to enable text extraction from images and scanned PDFs.

Is the batch-parse functionality available in the Node.js wrapper?

Yes. The Node.js wrapper exposes a batchParse method that accepts inputDir, outputDir, recursive, and extension parameters. This method delegates to the native Rust binary, providing identical functionality to the CLI without requiring developers to write Rust code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →