# LiteParse Batch-Parse: Processing Multiple Documents in One Command

> Streamline document processing with LiteParse batch-parse. Process multiple PDFs DOCXs and images with one command or API call minimizing resource overhead.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-25

---

**LiteParse's batch-parse functionality allows you to process entire directories of documents (PDFs, DOCXs, images) through a single CLI command or API call, reusing a single `LiteParse` instance to minimize resource overhead while preserving directory structure in the output.**

The run-llama/liteparse repository provides a dedicated batch-parse command that transforms document processing from a file-by-file operation into an efficient batch workflow. This functionality processes mixed document types automatically and outputs structured results without requiring manual orchestration of individual parse calls.

## How Batch-Parse Works in LiteParse

The implementation lives in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs), where the `batch-parse` subcommand orchestrates the pipeline. The process follows six distinct stages that optimize for throughput and error resilience.

### CLI Entry Point and Configuration

The `BatchParseCommand` struct defines all CLI flags including input directory, output directory, format selection, OCR toggles, DPI settings, and recursion options. This configuration is collected upfront in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) to avoid re-parsing arguments during the document loop.

### File Discovery and Filtering

The helper function `collect_files` walks the input directory recursively when requested, returning a sorted `Vec<String>` of file paths. An optional extension filter applied during collection ensures only relevant documents enter the processing queue.

### Output Path Resolution

For each input file, the `batch_output_path` function computes the corresponding output location by stripping the input directory prefix, preserving nested subfolder structures, and appending the appropriate extension based on the selected output format.

### The Parsing Loop and Resource Management

A single `LiteParse` instance is instantiated with the batch-specific configuration before entering the loop. This approach reuses the OCR worker pool and other heavy resources across all documents, significantly reducing per-file overhead. The CLI iterates over collected files, invoking `lp.parse(file_path).await` for each entry, then formats the `ParseResult` according to the requested output format.

## Command-Line Usage Examples

Use the `lit batch-parse` command to process documents without writing code.

Basic recursive parsing:

```bash
lit batch-parse \
    --input-dir ./documents \
    --output-dir ./out \
    --format markdown \
    --recursive \
    --extension .pdf

```

This command scans `./documents` recursively, extracts all PDFs, parses each with OCR enabled by default, and writes `.md` files mirroring the input directory structure under `./out`.

## Programmatic Batch Processing

Beyond the CLI, LiteParse exposes the same batch functionality through Rust and Node.js APIs.

### Rust Implementation

When building custom pipelines, instantiate `LiteParse` once and reuse it across multiple files:

```rust
use liteparse::{LiteParse, config::LiteParseConfig, parser::PdfInput};
use std::path::Path;

let config = LiteParseConfig {
    ocr_enabled: true,
    ocr_language: "eng".into(),
    output_format: liteparse::config::OutputFormat::Json,
    max_pages: 1000,
    ..Default::default()
};

let parser = LiteParse::new(config);

for entry in walkdir::WalkDir::new("./batch-input")
    .into_iter()
    .filter_map(Result::ok)
    .filter(|e| e.path().extension() == Some(std::ffi::OsStr::new("pdf")))
{
    let input = PdfInput::Path(entry.path().to_string_lossy().into_owned());
    let result = parser.parse(&input).await?;
    let json = liteparse::output::json::format_json(&result.pages)?;
    let out_path = entry.path()
        .strip_prefix("./batch-input")?
        .with_extension("json");
    std::fs::create_dir_all(out_path.parent().unwrap())?;
    std::fs::write(out_path, json)?;
}

```

This approach mirrors the CLI logic while providing full control over file discovery and error handling.

### Node.js Wrapper

The Node.js API provides a high-level `batchParse` method:

```ts
import { LiteParse } from "liteparse";

const lp = new LiteParse({ outputFormat: "json", ocrEnabled: false });

await lp.batchParse({
  inputDir: "src/pdfs",
  outputDir: "out/json",
  recursive: true,
  extension: ".pdf",
});

```

The wrapper forwards options to the native binary, using the same underlying Rust implementation described in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs).

## Key Source Files and Architecture

Understanding the codebase structure helps when extending batch functionality:

- **[`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs)**: Contains `BatchParseCommand`, `collect_files`, `batch_output_path`, and the main orchestration loop.
- **[`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs)**: Defines `LiteParseConfig` which drives parsing behavior across batch operations.
- **[`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs)**: Houses the core `LiteParse::parse` method reused for every document in the batch.
- **[`crates/liteparse/src/output/json.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/json.rs)**: JSON formatter used when `--format json` is specified.
- **[`crates/liteparse/src/output/text.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/text.rs)**: Plain-text formatter for `--format text`.
- **[`crates/liteparse/src/output/markdown.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/markdown.rs)**: Markdown formatter for `--format markdown`.

## Summary

- LiteParse's batch-parse functionality processes entire directories through a single command or API call.
- The `BatchParseCommand` struct in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) handles CLI option parsing for input directories, output formats, and filtering.
- The `collect_files` helper discovers documents recursively while `batch_output_path` preserves directory structure in outputs.
- A single `LiteParse` instance is reused across all files, keeping OCR workers and heavy resources allocated once for efficiency.
- Output formats include JSON, plain text, and Markdown, implemented in dedicated modules under `crates/liteparse/src/output/`.
- The command exits with non-zero status if any file fails, making it suitable for CI/CD pipelines and automated scripts.

## Frequently Asked Questions

### What file types does LiteParse batch-parse support?

LiteParse supports PDFs, DOCXs, and various image formats in batch mode. The `collect_files` helper can filter by extension using the `--extension` flag, allowing you to target specific document types within mixed directories while ignoring unsupported files.

### How does LiteParse handle errors during batch processing?

The CLI tracks successes and failures throughout the parsing loop. If any document fails to parse, the command prints the error, continues processing remaining files, and exits with a non-zero status code. This behavior makes the command suitable for shell scripts and CI/CD pipelines that need to detect processing failures.

### Can I use OCR in batch mode?

Yes. OCR is controlled via the `LiteParseConfig` settings passed to the batch command. When enabled, the same OCR worker pool is initialized once and reused across all documents, minimizing the performance penalty of repeated initialization. Set `ocr_enabled: true` in Rust or `ocrEnabled: true` in Node.js to enable text extraction from images and scanned PDFs.

### Is the batch-parse functionality available in the Node.js wrapper?

Yes. The Node.js wrapper exposes a `batchParse` method that accepts `inputDir`, `outputDir`, `recursive`, and `extension` parameters. This method delegates to the native Rust binary, providing identical functionality to the CLI without requiring developers to write Rust code.