LiteParse Batch-Parse: Processing Multiple Documents in One Command
LiteParse's batch-parse functionality allows you to process entire directories of documents (PDFs, DOCXs, images) through a single CLI command or API call, reusing a single LiteParse instance to minimize resource overhead while preserving directory structure in the output.
The run-llama/liteparse repository provides a dedicated batch-parse command that transforms document processing from a file-by-file operation into an efficient batch workflow. This functionality processes mixed document types automatically and outputs structured results without requiring manual orchestration of individual parse calls.
How Batch-Parse Works in LiteParse
The implementation lives in crates/liteparse/src/main.rs, where the batch-parse subcommand orchestrates the pipeline. The process follows six distinct stages that optimize for throughput and error resilience.
CLI Entry Point and Configuration
The BatchParseCommand struct defines all CLI flags including input directory, output directory, format selection, OCR toggles, DPI settings, and recursion options. This configuration is collected upfront in crates/liteparse/src/main.rs to avoid re-parsing arguments during the document loop.
File Discovery and Filtering
The helper function collect_files walks the input directory recursively when requested, returning a sorted Vec<String> of file paths. An optional extension filter applied during collection ensures only relevant documents enter the processing queue.
Output Path Resolution
For each input file, the batch_output_path function computes the corresponding output location by stripping the input directory prefix, preserving nested subfolder structures, and appending the appropriate extension based on the selected output format.
The Parsing Loop and Resource Management
A single LiteParse instance is instantiated with the batch-specific configuration before entering the loop. This approach reuses the OCR worker pool and other heavy resources across all documents, significantly reducing per-file overhead. The CLI iterates over collected files, invoking lp.parse(file_path).await for each entry, then formats the ParseResult according to the requested output format.
Command-Line Usage Examples
Use the lit batch-parse command to process documents without writing code.
Basic recursive parsing:
lit batch-parse \
--input-dir ./documents \
--output-dir ./out \
--format markdown \
--recursive \
--extension .pdf
This command scans ./documents recursively, extracts all PDFs, parses each with OCR enabled by default, and writes .md files mirroring the input directory structure under ./out.
Programmatic Batch Processing
Beyond the CLI, LiteParse exposes the same batch functionality through Rust and Node.js APIs.
Rust Implementation
When building custom pipelines, instantiate LiteParse once and reuse it across multiple files:
use liteparse::{LiteParse, config::LiteParseConfig, parser::PdfInput};
use std::path::Path;
let config = LiteParseConfig {
ocr_enabled: true,
ocr_language: "eng".into(),
output_format: liteparse::config::OutputFormat::Json,
max_pages: 1000,
..Default::default()
};
let parser = LiteParse::new(config);
for entry in walkdir::WalkDir::new("./batch-input")
.into_iter()
.filter_map(Result::ok)
.filter(|e| e.path().extension() == Some(std::ffi::OsStr::new("pdf")))
{
let input = PdfInput::Path(entry.path().to_string_lossy().into_owned());
let result = parser.parse(&input).await?;
let json = liteparse::output::json::format_json(&result.pages)?;
let out_path = entry.path()
.strip_prefix("./batch-input")?
.with_extension("json");
std::fs::create_dir_all(out_path.parent().unwrap())?;
std::fs::write(out_path, json)?;
}
This approach mirrors the CLI logic while providing full control over file discovery and error handling.
Node.js Wrapper
The Node.js API provides a high-level batchParse method:
import { LiteParse } from "liteparse";
const lp = new LiteParse({ outputFormat: "json", ocrEnabled: false });
await lp.batchParse({
inputDir: "src/pdfs",
outputDir: "out/json",
recursive: true,
extension: ".pdf",
});
The wrapper forwards options to the native binary, using the same underlying Rust implementation described in crates/liteparse/src/main.rs.
Key Source Files and Architecture
Understanding the codebase structure helps when extending batch functionality:
crates/liteparse/src/main.rs: ContainsBatchParseCommand,collect_files,batch_output_path, and the main orchestration loop.crates/liteparse/src/config.rs: DefinesLiteParseConfigwhich drives parsing behavior across batch operations.crates/liteparse/src/parser.rs: Houses the coreLiteParse::parsemethod reused for every document in the batch.crates/liteparse/src/output/json.rs: JSON formatter used when--format jsonis specified.crates/liteparse/src/output/text.rs: Plain-text formatter for--format text.crates/liteparse/src/output/markdown.rs: Markdown formatter for--format markdown.
Summary
- LiteParse's batch-parse functionality processes entire directories through a single command or API call.
- The
BatchParseCommandstruct incrates/liteparse/src/main.rshandles CLI option parsing for input directories, output formats, and filtering. - The
collect_fileshelper discovers documents recursively whilebatch_output_pathpreserves directory structure in outputs. - A single
LiteParseinstance is reused across all files, keeping OCR workers and heavy resources allocated once for efficiency. - Output formats include JSON, plain text, and Markdown, implemented in dedicated modules under
crates/liteparse/src/output/. - The command exits with non-zero status if any file fails, making it suitable for CI/CD pipelines and automated scripts.
Frequently Asked Questions
What file types does LiteParse batch-parse support?
LiteParse supports PDFs, DOCXs, and various image formats in batch mode. The collect_files helper can filter by extension using the --extension flag, allowing you to target specific document types within mixed directories while ignoring unsupported files.
How does LiteParse handle errors during batch processing?
The CLI tracks successes and failures throughout the parsing loop. If any document fails to parse, the command prints the error, continues processing remaining files, and exits with a non-zero status code. This behavior makes the command suitable for shell scripts and CI/CD pipelines that need to detect processing failures.
Can I use OCR in batch mode?
Yes. OCR is controlled via the LiteParseConfig settings passed to the batch command. When enabled, the same OCR worker pool is initialized once and reused across all documents, minimizing the performance penalty of repeated initialization. Set ocr_enabled: true in Rust or ocrEnabled: true in Node.js to enable text extraction from images and scanned PDFs.
Is the batch-parse functionality available in the Node.js wrapper?
Yes. The Node.js wrapper exposes a batchParse method that accepts inputDir, outputDir, recursive, and extension parameters. This method delegates to the native Rust binary, providing identical functionality to the CLI without requiring developers to write Rust code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →