JSON and Text Output Formats in LiteParse: Key Differences Explained
JSON output provides structured document metadata including bounding boxes and OCR confidence scores, while Text output returns a clean, human-readable plain-text string without formatting metadata.
LiteParse, the open-source document parsing engine maintained by the run-llama organization, exposes two distinct serialization formats through its LiteParseConfig structure. Understanding the difference between JSON and Text output formats in LiteParse is essential for selecting the right representation—whether your pipeline requires rich spatial metadata for programmatic analysis or simple extracted text for immediate consumption.
Structural Differences: Hierarchical Data vs. Plain Text
JSON: Rich, Programmatic Structure
When output_format is set to OutputFormat::Json in crates/liteparse/src/config.rs, the format_json function in crates/liteparse/src/output/json.rs serializes the parsed document as a fully structured JSON object. The hierarchy follows ParsedResult → pages[] → items[], where each text item includes:
bbox: Bounding box coordinates for spatial positioningconfidence: OCR confidence levels for quality auditingsource: Flags distinguishing native text from OCR-generated content
This structure enables downstream applications to reconstruct document layout, filter low-confidence OCR results, or render visual overlays.
Text: Flat, Human-Readable Output
Selecting OutputFormat::Text triggers format_text in crates/liteparse/src/output/text.rs, which traverses the same internal ParsedPage list but simply joins the text fields with \n characters. The result is a single plain-text string that preserves natural reading order with line breaks between lines and pages, but deliberately omits all metadata, coordinates, and confidence scores.
Implementation and Dispatch Flow
The core dispatch logic resides in crates/liteparse/src/main.rs, where the CLI match arm routes to the appropriate formatter:
match config.output_format {
OutputFormat::Json => json::format_json(&result.pages)?,
OutputFormat::Text => text::format_text(&result.pages),
}
Both formatters operate on the identical internal representation of parsed pages, ensuring consistent extraction quality regardless of output format. The JSON path uses serde_json for serialization, while the Text path performs simple string concatenation.
Configuration Examples
Command Line Interface
Use the --output-format flag to select your desired representation:
liteparse input.pdf --output-format json > result.json
liteparse input.pdf --output-format text > result.txt
Node.js Binding
import { LiteParse } from "liteparse";
const jsonParser = new LiteParse({ output_format: "json" });
const jsonResult = await jsonParser.parse("sample.pdf"); // JSON string
const textParser = new LiteParse({ output_format: "text" });
const textResult = await textParser.parse("sample.pdf"); // Plain text
Python Binding
from liteparse import LiteParse
json_parser = LiteParse(output_format="json")
json_output = json_parser.parse("sample.pdf") # JSON string
text_parser = LiteParse(output_format="text")
text_output = text_parser.parse("sample.pdf") # Plain text
WASM (Browser) Binding
import { LiteParse } from "liteparse-wasm";
const parser = new LiteParse({ output_format: "json" });
const json = await parser.parse(file); // Returns JSON string
Use Case Recommendations
Choose JSON when building document understanding pipelines that require spatial analysis, table extraction, or OCR quality auditing. The bounding box and confidence metadata enable precise layout reconstruction and filtering of uncertain text regions.
Choose Text when performing quick human inspection, piping content to standard Unix tools like grep or awk, or feeding downstream NLP models that only require raw text without spatial context. This format minimizes file size and parsing overhead.
Summary
- JSON output in LiteParse delivers hierarchical structured data via
format_jsonincrates/liteparse/src/output/json.rs, exposing bounding boxes, OCR confidence levels, and source flags for each text fragment. - Text output produces a flat, concatenated string via
format_textincrates/liteparse/src/output/text.rs, optimized for readability and simple command-line pipelines. - Both formats are configured through the
output_formatfield inLiteParseConfigand dispatched fromcrates/liteparse/src/main.rs, with consistent API exposure across Node.js, Python, and WASM bindings. - JSON suits metadata-rich programmatic workflows requiring spatial analysis, while Text is ideal for human inspection and text-only processing tasks.
Frequently Asked Questions
Can I convert between JSON and Text output without re-parsing the document?
No, you must specify the desired output_format before invoking the parser. While both format_json and format_text consume the same internal ParsedPage structures, the serialization occurs during the formatting phase, and the lightweight Text format deliberately discards metadata that cannot be recovered.
Does the Text output preserve document structure like paragraphs and columns?
The Text format maintains natural reading order and inserts line breaks between logical lines and pages, but it strips all structural metadata. Headers, paragraphs, and table cells appear as sequential text blocks without semantic markup or spatial coordinates that would enable layout reconstruction.
Which format is more performant for large document batches?
Text generation is marginally faster due to simple string concatenation in text.rs, while JSON serialization in json.rs using serde_json incurs additional overhead for struct traversal and metadata encoding. However, both formatting operations are typically dwarfed by the actual PDF parsing and OCR processing time.
Can I extract table data using the Text format?
While the Text format will output tabular content as sequential text, you should use JSON for reliable table extraction. The bbox coordinates and item hierarchy in JSON output enable precise reconstruction of row and column relationships that are lost when content is flattened to plain text.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →