JSON and Text Output Formats in LiteParse: Key Differences Explained

JSON output provides structured document metadata including bounding boxes and OCR confidence scores, while Text output returns a clean, human-readable plain-text string without formatting metadata.

LiteParse, the open-source document parsing engine maintained by the run-llama organization, exposes two distinct serialization formats through its LiteParseConfig structure. Understanding the difference between JSON and Text output formats in LiteParse is essential for selecting the right representation—whether your pipeline requires rich spatial metadata for programmatic analysis or simple extracted text for immediate consumption.

Structural Differences: Hierarchical Data vs. Plain Text

JSON: Rich, Programmatic Structure

When output_format is set to OutputFormat::Json in crates/liteparse/src/config.rs, the format_json function in crates/liteparse/src/output/json.rs serializes the parsed document as a fully structured JSON object. The hierarchy follows ParsedResult → pages[] → items[], where each text item includes:

  • bbox: Bounding box coordinates for spatial positioning
  • confidence: OCR confidence levels for quality auditing
  • source: Flags distinguishing native text from OCR-generated content

This structure enables downstream applications to reconstruct document layout, filter low-confidence OCR results, or render visual overlays.

Text: Flat, Human-Readable Output

Selecting OutputFormat::Text triggers format_text in crates/liteparse/src/output/text.rs, which traverses the same internal ParsedPage list but simply joins the text fields with \n characters. The result is a single plain-text string that preserves natural reading order with line breaks between lines and pages, but deliberately omits all metadata, coordinates, and confidence scores.

Implementation and Dispatch Flow

The core dispatch logic resides in crates/liteparse/src/main.rs, where the CLI match arm routes to the appropriate formatter:

match config.output_format {
    OutputFormat::Json => json::format_json(&result.pages)?,
    OutputFormat::Text => text::format_text(&result.pages),
}

Both formatters operate on the identical internal representation of parsed pages, ensuring consistent extraction quality regardless of output format. The JSON path uses serde_json for serialization, while the Text path performs simple string concatenation.

Configuration Examples

Command Line Interface

Use the --output-format flag to select your desired representation:

liteparse input.pdf --output-format json > result.json
liteparse input.pdf --output-format text > result.txt

Node.js Binding

import { LiteParse } from "liteparse";

const jsonParser = new LiteParse({ output_format: "json" });
const jsonResult = await jsonParser.parse("sample.pdf");  // JSON string

const textParser = new LiteParse({ output_format: "text" });
const textResult = await textParser.parse("sample.pdf"); // Plain text

Python Binding

from liteparse import LiteParse

json_parser = LiteParse(output_format="json")
json_output = json_parser.parse("sample.pdf")   # JSON string

text_parser = LiteParse(output_format="text")
text_output = text_parser.parse("sample.pdf")  # Plain text

WASM (Browser) Binding

import { LiteParse } from "liteparse-wasm";

const parser = new LiteParse({ output_format: "json" });
const json = await parser.parse(file);   // Returns JSON string

Use Case Recommendations

Choose JSON when building document understanding pipelines that require spatial analysis, table extraction, or OCR quality auditing. The bounding box and confidence metadata enable precise layout reconstruction and filtering of uncertain text regions.

Choose Text when performing quick human inspection, piping content to standard Unix tools like grep or awk, or feeding downstream NLP models that only require raw text without spatial context. This format minimizes file size and parsing overhead.

Summary

  • JSON output in LiteParse delivers hierarchical structured data via format_json in crates/liteparse/src/output/json.rs, exposing bounding boxes, OCR confidence levels, and source flags for each text fragment.
  • Text output produces a flat, concatenated string via format_text in crates/liteparse/src/output/text.rs, optimized for readability and simple command-line pipelines.
  • Both formats are configured through the output_format field in LiteParseConfig and dispatched from crates/liteparse/src/main.rs, with consistent API exposure across Node.js, Python, and WASM bindings.
  • JSON suits metadata-rich programmatic workflows requiring spatial analysis, while Text is ideal for human inspection and text-only processing tasks.

Frequently Asked Questions

Can I convert between JSON and Text output without re-parsing the document?

No, you must specify the desired output_format before invoking the parser. While both format_json and format_text consume the same internal ParsedPage structures, the serialization occurs during the formatting phase, and the lightweight Text format deliberately discards metadata that cannot be recovered.

Does the Text output preserve document structure like paragraphs and columns?

The Text format maintains natural reading order and inserts line breaks between logical lines and pages, but it strips all structural metadata. Headers, paragraphs, and table cells appear as sequential text blocks without semantic markup or spatial coordinates that would enable layout reconstruction.

Which format is more performant for large document batches?

Text generation is marginally faster due to simple string concatenation in text.rs, while JSON serialization in json.rs using serde_json incurs additional overhead for struct traversal and metadata encoding. However, both formatting operations are typically dwarfed by the actual PDF parsing and OCR processing time.

Can I extract table data using the Text format?

While the Text format will output tabular content as sequential text, you should use JSON for reliable table extraction. The bbox coordinates and item hierarchy in JSON output enable precise reconstruction of row and column relationships that are lost when content is flattened to plain text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →