LiteParse Data Flow: From PdfInput Through Conversion to ParseResult
LiteParse transforms documents from initial PdfInput through an eight-stage pipeline—conversion, PDF loading, text extraction, optional OCR, spatial projection, and result serialization—to produce a structured ParseResult.
The run-llama/liteparse repository provides a Rust-native document parser that normalizes diverse file formats into machine-readable structures. Understanding the data flow from PdfInput through conversion to ParseResult reveals how the library handles format abstraction, text extraction, and spatial layout analysis. Each stage is implemented as a discrete module, allowing precise control over processing behavior.
Entry Point: LiteParse::parse in lib.rs
The public API surface resides in crates/liteparse/src/lib.rs, where the LiteParse::parse method accepts an Input type and delegates execution to the orchestrator in crates/liteparse/src/parser.rs. This method signature pub async fn parse(&self, input: Input) -> Result<ParseResult> defines the contract that all language bindings ultimately consume. The orchestrator coordinates the subsequent pipeline stages, beginning with format detection and conversion.
Format Conversion Layer (conversion.rs)
Non-PDF inputs undergo normalization in crates/liteparse/src/conversion.rs before entering the extraction pipeline. The Conversion::to_pdf method inspects file extensions and invokes external tools—LibreOffice for Office documents or ImageMagick for images—to generate a temporary PDF.
if !input_path.extension().map_or(false, |e| e == "pdf") {
// run external tool → temporary PDF
let tmp_pdf = self.run_external_converter(&input_path)?;
Ok(tmp_pdf)
} else {
Ok(input_path.to_path_buf())
}
This ensures that downstream components always process a valid PDF, regardless of the original input format.
PDF Loading and Raw Text Extraction
The pipeline loads the normalized PDF using the pdfium wrapper and extracts raw text elements.
PDFium Document Loading
Located in crates/pdfium/src/document.rs, the pdfium::Document::load function opens the PDF and provides access to page structures, text streams, and embedded images. This Rust binding to the PDFium library handles low-level PDF parsing, font decoding, and metadata retrieval.
Text Item Extraction (extract.rs)
The crates/liteparse/src/extract.rs module contains Extractor::extract_page, which iterates over every page and text run to build a flat vector of TextItem structs. Each TextItem captures the raw text string, bounding box coordinates, and font information.
for page in document.pages() {
let items = self.extract_page(page)?;
all_items.extend(items);
}
At this stage, the data consists of an unordered collection of text fragments with spatial coordinates.
Optional OCR Pipeline
When LiteParseConfig enables OCR, the system processes image-based content through a dedicated sub-pipeline to extract machine-readable text from visuals.
Page Rasterization (render.rs)
The crates/liteparse/src/render.rs module rasterizes PDF pages into PNG images, providing the bitmap data required by OCR engines. This step handles resolution scaling and color space conversion to optimize recognition accuracy.
OCR Engine Abstraction (ocr/mod.rs)
The crates/liteparse/src/ocr/mod.rs trait abstracts specific OCR implementations like ocr::Tesseract or ocr::HttpSimple. The engine processes rasterized images and returns text strings with bounding box annotations, which are converted back into TextItem instances.
Merging Native and OCR Text (ocr_merge.rs)
The crates/liteparse/src/ocr_merge.rs module combines native PDF text with OCR-derived results, removing duplicates and tagging items with their source type. This prevents double-counting text that exists both as PDF commands and as rendered glyphs in images.
Spatial Projection and Layout Analysis (projection.rs)
The flat list of TextItem instances receives structural organization in crates/liteparse/src/projection.rs. The Projector::project method implements spatial analysis algorithms that detect reading direction, column boundaries, and text rotation (0°, 90°, 180°, 270°).
The algorithm constructs forward anchors that preserve column alignment across lines, ensuring that tabular data and multi-column layouts resolve into correct reading order. This stage transforms coordinate-based text fragments into a semantically ordered stream.
Result Construction and Serialization (types.rs)
The final stage packages processed data into the ParseResult struct defined in crates/liteparse/src/types.rs. This structure contains:
pages: Vec<PageResult>storing per-page dimensions and rotation metadataitems: Vec<TextItem>containing the spatially ordered text stream- Optional OCR confidence scores for items derived from image recognition
Output Formatters
The crates/liteparse/src/output/json.rs and crates/liteparse/src/output/text.rs modules implement the Formatter trait to serialize ParseResult into JSON or plain text. Both formatters consume the ordered TextItem list and produce language-specific output suitable for downstream applications.
Practical Implementation Example
The following Rust example demonstrates the complete data flow from file path to structured result:
use liteparse::{LiteParse, LiteParseConfig};
#[tokio::main]
async fn main() -> anyhow::Result<()> {
// Configure parser with default settings
let config = LiteParseConfig::default();
let parser = LiteParse::new(config).await?;
// Triggers: conversion → PDFium → extract → projection → ParseResult
let result = parser.parse_path("document.docx").await?;
// Serialize to JSON
let json = serde_json::to_string_pretty(&result)?;
println!("{}", json);
Ok(())
}
This single parse_path call executes the full pipeline: format conversion (for .docx), PDF loading via pdfium, text extraction, spatial projection, and final ParseResult construction.
Summary
- Format normalization in
conversion.rsensures all inputs become PDFs before processing, using external tools for Office documents and images. - Text extraction relies on
pdfium::Documentandextract.rsto pull rawTextItemdata including bounding boxes and font metadata. - OCR integration through
render.rs,ocr/mod.rs, andocr_merge.rshandles image-based text extraction and deduplication. - Spatial projection in
projection.rsreorders flat text fragments into reading order using column detection and rotation analysis. - Result construction encapsulates ordered text, page metadata, and confidence scores in the
ParseResulttype fromtypes.rs.
Frequently Asked Questions
What happens if the input is already a PDF?
If the input carries a .pdf extension, Conversion::to_pdf in crates/liteparse/src/conversion.rs bypasses external conversion tools and returns the original path immediately. The pipeline proceeds directly to PDF loading via pdfium::Document::load without creating temporary files.
How does LiteParse determine when to trigger OCR?
OCR activation depends on the LiteParseConfig settings passed to LiteParse::new. When enabled, the pipeline automatically rasterizes pages through render.rs and processes them through the OCR engine whenever extract.rs detects image content or when the configuration mandates full-page OCR, merging results in ocr_merge.rs.
What information does the ParseResult struct contain?
According to crates/liteparse/src/types.rs, ParseResult contains a vector of PageResult objects with dimensions and rotation angles, a chronologically ordered vector of TextItem objects with text content and bounding boxes, and optional OCR confidence metrics indicating recognition certainty for specific text blocks.
How does spatial projection handle multi-column layouts?
The Projector::project implementation in crates/liteparse/src/projection.rs analyzes horizontal alignment patterns to construct forward anchors that track column positions across lines. This algorithm detects when text items align vertically with previous lines, preserving column relationships even when content wraps across pages or contains mixed rotations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →