How to Use the LiteParse Rust API Directly for Custom Applications

The LiteParse Rust API centers on the LiteParse struct in src/parser.rs, which accepts a LiteParseConfig and provides async methods parse() for files and parse_input() for bytes, returning structured ParseResult data with full thread safety.

The run-llama/liteparse crate exposes a concise yet powerful Rust API for document parsing and OCR tasks. By interacting directly with the liteparse core crate, you can build custom applications that process PDFs from disk or memory, configure OCR engines, and extract structured text with spatial layout preservation.

Core Architecture and Configuration Types

The public API is organized around three primary types defined in the liteparse crate:

  • LiteParse in src/parser.rs — The main orchestrator struct that handles the parsing lifecycle
  • LiteParseConfig in src/config.rs — Configuration container for all tunable options including OCR language, DPI, and output format
  • ParseResult in src/parser.rs — The structured output containing pages, text, outlines, and images

Configuration defaults are sensible: OCR language is "eng", output format is JSON, and image handling is enabled via embedding modes.

Initializing the Parser with LiteParseConfig

Create a parser by first configuring LiteParseConfig, then instantiating LiteParse with LiteParse::new(config).

For basic file parsing with default settings:

use liteparse::config::LiteParseConfig;
use liteparse::LiteParse;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Use the default configuration (JSON output, OCR enabled, etc.)
    let config = LiteParseConfig::default();

    // Create the parser
    let parser = LiteParse::new(config);

    // Parse a file on disk
    let result = parser.parse("example.pdf").await?;

    // The concatenated text is in `result.text`
    println!("{}", result.text);
    Ok(())
}

Key implementation details from src/parser.rs:

  • LiteParse::new(config) takes ownership of the configuration
  • The struct implements Send + Sync, allowing safe sharing across threads
  • PDFium operations are internally serialized by a global lock, while OCR and projection run concurrently

Parsing from Different Input Sources

The API supports two primary input methods defined in src/parser.rs: parse() for filesystem paths and parse_input() for abstraction over bytes.

Processing Files from Disk

The parse(&self, path: &str) method provides a convenience wrapper for on-disk files (non-WASM builds). It handles input resolution including non-PDF conversion and password protection as implemented in src/parser.rs lines 84-103.

Handling In-Memory PDF Bytes

For web services or testing scenarios, use parse_input() with PdfInput::Bytes to avoid filesystem I/O entirely. This is defined in src/types.rs and accepts raw PDF bytes directly.

use liteparse::config::LiteParseConfig;
use liteparse::types::PdfInput;
use liteparse::LiteParse;
use std::fs;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Load PDF data into memory (e.g., read from a network request)
    let pdf_bytes = fs::read("report.pdf")?;

    // Configuration with Markdown output and embedded images
    let mut cfg = LiteParseConfig::default();
    cfg.output_format = liteparse::config::OutputFormat::Markdown;
    cfg.image_mode = liteparse::config::ImageMode::Embed;

    let parser = LiteParse::new(cfg);

    // Parse from raw bytes
    let result = parser
        .parse_input(PdfInput::Bytes(pdf_bytes))
        .await?;

    // The markdown representation includes image references like ![](image_p1_0.png)
    println!("{}", result.text);
    Ok(())
}

Advanced Customization

Integrating Custom OCR Engines

Override the built-in OCR selection logic by chaining .with_ocr_engine() before parsing. According to src/parser.rs lines 68-82, this method accepts any Arc<dyn OcrEngine> and forces the parser to use your implementation regardless of the tesseract feature or ocr_server_url configuration.

use std::sync::Arc;
use liteparse::config::LiteParseConfig;
use liteparse::ocr::http_simple::HttpOcrEngine;
use liteparse::LiteParse;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut cfg = LiteParseConfig::default();
    cfg.ocr_enabled = true;               // ensure OCR runs
    cfg.ocr_server_url = Some(
        "http://my-ocr.example.com/ocr".into()
    );

    // Build the HTTP OCR engine (adds any needed headers)
    let http_engine = Arc::new(HttpOcrEngine::with_headers(
        cfg.ocr_server_url.clone().unwrap(),
        vec![],
    ));

    // Override the built-in selection logic
    let parser = LiteParse::new(cfg).with_ocr_engine(http_engine);

    let result = parser.parse("scanned.pdf").await?;
    println!("Extracted text (OCR):\n{}", result.text);
    Ok(())
}

Selective Page Extraction

Configure target_pages in LiteParseConfig to parse only specific ranges, implemented in src/config.rs. This accepts strings like "1-3,7" (1-based indexing) and validates them via the internal parse_target_pages logic.

use liteparse::config::{LiteParseConfig, OutputFormat};
use liteparse::LiteParse;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut cfg = LiteParseConfig::default();
    cfg.target_pages = Some("1-3,7".into()); // page numbers are 1‑based
    cfg.output_format = OutputFormat::Text;

    let parser = LiteParse::new(cfg);
    let res = parser.parse("big_document.pdf").await?;
    println!("{}", res.text);
    Ok(())
}

Understanding the Parsing Pipeline and Results

The ParseResult struct returned by both parsing methods contains:

  • pages — A vector of ParsedPage objects with spatially projected text items
  • text — The full concatenated representation (JSON, plain text, or Markdown depending on output_format)
  • outline — Document bookmarks extracted via PDFium, useful for hierarchical navigation
  • images — Raster images when ImageMode::Embed is enabled

The internal flow implemented across src/extract.rs, src/ocr_merge.rs, and src/projection.rs follows five stages:

  1. Input Resolution — Handles non-PDF conversion and password protection
  2. PDFium Extraction — Pulls raw text items, images, and outline data
  3. Optional OCR — Renders pages and merges results back into the text layout via the OCR engine interface
  4. Grid Projection — Reconstructs reading order using the spatial algorithm in src/projection.rs
  5. Formatting — Serializes output via modules in src/output/

Error handling uses LiteParseError, which implements std::error::Error for seamless integration with Rust error handling patterns.

Summary

  • The LiteParse Rust API entry point is the LiteParse struct in src/parser.rs, configured via LiteParseConfig from src/config.rs
  • Use LiteParse::new(config) to create instances, and with_ocr_engine() to override default OCR selection
  • Parse files with parse() or raw bytes with parse_input() using PdfInput::Bytes from src/types.rs
  • The ParseResult struct provides structured access to pages, text, outlines, and embedded images
  • The API is fully thread-safe (Send + Sync), making it suitable for concurrent processing in tokio task pools or multi-threaded applications
  • All errors are wrapped in LiteParseError for consistent error handling

Frequently Asked Questions

What is the entry point for the LiteParse Rust API?

The entry point is the LiteParse struct defined in src/parser.rs. You create an instance by calling LiteParse::new(config) with a LiteParseConfig object, which encapsulates all parsing parameters including OCR settings and output formats.

How do I parse PDFs from memory instead of disk?

Use the parse_input() method with PdfInput::Bytes(pdf_bytes) instead of parse(). This variant accepts raw byte vectors and is defined in src/parser.rs, avoiding filesystem I/O entirely. This is particularly useful for web servers processing uploaded files.

Is LiteParse thread-safe for concurrent processing?

Yes, LiteParse implements both Send and Sync traits. You can safely share a single instance across multiple threads, such as inside a tokio task pool. The implementation serializes PDFium access via a global lock while allowing OCR and projection phases to run concurrently.

Can I replace the default Tesseract OCR with a custom implementation?

Yes, chain .with_ocr_engine(my_engine) when building the LiteParse instance, passing an Arc<dyn OcrEngine>. This overrides the automatic engine selection logic in src/parser.rs and forces the parser to use your custom OCR implementation, whether it's an HTTP service or a local alternative.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →