How to Use LiteParse Programmatically in Python vs Rust: Complete API Guide

LiteParse exposes the same Rust-based parsing engine through idiomatic APIs in both languages, with Python offering synchronous convenience methods and Rust providing full async control.

To use LiteParse programmatically, you interact with the run-llama/liteparse repository’s core logic through language-specific bindings. Whether you choose Python for rapid prototyping or Rust for performance-critical applications, you access identical PDF extraction, OCR, and grid projection capabilities. Both implementations rely on the shared LiteParseConfig struct defined in the core Rust crate.

Architecture Overview

LiteParse follows a split architecture where the heavy lifting happens in Rust, and language bindings provide thin, ergonomic wrappers.

Core Rust Engine

All parsing logic lives in crates/liteparse/src/parser.rs, which handles PDF extraction, OCR merging, and coordinate projection. The LiteParse struct in this file exposes the primary parse_input async method and a convenience parse method for file paths. The core data structures—ParsedPage, TextItem, and ParseResult—are defined in crates/liteparse/src/types.rs and re-exported through crates/liteparse/src/lib.rs.

Python Bindings Layer

The Python interface resides in crates/liteparse-python/src/lib.rs. This wrapper translates Python method calls to the Rust core by running an internal Tokio runtime, making the API appear synchronous to Python callers. It converts ParseResult into PyParseResult, exposing Python-friendly attributes like pages, text, and num_pages.

Using LiteParse Programmatically in Python

The Python API prioritizes ease of use with synchronous execution while maintaining access to all configuration options.

First, install the wheel:

pip install liteparse

Then instantiate the parser and process documents:

from liteparse import LiteParse

# Configure the parser instance

parser = LiteParse(
    ocr_enabled=True,          # Requires tesseract or an OCR server

    dpi=300,                   # Rendering resolution for OCR

    output_format="json",      # Options: "json" or "text"

)

# parse() is synchronous from Python's perspective

result = parser.parse("example.pdf")

print("Full text:")
print(result.text)

print(f"\nTotal pages: {result.num_pages}")
for page in result.pages:
    print(f"Page {page.page_num}: {len(page.text_items)} items, dimensions {page.width}×{page.height}")

Under the hood, this calls PyLiteParse defined in crates/liteparse-python/src/lib.rs, which builds a LiteParseConfig and manages a Tokio runtime to execute the Rust async code.

Using LiteParse Programmatically in Rust

The Rust API provides direct async/await access to the parsing engine with zero overhead from language translation.

Add the dependency to your Cargo.toml:

[dependencies]
liteparse = "0.1"   # Use exact version from crates.io

tokio = { version = "1", features = ["full"] }

Implement async parsing:

use liteparse::{LiteParse, LiteParseConfig, OutputFormat};

#[tokio::main]
async fn main() -> Result<(), liteparse::LiteParseError> {
    // Build configuration using idiomatic Rust patterns
    let mut cfg = LiteParseConfig::default();
    cfg.ocr_enabled = true;
    cfg.dpi = 300.0;
    cfg.output_format = OutputFormat::Json;

    // Instantiate the parser
    let parser = LiteParse::new(cfg);

    // parse() returns a Future; .await yields the ParseResult
    let result = parser.parse("example.pdf").await?;

    println!("Extracted text:\n{}", result.text);

    for page in result.pages {
        println!(
            "Page {} – {} text items – size {:.2}×{:.2}",
            page.page_number,
            page.text_items.len(),
            page.page_width,
            page.page_height
        );
    }

    Ok(())
}

The LiteParse::new constructor and parse method are implemented in crates/liteparse/src/parser.rs at lines 44 and 71 respectively.

Key Differences Between APIs

While both languages access the same core functionality, their invocation patterns differ:

  • Execution Model: Python methods in liteparse-python/src/lib.rs block the caller while running a Tokio runtime internally; Rust exposes native async/await requiring an explicit runtime like tokio::main.
  • Configuration: Python accepts keyword arguments directly in the constructor; Rust requires mutating a LiteParseConfig struct before passing it to LiteParse::new.
  • Result Access: Python returns a PyParseResult with snake_case attributes (page_num, text_items); Rust returns ParseResult with fields matching the types.rs definitions (page_number).

Summary

  • LiteParse provides a unified parsing engine written in Rust, exposed through crates/liteparse/src/parser.rs.
  • Python users import from liteparse and call synchronous methods that wrap the Rust async runtime, defined in crates/liteparse-python/src/lib.rs.
  • Rust users interact directly with LiteParse and LiteParseConfig from crates/liteparse/src/lib.rs, using standard async/await patterns.
  • Both APIs support identical OCR, DPI, and output format configurations through the shared LiteParseConfig struct in crates/liteparse/src/config.rs.
  • Performance: Rust offers zero-cost async; Python incurs minimal overhead from the PyO3 binding layer and internal Tokio management.

Frequently Asked Questions

Does the Python API support asynchronous programming?

No, the Python binding in liteparse-python/src/lib.rs intentionally provides a synchronous interface. It manages an internal Tokio runtime to execute the underlying Rust async code, blocking until the parse completes. For async Python workflows, run LiteParse calls in a thread pool using asyncio.to_thread or similar utilities.

What OCR dependencies are required for programmatic usage?

When ocr_enabled is set to true, LiteParse requires either Tesseract OCR installed locally or access to an OCR server endpoint. The Rust core in parser.rs handles OCR merging logic in ocr_merge.rs, but the actual recognition depends on external binaries or services not bundled with the library.

How do I configure advanced parsing options programmatically?

Both APIs expose the same configuration struct. In Python, pass parameters like dpi, ocr_enabled, and output_format directly to the LiteParse constructor. In Rust, instantiate LiteParseConfig::default() and mutate fields before calling LiteParse::new(cfg). All options are validated in crates/liteparse/src/config.rs regardless of language.

Is there a performance difference between Python and Rust bindings?

Yes. The Rust implementation offers native performance with direct async/await support and no serialization overhead. The Python binding adds minimal latency through PyO3 translation and the internal Tokio runtime management, though the core PDF processing speed remains identical since both execute the same Rust code from parser.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →