# How to Use the LiteParse Rust API Directly for Custom Applications

> Learn to use the LiteParse Rust API directly in your custom applications. Explore the LiteParse struct and its async parse methods for efficient file and byte data processing.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-25

---

**The LiteParse Rust API centers on the `LiteParse` struct in [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs), which accepts a `LiteParseConfig` and provides async methods `parse()` for files and `parse_input()` for bytes, returning structured `ParseResult` data with full thread safety.**

The `run-llama/liteparse` crate exposes a concise yet powerful Rust API for document parsing and OCR tasks. By interacting directly with the `liteparse` core crate, you can build custom applications that process PDFs from disk or memory, configure OCR engines, and extract structured text with spatial layout preservation.

## Core Architecture and Configuration Types

The public API is organized around three primary types defined in the `liteparse` crate:

- `LiteParse` in [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs) — The main orchestrator struct that handles the parsing lifecycle
- `LiteParseConfig` in [`src/config.rs`](https://github.com/run-llama/liteparse/blob/main/src/config.rs) — Configuration container for all tunable options including OCR language, DPI, and output format
- `ParseResult` in [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs) — The structured output containing pages, text, outlines, and images

Configuration defaults are sensible: OCR language is `"eng"`, output format is JSON, and image handling is enabled via embedding modes.

## Initializing the Parser with LiteParseConfig

Create a parser by first configuring `LiteParseConfig`, then instantiating `LiteParse` with `LiteParse::new(config)`.

For basic file parsing with default settings:

```rust
use liteparse::config::LiteParseConfig;
use liteparse::LiteParse;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Use the default configuration (JSON output, OCR enabled, etc.)
    let config = LiteParseConfig::default();

    // Create the parser
    let parser = LiteParse::new(config);

    // Parse a file on disk
    let result = parser.parse("example.pdf").await?;

    // The concatenated text is in `result.text`
    println!("{}", result.text);
    Ok(())
}

```

Key implementation details from [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs):

- `LiteParse::new(config)` takes ownership of the configuration
- The struct implements `Send + Sync`, allowing safe sharing across threads
- PDFium operations are internally serialized by a global lock, while OCR and projection run concurrently

## Parsing from Different Input Sources

The API supports two primary input methods defined in [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs): `parse()` for filesystem paths and `parse_input()` for abstraction over bytes.

### Processing Files from Disk

The `parse(&self, path: &str)` method provides a convenience wrapper for on-disk files (non-WASM builds). It handles input resolution including non-PDF conversion and password protection as implemented in [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs) lines 84-103.

### Handling In-Memory PDF Bytes

For web services or testing scenarios, use `parse_input()` with `PdfInput::Bytes` to avoid filesystem I/O entirely. This is defined in [`src/types.rs`](https://github.com/run-llama/liteparse/blob/main/src/types.rs) and accepts raw PDF bytes directly.

```rust
use liteparse::config::LiteParseConfig;
use liteparse::types::PdfInput;
use liteparse::LiteParse;
use std::fs;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Load PDF data into memory (e.g., read from a network request)
    let pdf_bytes = fs::read("report.pdf")?;

    // Configuration with Markdown output and embedded images
    let mut cfg = LiteParseConfig::default();
    cfg.output_format = liteparse::config::OutputFormat::Markdown;
    cfg.image_mode = liteparse::config::ImageMode::Embed;

    let parser = LiteParse::new(cfg);

    // Parse from raw bytes
    let result = parser
        .parse_input(PdfInput::Bytes(pdf_bytes))
        .await?;

    // The markdown representation includes image references like ![](image_p1_0.png)
    println!("{}", result.text);
    Ok(())
}

```

## Advanced Customization

### Integrating Custom OCR Engines

Override the built-in OCR selection logic by chaining `.with_ocr_engine()` before parsing. According to [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs) lines 68-82, this method accepts any `Arc<dyn OcrEngine>` and forces the parser to use your implementation regardless of the `tesseract` feature or `ocr_server_url` configuration.

```rust
use std::sync::Arc;
use liteparse::config::LiteParseConfig;
use liteparse::ocr::http_simple::HttpOcrEngine;
use liteparse::LiteParse;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut cfg = LiteParseConfig::default();
    cfg.ocr_enabled = true;               // ensure OCR runs
    cfg.ocr_server_url = Some(
        "http://my-ocr.example.com/ocr".into()
    );

    // Build the HTTP OCR engine (adds any needed headers)
    let http_engine = Arc::new(HttpOcrEngine::with_headers(
        cfg.ocr_server_url.clone().unwrap(),
        vec![],
    ));

    // Override the built-in selection logic
    let parser = LiteParse::new(cfg).with_ocr_engine(http_engine);

    let result = parser.parse("scanned.pdf").await?;
    println!("Extracted text (OCR):\n{}", result.text);
    Ok(())
}

```

### Selective Page Extraction

Configure `target_pages` in `LiteParseConfig` to parse only specific ranges, implemented in [`src/config.rs`](https://github.com/run-llama/liteparse/blob/main/src/config.rs). This accepts strings like `"1-3,7"` (1-based indexing) and validates them via the internal `parse_target_pages` logic.

```rust
use liteparse::config::{LiteParseConfig, OutputFormat};
use liteparse::LiteParse;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut cfg = LiteParseConfig::default();
    cfg.target_pages = Some("1-3,7".into()); // page numbers are 1‑based
    cfg.output_format = OutputFormat::Text;

    let parser = LiteParse::new(cfg);
    let res = parser.parse("big_document.pdf").await?;
    println!("{}", res.text);
    Ok(())
}

```

## Understanding the Parsing Pipeline and Results

The `ParseResult` struct returned by both parsing methods contains:

- `pages` — A vector of `ParsedPage` objects with spatially projected text items
- `text` — The full concatenated representation (JSON, plain text, or Markdown depending on `output_format`)
- `outline` — Document bookmarks extracted via PDFium, useful for hierarchical navigation
- `images` — Raster images when `ImageMode::Embed` is enabled

The internal flow implemented across [`src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/src/extract.rs), [`src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/src/ocr_merge.rs), and [`src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/src/projection.rs) follows five stages:

1. **Input Resolution** — Handles non-PDF conversion and password protection
2. **PDFium Extraction** — Pulls raw text items, images, and outline data
3. **Optional OCR** — Renders pages and merges results back into the text layout via the OCR engine interface
4. **Grid Projection** — Reconstructs reading order using the spatial algorithm in [`src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/src/projection.rs)
5. **Formatting** — Serializes output via modules in `src/output/`

Error handling uses `LiteParseError`, which implements `std::error::Error` for seamless integration with Rust error handling patterns.

## Summary

- The **LiteParse Rust API** entry point is the `LiteParse` struct in [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs), configured via `LiteParseConfig` from [`src/config.rs`](https://github.com/run-llama/liteparse/blob/main/src/config.rs)
- Use **`LiteParse::new(config)`** to create instances, and **`with_ocr_engine()`** to override default OCR selection
- Parse files with **`parse()`** or raw bytes with **`parse_input()`** using `PdfInput::Bytes` from [`src/types.rs`](https://github.com/run-llama/liteparse/blob/main/src/types.rs)
- The **`ParseResult`** struct provides structured access to pages, text, outlines, and embedded images
- The API is fully **thread-safe** (`Send + Sync`), making it suitable for concurrent processing in `tokio` task pools or multi-threaded applications
- All errors are wrapped in **`LiteParseError`** for consistent error handling

## Frequently Asked Questions

### What is the entry point for the LiteParse Rust API?

The entry point is the `LiteParse` struct defined in [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs). You create an instance by calling `LiteParse::new(config)` with a `LiteParseConfig` object, which encapsulates all parsing parameters including OCR settings and output formats.

### How do I parse PDFs from memory instead of disk?

Use the `parse_input()` method with `PdfInput::Bytes(pdf_bytes)` instead of `parse()`. This variant accepts raw byte vectors and is defined in [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs), avoiding filesystem I/O entirely. This is particularly useful for web servers processing uploaded files.

### Is LiteParse thread-safe for concurrent processing?

Yes, `LiteParse` implements both `Send` and `Sync` traits. You can safely share a single instance across multiple threads, such as inside a `tokio` task pool. The implementation serializes PDFium access via a global lock while allowing OCR and projection phases to run concurrently.

### Can I replace the default Tesseract OCR with a custom implementation?

Yes, chain `.with_ocr_engine(my_engine)` when building the `LiteParse` instance, passing an `Arc<dyn OcrEngine>`. This overrides the automatic engine selection logic in [`src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/src/parser.rs) and forces the parser to use your custom OCR implementation, whether it's an HTTP service or a local alternative.