How to Use LiteParse from Python with PyO3 Bindings: A Complete Guide

LiteParse exposes its Rust PDF parsing engine to Python through PyO3 bindings, providing a LiteParse class that instantiates a native Rust struct and converts results into Python dataclasses for text extraction, OCR, and layout analysis.

The run-llama/liteparse repository ships a high-performance PDF parser with first-class Python support generated via PyO3. Using LiteParse from Python with PyO3 bindings gives you access to the same Rust core used by the CLI—including PDFium extraction, optional OCR merging, and layout reconstruction—while working with idiomatic Python objects defined in types.py.

Architecture of the Python Bindings

The Native Module Wrapper

In packages/python/liteparse/parser.py, the wrapper imports the compiled extension as _NativeLiteParse from liteparse._liteparse. This generated module contains the Rust LiteParse struct exposed as a Python class. When you instantiate liteparse.LiteParse, the wrapper creates an instance of _NativeLiteParse and stores it as self._native.

Configuration Forwarding

The Python constructor collects all non-None keyword arguments into a kwargs dictionary and passes them directly to the Rust constructor. This mirrors the LiteParseConfig struct used internally in crates/liteparse/src/lib.rs. Configuration options like ocr_enabled, ocr_server_url, dpi, and output_format are forwarded without modification, ensuring the Python API matches the Rust CLI experience.

Result Conversion Pipeline

After the native parsing completes, the Rust side returns a PyParseResult object. The wrapper method _convert_native_result (defined in packages/python/liteparse/parser.py) walks the native pages and images, constructing Python dataclasses: ParsedPage, ExtractedImage, and finally the top-level ParseResult. These pure-Python objects are defined in packages/python/liteparse/types.py and provide type-safe access to layout data.

Installation and Basic Parsing

Install the precompiled wheel from PyPI, which includes the PyO3 extension module:

pip install liteparse

The following example demonstrates basic document parsing with OCR enabled:

from pathlib import Path
from liteparse import LiteParse, ParseError

# Initialize with configuration forwarded to Rust

parser = LiteParse(
    ocr_enabled=True,          # Enable OCR for scanned PDFs

    output_format="text",      # Choose plain-text output

    dpi=300,                   # Resolution for extraction

)

try:
    result = parser.parse("sample.pdf")
    print("Full document text:")
    print(result.text)                     # Single string with concatenated page text

    print(f"Pages parsed: {result.num_pages}")
except ParseError as exc:
    print(f"Parsing failed: {exc}")

Working with Structured Results

The ParseResult object contains rich page-level data accessible through Pythonic accessors. Each page is represented as a ParsedPage dataclass containing TextItem objects with spatial coordinates:


# Access page-level data

first_page = result.get_page(1)
if first_page:
    print("\n--- Page 1 ---")
    print(first_page.text)                 # Raw text of page 1

    print("Number of text items:", len(first_page.text_items))
    
    # Iterate over individual text items with coordinates

    for item in first_page.text_items[:5]:
        print(f"[{item.x:.1f},{item.y:.1f}] {item.text}")

For documents parsed with image_mode="embed", extracted images are available as ExtractedImage dataclasses:


# Save embedded images

for img in result.images:
    out_path = Path(f"image_{img.id}.{img.format}")
    out_path.write_bytes(img.bytes)
    print(f"Saved image {img.id} to {out_path}")

Screenshots and Search Utilities

The bindings expose additional Rust functionality through wrapper methods. The screenshot method calls the native Rust screenshot API (implemented in crates/liteparse/src/parser.rs) to render pages via PDFium:


# Take screenshots of specific pages

screenshots = parser.screenshot(
    "sample.pdf",
    page_numbers=[1, 2]      # Optional: limit to specific pages

)
for ss in screenshots:
    img_path = Path(f"screenshot_page_{ss.page_num}.png")
    img_path.write_bytes(ss.image_bytes)
    print(f"Saved screenshot for page {ss.page_num}")

For text search, the module exposes a pure-Python search_items function that forwards requests to the native search_items implementation and returns matching TextItem objects:

from liteparse import search_items

# Search for a phrase across page text items

matches = search_items(result.pages[0].text_items, "Lorem ipsum")
print(f"Found {len(matches)} matches on page 1")
for m in matches:
    print(f"Match at ({m.x:.1f},{m.y:.1f}) → {m.text}")

Configuration and Rust Integration

All configuration options available in the Rust LiteParseConfig struct are exposed as optional arguments to the Python class. The heavy lifting—PDFium document extraction, OCR processing via the OcrEngine trait (defined in crates/liteparse/src/ocr/mod.rs), and async I/O via tokio—remains in Rust. You can inspect the resolved configuration using get_config():

config = parser.get_config()
print(config)  # Returns the effective configuration from the native object

Summary

  • LiteParse uses PyO3 to expose a native Rust class as _liteparse.LiteParse, instantiated through the Python wrapper in packages/python/liteparse/parser.py.
  • The wrapper converts between Rust structs (PyParseResult) and Python dataclasses (ParseResult, ParsedPage, TextItem, ExtractedImage) defined in packages/python/liteparse/types.py.
  • All heavy computation—PDFium integration, OCR merging, and async I/O via tokio—remains in the Rust core while the API feels native to Python.
  • The parse() method supports both file paths (delegating to _native.parse) and raw bytes (delegating to _native.parse_bytes).
  • Advanced features like page screenshots and phrase search are available through the same lightweight binding layer.

Frequently Asked Questions

What is the relationship between the liteparse Python package and the Rust core?

The Python package is a thin wrapper around the Rust library. When you install liteparse from PyPI, you receive a precompiled extension module (_liteparse) built with PyO3 that exposes the Rust LiteParse struct directly to Python. According to the source in packages/python/liteparse/parser.py, the wrapper handles configuration marshalling and result conversion while the Rust side in crates/liteparse/src/lib.rs performs all parsing logic.

How does LiteParse handle OCR when called from Python?

The Python wrapper forwards ocr_enabled and ocr_server_url parameters to the Rust constructor. As implemented in crates/liteparse/src/ocr/mod.rs, the Rust core defines the OcrEngine trait and processes scanned PDFs within native code before returning structured results to Python. The OCR processing happens entirely in Rust; Python only receives the final extracted text items.

Can I parse PDFs from memory (bytes) instead of file paths?

Yes. The LiteParse.parse method detects input types and delegates to either _native.parse for file paths or _native.parse_bytes for raw bytes. Both methods return the same PyParseResult type, which the wrapper converts to Python dataclasses. This allows you to process PDFs received from network streams or databases without writing temporary files.

Are the Python bindings asynchronous?

While the underlying Rust code uses tokio for asynchronous operations, the Python bindings currently expose a synchronous API. The heavy lifting happens in Rust's async runtime, but your Python code blocks until parsing completes. This design makes the API compatible with standard Python workflows while still leveraging Rust's async performance for I/O-bound operations like OCR server requests.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →