How to Use LiteParse from Python with PyO3 Bindings: A Complete Guide
LiteParse exposes its Rust PDF parsing engine to Python through PyO3 bindings, providing a LiteParse class that instantiates a native Rust struct and converts results into Python dataclasses for text extraction, OCR, and layout analysis.
The run-llama/liteparse repository ships a high-performance PDF parser with first-class Python support generated via PyO3. Using LiteParse from Python with PyO3 bindings gives you access to the same Rust core used by the CLI—including PDFium extraction, optional OCR merging, and layout reconstruction—while working with idiomatic Python objects defined in types.py.
Architecture of the Python Bindings
The Native Module Wrapper
In packages/python/liteparse/parser.py, the wrapper imports the compiled extension as _NativeLiteParse from liteparse._liteparse. This generated module contains the Rust LiteParse struct exposed as a Python class. When you instantiate liteparse.LiteParse, the wrapper creates an instance of _NativeLiteParse and stores it as self._native.
Configuration Forwarding
The Python constructor collects all non-None keyword arguments into a kwargs dictionary and passes them directly to the Rust constructor. This mirrors the LiteParseConfig struct used internally in crates/liteparse/src/lib.rs. Configuration options like ocr_enabled, ocr_server_url, dpi, and output_format are forwarded without modification, ensuring the Python API matches the Rust CLI experience.
Result Conversion Pipeline
After the native parsing completes, the Rust side returns a PyParseResult object. The wrapper method _convert_native_result (defined in packages/python/liteparse/parser.py) walks the native pages and images, constructing Python dataclasses: ParsedPage, ExtractedImage, and finally the top-level ParseResult. These pure-Python objects are defined in packages/python/liteparse/types.py and provide type-safe access to layout data.
Installation and Basic Parsing
Install the precompiled wheel from PyPI, which includes the PyO3 extension module:
pip install liteparse
The following example demonstrates basic document parsing with OCR enabled:
from pathlib import Path
from liteparse import LiteParse, ParseError
# Initialize with configuration forwarded to Rust
parser = LiteParse(
ocr_enabled=True, # Enable OCR for scanned PDFs
output_format="text", # Choose plain-text output
dpi=300, # Resolution for extraction
)
try:
result = parser.parse("sample.pdf")
print("Full document text:")
print(result.text) # Single string with concatenated page text
print(f"Pages parsed: {result.num_pages}")
except ParseError as exc:
print(f"Parsing failed: {exc}")
Working with Structured Results
The ParseResult object contains rich page-level data accessible through Pythonic accessors. Each page is represented as a ParsedPage dataclass containing TextItem objects with spatial coordinates:
# Access page-level data
first_page = result.get_page(1)
if first_page:
print("\n--- Page 1 ---")
print(first_page.text) # Raw text of page 1
print("Number of text items:", len(first_page.text_items))
# Iterate over individual text items with coordinates
for item in first_page.text_items[:5]:
print(f"[{item.x:.1f},{item.y:.1f}] {item.text}")
For documents parsed with image_mode="embed", extracted images are available as ExtractedImage dataclasses:
# Save embedded images
for img in result.images:
out_path = Path(f"image_{img.id}.{img.format}")
out_path.write_bytes(img.bytes)
print(f"Saved image {img.id} to {out_path}")
Screenshots and Search Utilities
The bindings expose additional Rust functionality through wrapper methods. The screenshot method calls the native Rust screenshot API (implemented in crates/liteparse/src/parser.rs) to render pages via PDFium:
# Take screenshots of specific pages
screenshots = parser.screenshot(
"sample.pdf",
page_numbers=[1, 2] # Optional: limit to specific pages
)
for ss in screenshots:
img_path = Path(f"screenshot_page_{ss.page_num}.png")
img_path.write_bytes(ss.image_bytes)
print(f"Saved screenshot for page {ss.page_num}")
For text search, the module exposes a pure-Python search_items function that forwards requests to the native search_items implementation and returns matching TextItem objects:
from liteparse import search_items
# Search for a phrase across page text items
matches = search_items(result.pages[0].text_items, "Lorem ipsum")
print(f"Found {len(matches)} matches on page 1")
for m in matches:
print(f"Match at ({m.x:.1f},{m.y:.1f}) → {m.text}")
Configuration and Rust Integration
All configuration options available in the Rust LiteParseConfig struct are exposed as optional arguments to the Python class. The heavy lifting—PDFium document extraction, OCR processing via the OcrEngine trait (defined in crates/liteparse/src/ocr/mod.rs), and async I/O via tokio—remains in Rust. You can inspect the resolved configuration using get_config():
config = parser.get_config()
print(config) # Returns the effective configuration from the native object
Summary
- LiteParse uses PyO3 to expose a native Rust class as
_liteparse.LiteParse, instantiated through the Python wrapper inpackages/python/liteparse/parser.py. - The wrapper converts between Rust structs (
PyParseResult) and Python dataclasses (ParseResult,ParsedPage,TextItem,ExtractedImage) defined inpackages/python/liteparse/types.py. - All heavy computation—PDFium integration, OCR merging, and async I/O via
tokio—remains in the Rust core while the API feels native to Python. - The
parse()method supports both file paths (delegating to_native.parse) and raw bytes (delegating to_native.parse_bytes). - Advanced features like page screenshots and phrase search are available through the same lightweight binding layer.
Frequently Asked Questions
What is the relationship between the liteparse Python package and the Rust core?
The Python package is a thin wrapper around the Rust library. When you install liteparse from PyPI, you receive a precompiled extension module (_liteparse) built with PyO3 that exposes the Rust LiteParse struct directly to Python. According to the source in packages/python/liteparse/parser.py, the wrapper handles configuration marshalling and result conversion while the Rust side in crates/liteparse/src/lib.rs performs all parsing logic.
How does LiteParse handle OCR when called from Python?
The Python wrapper forwards ocr_enabled and ocr_server_url parameters to the Rust constructor. As implemented in crates/liteparse/src/ocr/mod.rs, the Rust core defines the OcrEngine trait and processes scanned PDFs within native code before returning structured results to Python. The OCR processing happens entirely in Rust; Python only receives the final extracted text items.
Can I parse PDFs from memory (bytes) instead of file paths?
Yes. The LiteParse.parse method detects input types and delegates to either _native.parse for file paths or _native.parse_bytes for raw bytes. Both methods return the same PyParseResult type, which the wrapper converts to Python dataclasses. This allows you to process PDFs received from network streams or databases without writing temporary files.
Are the Python bindings asynchronous?
While the underlying Rust code uses tokio for asynchronous operations, the Python bindings currently expose a synchronous API. The heavy lifting happens in Rust's async runtime, but your Python code blocks until parsing completes. This design makes the API compatible with standard Python workflows while still leveraging Rust's async performance for I/O-bound operations like OCR server requests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →