# The Underlying Mechanism for Text Extraction from PDFium in LiteParse

> Discover how LiteParse extracts text from PDFs using PDFium and Rust. Learn its character-by-character process, handling ligatures, invisible text, and layout reconstruction.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: internals
- Published: 2026-06-07

---

**LiteParse extracts text by initializing the PDFium C library through Rust bindings, then iterates character-by-character across PDF pages to build structured `TextItem` objects while handling ligatures, invisible text layers, and spatial layout reconstruction.**

The **run-llama/liteparse** repository implements a sophisticated text extraction pipeline that bridges Google's PDFium C library with high-level Rust abstractions. Understanding this mechanism requires examining how the library transforms low-level PDF character streams into clean, layout-aware text structures.

## The PDFium Foundation in LiteParse

At the core of LiteParse's extraction capability lies the **pdfium** Rust crate, which provides safe bindings to Google's PDFium library. Unlike higher-level PDF parsers that rely on intermediate representations, LiteParse calls PDFium directly to access character-level positioning, font metadata, and rendering modes.

The entry point for all extraction operations is `extract::load_document_from_input` in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs). This function initializes the PDFium library and handles both file-based and in-memory PDF sources:

```rust
// crates/liteparse/src/extract.rs
let lib = Library::init();
let document = match input {
    PdfInput::Path(p) => lib.load_document(p, password)?,
    PdfInput::Bytes(b) => lib.load_document_from_bytes(b, password)?,
};

```

This initialization creates a `pdfium::Library` instance capable of handling password-protected documents when credentials are provided.

## The Three-Stage Text Extraction Pipeline

LiteParse processes PDF text through a sequential pipeline that moves from raw binary data to structured content. Each stage addresses specific challenges in PDF text extraction.

### Stage 1: Document Loading and Initialization

The first stage establishes the PDFium context and validates document accessibility. The `extract_pages_from_document` function iterates over the document's page count, retrieving individual `pdfium::Page` objects. For each page, the code obtains a `TextPage` handle—the primary interface for accessing character data:

```rust
let page = document.page(page_index)?;
let text_page = page.text()?;               // PDFium TextPage
let items = extract_page_text_items(&page, &text_page, &view_box)?;

```

According to the source in `crates/liteparse/src/extract.rs#L50-L59`, this step maps the PDF's internal text objects to PDFium's `FPDF_TEXTPAGE` structure, which provides methods like `char_count()` and `char_at_unchecked()` for indexed character access.

### Stage 2: Character-Level Page Processing

The `extract_page_text_items` function performs the heavy lifting of character extraction. Operating inside `crates/liteparse/src/extract.rs#L79-L150`, this function loops through every character index in the `TextPage`, extracting:

- **Unicode values** via `ch.unicode()`
- **Rendering modes** via `ch.text_render_mode()` 
- **Bounding boxes** via `ch.loose_char_box()`
- **Font and color metadata** for styling preservation

The code handles **ligature decomposition** by mapping PDFium control-code Unicode points to their expanded string equivalents:

```rust
// Core of the per-character loop (excerpt)
for i in 0..char_count {
    let ch = text_page.char_at_unchecked(i);
    let unicode = ch.unicode();

    // Skip invisible text when visible text dominates
    if skip_invisible && ch.text_render_mode() == Some(3) { continue; }

    // Map control-code ligatures
    let (c, ligature_tail) = match unicode {
        0x02 => ('-', ""),
        0x1A => ('f', "f"),
        0x1C => ('f', "i"),  // fi ligature
        // …
        _ => (char::from_u32(unicode).unwrap_or('?'), ""),
    };
    // ...
}

```

### Stage 3: Segmentation and TextItem Construction

Raw characters undergo aggregation through a custom **SegmentBuilder** that groups glyphs into logical units. This stage implements spatial analysis to detect:

- **Line breaks** using vertical tolerance thresholds
- **Column breaks** by detecting leftward horizontal jumps
- **Dot-leader tables** through gap analysis (`MAX_INLINE_GAP`)

The builder merges consecutive characters into unified `TextItem` structures containing consolidated bounding boxes, font metadata, rotation angles, and color information.

## Handling Edge Cases and Layout Complexity

LiteParse implements specialized logic for problematic PDF text patterns that simple extraction misses.

### Invisible Text Layer Detection

PDFs often contain invisible text (render mode 3) for PDF/A compliance or OCR backing layers. The extraction logic in `extract_page_text_items` skips these characters unless the page consists entirely of invisible text, preventing OCR artifacts from polluting the output while preserving legitimate hidden metadata layers.

### Duplicate Suppression

After character aggregation, `dedup_overlapping_items` removes duplicate text items that share identical content and overlapping bounding boxes. This eliminates redundant text streams that some PDF generators create for compatibility purposes.

### Ligature Expansion

Beyond simple character mapping, the pipeline handles complex typographic ligatures (e.g., "ff", "fi", "fl") by splitting single Unicode control codes into multiple constituent characters, ensuring searchability and text fidelity.

## From Raw Characters to Structured Layout

Following extraction and deduplication, raw `TextItem` objects flow into the **grid-projection** stage via `projection::project_pages_to_grid` in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs). This module reconstructs the spatial layout by:

- Projecting text onto a 2D grid to identify column boundaries
- Handling rotated text through coordinate transformation
- Establishing reading order independent of PDF content stream order

The final output produces `ParsedPage` structures containing layout-preserved text ready for downstream consumption.

## Usage Example

To extract text using the complete LiteParse pipeline:

```rust
use liteparse::parser::{LiteParse, LiteParseConfig};

#[tokio::main]
async fn main() -> Result<(), liteparse::error::LiteParseError> {
    // Minimal configuration (no OCR, default DPI)
    let cfg = LiteParseConfig::default();
    let parser = LiteParse::new(cfg);

    // Parse a file on disk
    let result = parser.parse("example.pdf").await?;

    // Full concatenated document text
    println!("Document text:\n{}", result.text);

    // Per-page info (bounding boxes, fonts, etc.)
    for page in result.pages {
        println!("Page {} – {} items", page.page_number, page.items.len());
    }
    Ok(())
}

```

The `parser::LiteParse::parse_input` method orchestrates this workflow, calling `extract::extract_pages_from_input` followed by `projection::project_pages_to_grid` as found in `crates/liteparse/src/parser.rs#L70-L90`.

## Summary

- **PDFium Integration**: LiteParse uses the `pdfium` Rust crate to bind Google's PDFium C library, loading documents via `extract::load_document_from_input` with support for both file paths and byte arrays.
- **Character-Level Processing**: Text extraction occurs at the glyph level using `TextPage.char_at_unchecked()` to access Unicode values, rendering modes, and bounding boxes.
- **Intelligent Filtering**: The pipeline skips invisible render-mode-3 characters (unless the page is entirely invisible) and expands ligature control codes into standard Unicode characters.
- **Spatial Reconstruction**: A custom `SegmentBuilder` groups characters into `TextItem` objects using spatial gap analysis (`MAX_INLINE_GAP`), followed by `dedup_overlapping_items` to remove duplicate layers.
- **Layout Awareness**: Final processing through `projection::project_pages_to_grid` reconstructs columns, rows, and rotated text into structured `ParsedPage` objects.

## Frequently Asked Questions

### How does LiteParse handle invisible text layers in PDFs?

LiteParse detects invisible text through PDFium's `text_render_mode()` method, specifically checking for render mode 3 (invisible). During extraction in `extract_page_text_items`, the code skips these characters unless the entire page consists of invisible text, which indicates an OCR backing layer. This prevents hidden metadata or PDF/A compliance text from appearing in the final output while preserving legitimate OCR content.

### What is the difference between TextPage and TextItem in LiteParse?

**TextPage** is a PDFium-internal structure (`FPDF_TEXTPAGE`) retrieved via `page.text()` that provides raw character access methods like `char_count()` and `char_at_unchecked()`. **TextItem** is LiteParse's high-level abstraction created by `extract_page_text_items`, representing aggregated groups of characters with unified bounding boxes, font metadata, and rotation information. While TextPage exposes individual glyphs, TextItem represents logical text segments like words or lines.

### How does LiteParse detect line breaks and column boundaries?

The `SegmentBuilder` inside `extract_page_text_items` analyzes spatial relationships between consecutive characters using constants like `MAX_INLINE_GAP` and vertical tolerance thresholds. When the vertical distance between characters exceeds the tolerance, or when text jumps leftward significantly, the builder triggers a new line or column break. This geometric analysis allows LiteParse to reconstruct reading order independent of the PDF's internal content stream ordering.

### Can LiteParse extract text from password-protected PDFs?

Yes. The `load_document_from_input` function in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs) accepts an optional password parameter that passes through to PDFium's `load_document` and `load_document_from_bytes` methods. When provided, PDFium handles the decryption internally before LiteParse begins character extraction, supporting both user and owner passwords for encrypted documents.