Structural Difference Between the Page Struct and ParsedPage in LiteParse

The Page struct stores raw PDFium extraction results as an internal intermediate type, while ParsedPage represents the finalized, spatially-projected output containing concatenated reading-order text intended for public API consumption.

Understanding the distinction between these two core structures is essential when working with the run-llama/liteparse PDF processing pipeline. While both definitions reside in crates/liteparse/src/types.rs, they serve fundamentally different roles in the document extraction lifecycle. This article examines the structural difference between the Page struct and ParsedPage to clarify their fields, visibility, and positions within the processing chain.

Purpose and API Visibility

LiteParse maintains a strict separation between raw extraction data and processed output.

Page functions as a low-level container that holds glyph data straight from PDFium. It is marked with #[doc(hidden)] in crates/liteparse/src/types.rs (lines 60-65), indicating it is not part of the stable public API. This struct exists solely to transport raw extraction results before layout reconstruction occurs.

ParsedPage serves as the high-level, user-facing representation defined at lines 68-75 of the same file. Unlike its counterpart, this struct is publicly exported and serialized for consumption across JavaScript, Python, and Rust interfaces. It contains the final output after LiteParse applies its spatial-grid projection algorithm and optional OCR merging.

Field-by-Field Comparison

Although both structs share dimensional metadata, their data payloads differ significantly:

  • Page contains:

    • page_number: usize – The original PDF page index
    • page_width: f32 and page_height: f32 – Dimensions in PDF points
    • text_items: Vec<TextItem> – Raw extracted text fragments with position, font, and rotation data
  • ParsedPage contains:

    • page_number: usize – Same original index
    • page_width: f32 and page_height: f32 – Unchanged dimensions
    • text: String – Concatenated text assembled in reading order by the projection algorithm
    • text_items: Vec<TextItem> – The same collection enriched with projection metadata (snapping, anchors) and potentially OCR-derived items

The critical structural difference is the addition of the text field in ParsedPage, which provides the full page content as a single string rather than requiring consumers to manually reconstruct reading order from individual items.

Source Code Definitions

According to the run-llama/liteparse source code, the definitions appear sequentially in the types module:

// crates/liteparse/src/types.rs (lines 60-65)
#[doc(hidden)]
#[derive(Debug, Serialize)]
pub struct Page {
    pub page_number: usize,
    pub page_width: f32,
    pub page_height: f32,
    pub text_items: Vec<TextItem>,
}
// crates/liteparse/src/types.rs (lines 68-75)
#[derive(Debug, Serialize)]
pub struct ParsedPage {
    pub page_number: usize,
    pub page_width: f32,
    pub page_height: f32,
    pub text: String,
    pub text_items: Vec<TextItem>,
}

The Processing Pipeline

The lifecycle follows a strict transformation path: extraction → projection → output.

  1. crates/liteparse/src/extract.rs queries PDFium and constructs a Page struct containing raw glyph data and positioning information.

  2. crates/liteparse/src/projection.rs receives the Page and executes the project_page() function, which runs the spatial-grid algorithm to determine reading order and concatenates text fragments into the final text string.

  3. If OCR is enabled, ocr_merge.rs further enriches the text_items vector before finalizing the ParsedPage.

Usage Examples

Accessing ParsedPage via Public API

When using the Node.js bindings exposed in packages/node/src/lib.ts, you interact exclusively with ParsedPage structures:

import { LiteParse } from "liteparse";

(async () => {
  const parser = new LiteParse();
  const result = await parser.parse("sample.pdf");
  // result.pages is ParsedPage[]
  const first = result.pages[0];
  console.log(first.text);               // Full page text
  console.log(first.text_items.length); // Number of items
})();

Working with Intermediate Page Data

In Rust-only contexts or unit tests, you may encounter the raw Page type before projection occurs:

#[cfg(test)]
mod tests {
    use liteparse::types::Page;
    
    #[test]
    fn inspect_raw_extraction() {
        // extract_page returns a Page from PDFium
        let raw_page: Page = extract_page("example.pdf", 0);
        assert_eq!(raw_page.text_items.len(), 42);
        // Note: raw_page.text does not exist here
    }
}

Converting Page to ParsedPage

Internal library code performs the transformation using the projection module:

use liteparse::{types::{Page, ParsedPage}, projection::project_page};

fn finalize_page(raw: Page) -> ParsedPage {
    project_page(raw) // Builds the text string and enriches metadata
}

Summary

Frequently Asked Questions

Why does ParsedPage include both text and text_items fields?

ParsedPage includes the text string for convenience when you need the full page content, while retaining text_items for granular access to positioning, font metadata, and individual word coordinates. This dual structure allows simple use cases to read page.text directly while supporting complex layout analysis through the itemized array.

Can I access the intermediate Page struct when using the Python or Node.js bindings?

No. The Page struct is #[doc(hidden)] and internal to the Rust core. Both the Python wrapper (packages/python/liteparse/parser.py) and Node.js bindings (packages/node/src/lib.ts) only expose ParsedPage results after the projection step has completed.

What happens to text_items during the projection step?

During projection in crates/liteparse/src/projection.rs, the original text_items from the Page struct are preserved but enriched with spatial metadata such as grid snapping and reading-order anchors. If OCR is enabled, additional items from ocr_merge.rs are also appended to this vector before being stored in the final ParsedPage.

Is the Page struct considered part of the stable public API?

No. As indicated by the #[doc(hidden)] attribute in crates/liteparse/src/types.rs, Page is an implementation detail subject to change. Only ParsedPage and its fields are covered by semantic versioning guarantees in the run-llama/liteparse repository.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →