Structural Difference Between the Page Struct and ParsedPage in LiteParse
The Page struct stores raw PDFium extraction results as an internal intermediate type, while ParsedPage represents the finalized, spatially-projected output containing concatenated reading-order text intended for public API consumption.
Understanding the distinction between these two core structures is essential when working with the run-llama/liteparse PDF processing pipeline. While both definitions reside in crates/liteparse/src/types.rs, they serve fundamentally different roles in the document extraction lifecycle. This article examines the structural difference between the Page struct and ParsedPage to clarify their fields, visibility, and positions within the processing chain.
Purpose and API Visibility
LiteParse maintains a strict separation between raw extraction data and processed output.
Page functions as a low-level container that holds glyph data straight from PDFium. It is marked with #[doc(hidden)] in crates/liteparse/src/types.rs (lines 60-65), indicating it is not part of the stable public API. This struct exists solely to transport raw extraction results before layout reconstruction occurs.
ParsedPage serves as the high-level, user-facing representation defined at lines 68-75 of the same file. Unlike its counterpart, this struct is publicly exported and serialized for consumption across JavaScript, Python, and Rust interfaces. It contains the final output after LiteParse applies its spatial-grid projection algorithm and optional OCR merging.
Field-by-Field Comparison
Although both structs share dimensional metadata, their data payloads differ significantly:
-
Pagecontains:page_number: usize– The original PDF page indexpage_width: f32andpage_height: f32– Dimensions in PDF pointstext_items: Vec<TextItem>– Raw extracted text fragments with position, font, and rotation data
-
ParsedPagecontains:page_number: usize– Same original indexpage_width: f32andpage_height: f32– Unchanged dimensionstext: String– Concatenated text assembled in reading order by the projection algorithmtext_items: Vec<TextItem>– The same collection enriched with projection metadata (snapping, anchors) and potentially OCR-derived items
The critical structural difference is the addition of the text field in ParsedPage, which provides the full page content as a single string rather than requiring consumers to manually reconstruct reading order from individual items.
Source Code Definitions
According to the run-llama/liteparse source code, the definitions appear sequentially in the types module:
// crates/liteparse/src/types.rs (lines 60-65)
#[doc(hidden)]
#[derive(Debug, Serialize)]
pub struct Page {
pub page_number: usize,
pub page_width: f32,
pub page_height: f32,
pub text_items: Vec<TextItem>,
}
// crates/liteparse/src/types.rs (lines 68-75)
#[derive(Debug, Serialize)]
pub struct ParsedPage {
pub page_number: usize,
pub page_width: f32,
pub page_height: f32,
pub text: String,
pub text_items: Vec<TextItem>,
}
The Processing Pipeline
The lifecycle follows a strict transformation path: extraction → projection → output.
-
crates/liteparse/src/extract.rsqueries PDFium and constructs aPagestruct containing raw glyph data and positioning information. -
crates/liteparse/src/projection.rsreceives thePageand executes theproject_page()function, which runs the spatial-grid algorithm to determine reading order and concatenates text fragments into the finaltextstring. -
If OCR is enabled,
ocr_merge.rsfurther enriches thetext_itemsvector before finalizing theParsedPage.
Usage Examples
Accessing ParsedPage via Public API
When using the Node.js bindings exposed in packages/node/src/lib.ts, you interact exclusively with ParsedPage structures:
import { LiteParse } from "liteparse";
(async () => {
const parser = new LiteParse();
const result = await parser.parse("sample.pdf");
// result.pages is ParsedPage[]
const first = result.pages[0];
console.log(first.text); // Full page text
console.log(first.text_items.length); // Number of items
})();
Working with Intermediate Page Data
In Rust-only contexts or unit tests, you may encounter the raw Page type before projection occurs:
#[cfg(test)]
mod tests {
use liteparse::types::Page;
#[test]
fn inspect_raw_extraction() {
// extract_page returns a Page from PDFium
let raw_page: Page = extract_page("example.pdf", 0);
assert_eq!(raw_page.text_items.len(), 42);
// Note: raw_page.text does not exist here
}
}
Converting Page to ParsedPage
Internal library code performs the transformation using the projection module:
use liteparse::{types::{Page, ParsedPage}, projection::project_page};
fn finalize_page(raw: Page) -> ParsedPage {
project_page(raw) // Builds the text string and enriches metadata
}
Summary
Pageis an internal#[doc(hidden)]struct incrates/liteparse/src/types.rsthat stores raw PDFium extraction results without reading-order text concatenation.ParsedPageis the public output struct that adds atext: Stringfield containing the full page content assembled by the spatial-grid algorithm.- The conversion occurs in
crates/liteparse/src/projection.rsvia theproject_page()function, which transforms raw glyph data into structured, readable output. - Language bindings in
packages/node/src/lib.tsandpackages/python/liteparse/parser.pyexclusively exposeParsedPageto end users.
Frequently Asked Questions
Why does ParsedPage include both text and text_items fields?
ParsedPage includes the text string for convenience when you need the full page content, while retaining text_items for granular access to positioning, font metadata, and individual word coordinates. This dual structure allows simple use cases to read page.text directly while supporting complex layout analysis through the itemized array.
Can I access the intermediate Page struct when using the Python or Node.js bindings?
No. The Page struct is #[doc(hidden)] and internal to the Rust core. Both the Python wrapper (packages/python/liteparse/parser.py) and Node.js bindings (packages/node/src/lib.ts) only expose ParsedPage results after the projection step has completed.
What happens to text_items during the projection step?
During projection in crates/liteparse/src/projection.rs, the original text_items from the Page struct are preserved but enriched with spatial metadata such as grid snapping and reading-order anchors. If OCR is enabled, additional items from ocr_merge.rs are also appended to this vector before being stored in the final ParsedPage.
Is the Page struct considered part of the stable public API?
No. As indicated by the #[doc(hidden)] attribute in crates/liteparse/src/types.rs, Page is an implementation detail subject to change. Only ParsedPage and its fields are covered by semantic versioning guarantees in the run-llama/liteparse repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →