# Structural Difference Between the Page Struct and ParsedPage in LiteParse

> Discover the structural difference between Page and ParsedPage in LiteParse. Understand how raw PDFium data transforms into user-ready text for your applications.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-06

---

**The `Page` struct stores raw PDFium extraction results as an internal intermediate type, while `ParsedPage` represents the finalized, spatially-projected output containing concatenated reading-order text intended for public API consumption.**

Understanding the distinction between these two core structures is essential when working with the `run-llama/liteparse` PDF processing pipeline. While both definitions reside in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs), they serve fundamentally different roles in the document extraction lifecycle. This article examines the structural difference between the `Page` struct and `ParsedPage` to clarify their fields, visibility, and positions within the processing chain.

## Purpose and API Visibility

LiteParse maintains a strict separation between raw extraction data and processed output.

**`Page`** functions as a low-level container that holds glyph data straight from PDFium. It is marked with `#[doc(hidden)]` in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) (lines 60-65), indicating it is not part of the stable public API. This struct exists solely to transport raw extraction results before layout reconstruction occurs.

**`ParsedPage`** serves as the high-level, user-facing representation defined at lines 68-75 of the same file. Unlike its counterpart, this struct is publicly exported and serialized for consumption across JavaScript, Python, and Rust interfaces. It contains the final output after LiteParse applies its spatial-grid projection algorithm and optional OCR merging.

## Field-by-Field Comparison

Although both structs share dimensional metadata, their data payloads differ significantly:

- **`Page`** contains:
  - `page_number: usize` – The original PDF page index
  - `page_width: f32` and `page_height: f32` – Dimensions in PDF points
  - `text_items: Vec<TextItem>` – Raw extracted text fragments with position, font, and rotation data

- **`ParsedPage`** contains:
  - `page_number: usize` – Same original index
  - `page_width: f32` and `page_height: f32` – Unchanged dimensions
  - `text: String` – **Concatenated text** assembled in reading order by the projection algorithm
  - `text_items: Vec<TextItem>` – The same collection enriched with projection metadata (snapping, anchors) and potentially OCR-derived items

The critical structural difference is the addition of the `text` field in `ParsedPage`, which provides the full page content as a single string rather than requiring consumers to manually reconstruct reading order from individual items.

## Source Code Definitions

According to the `run-llama/liteparse` source code, the definitions appear sequentially in the types module:

```rust
// crates/liteparse/src/types.rs (lines 60-65)
#[doc(hidden)]
#[derive(Debug, Serialize)]
pub struct Page {
    pub page_number: usize,
    pub page_width: f32,
    pub page_height: f32,
    pub text_items: Vec<TextItem>,
}

```

```rust
// crates/liteparse/src/types.rs (lines 68-75)
#[derive(Debug, Serialize)]
pub struct ParsedPage {
    pub page_number: usize,
    pub page_width: f32,
    pub page_height: f32,
    pub text: String,
    pub text_items: Vec<TextItem>,
}

```

## The Processing Pipeline

The lifecycle follows a strict transformation path: **extraction → projection → output**.

1. **[`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs)** queries PDFium and constructs a `Page` struct containing raw glyph data and positioning information.

2. **[`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs)** receives the `Page` and executes the `project_page()` function, which runs the spatial-grid algorithm to determine reading order and concatenates text fragments into the final `text` string.

3. If OCR is enabled, [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) further enriches the `text_items` vector before finalizing the `ParsedPage`.

## Usage Examples

### Accessing ParsedPage via Public API

When using the Node.js bindings exposed in [`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts), you interact exclusively with `ParsedPage` structures:

```typescript
import { LiteParse } from "liteparse";

(async () => {
  const parser = new LiteParse();
  const result = await parser.parse("sample.pdf");
  // result.pages is ParsedPage[]
  const first = result.pages[0];
  console.log(first.text);               // Full page text
  console.log(first.text_items.length); // Number of items
})();

```

### Working with Intermediate Page Data

In Rust-only contexts or unit tests, you may encounter the raw `Page` type before projection occurs:

```rust
#[cfg(test)]
mod tests {
    use liteparse::types::Page;
    
    #[test]
    fn inspect_raw_extraction() {
        // extract_page returns a Page from PDFium
        let raw_page: Page = extract_page("example.pdf", 0);
        assert_eq!(raw_page.text_items.len(), 42);
        // Note: raw_page.text does not exist here
    }
}

```

### Converting Page to ParsedPage

Internal library code performs the transformation using the projection module:

```rust
use liteparse::{types::{Page, ParsedPage}, projection::project_page};

fn finalize_page(raw: Page) -> ParsedPage {
    project_page(raw) // Builds the text string and enriches metadata
}

```

## Summary

- **`Page`** is an internal `#[doc(hidden)]` struct in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) that stores raw PDFium extraction results without reading-order text concatenation.
- **`ParsedPage`** is the public output struct that adds a `text: String` field containing the full page content assembled by the spatial-grid algorithm.
- The conversion occurs in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) via the `project_page()` function, which transforms raw glyph data into structured, readable output.
- Language bindings in [`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts) and [`packages/python/liteparse/parser.py`](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py) exclusively expose `ParsedPage` to end users.

## Frequently Asked Questions

### Why does ParsedPage include both `text` and `text_items` fields?

**`ParsedPage`** includes the `text` string for convenience when you need the full page content, while retaining `text_items` for granular access to positioning, font metadata, and individual word coordinates. This dual structure allows simple use cases to read `page.text` directly while supporting complex layout analysis through the itemized array.

### Can I access the intermediate `Page` struct when using the Python or Node.js bindings?

No. The `Page` struct is `#[doc(hidden)]` and internal to the Rust core. Both the Python wrapper ([`packages/python/liteparse/parser.py`](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py)) and Node.js bindings ([`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts)) only expose `ParsedPage` results after the projection step has completed.

### What happens to `text_items` during the projection step?

During projection in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs), the original `text_items` from the `Page` struct are preserved but enriched with spatial metadata such as grid snapping and reading-order anchors. If OCR is enabled, additional items from [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs) are also appended to this vector before being stored in the final `ParsedPage`.

### Is the `Page` struct considered part of the stable public API?

No. As indicated by the `#[doc(hidden)]` attribute in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs), `Page` is an implementation detail subject to change. Only `ParsedPage` and its fields are covered by semantic versioning guarantees in the `run-llama/liteparse` repository.