# Page Offset Calculation in PageIndex: Mapping TOC Page Numbers to Physical Indices

> Learn how PageIndex calculates page offset for TOC entries by finding the common difference between TOC page numbers and physical indices, ensuring accurate alignment.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: how-to-guide
- Published: 2026-02-16

---

**PageIndex calculates the page offset by finding the most common difference between printed TOC page numbers and their corresponding physical indices, then applies this integer to align all entries with the document's internal structure.**

The `VectifyAI/PageIndex` repository provides robust PDF indexing capabilities that reconcile printed page numbers with internal document structures. Understanding the **page offset calculation** is essential for developers working with Table of Contents (TOC) entries that contain pre-existing page numbers, as this process bridges the gap between human-readable pagination and zero-based physical indices.

## How Page Offset Calculation Works in [`page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index.py)

The core logic resides in `process_toc_with_page_numbers` within **[`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)**. This method executes a three-step statistical alignment process to determine the correct offset between printed page numbers and physical document indices.

### Step 1: Extract Matching Page Pairs with `extract_matching_page_pairs`

First, the system identifies corresponding entries between the original TOC JSON (containing `page` fields) and the intermediate TOC structure (containing LLM-derived `physical_index` values). The `extract_matching_page_pairs` function filters for titles appearing in both structures where the physical index falls after the last TOC page.

```python
matching_pairs = extract_matching_page_pairs(
    toc_with_page_number,
    toc_with_physical_index,
    start_page_index
)

```

This function operates on lines 371-383 of [`page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index.py), returning a list of dictionaries with `title`, `page`, and `physical_index` keys.

### Step 2: Calculate the Statistical Offset with `calculate_page_offset`

Next, `calculate_page_offset` computes the difference between physical indices and printed page numbers for each matched pair:

```

difference = physical_index - page

```

The function aggregates all differences, counts their frequencies, and selects the **most common value** as the page offset. This statistical approach tolerates occasional OCR errors or mismatches by choosing the mode rather than the mean.

```python
offset = calculate_page_offset(matching_pairs)

```

Located at lines 386-406 in [`page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index.py), this method ensures robust alignment even when individual entries contain noise.

### Step 3: Apply the Offset with `add_page_offset_to_toc_json`

Finally, `add_page_offset_to_toc_json` transforms the original TOC structure by replacing each `page` value with `physical_index = page + offset`, then removes the obsolete `page` field.

```python
toc_with_page_number = add_page_offset_to_toc_json(
    toc_with_page_number,
    offset
)

```

This function appears at lines 408-413 of [`page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index.py), producing a corrected TOC where all entries reference accurate physical indices suitable for downstream chunking and retrieval.

## Practical Implementation Example

The following example demonstrates the complete **page offset calculation** pipeline using the actual PageIndex API:

```python
from pageindex.page_index import (
    extract_matching_page_pairs,
    calculate_page_offset,
    add_page_offset_to_toc_json,
)

# Sample data representing matched TOC entries

pairs = [
    {"title": "Executive Summary", "page": 3, "physical_index": 12},
    {"title": "Methodology", "page": 5, "physical_index": 14},
    {"title": "Results", "page": 10, "physical_index": 19},
    # Noisy entry (off by 1, will be filtered out by mode)

    {"title": "Appendix", "page": 45, "physical_index": 55},
]

# Calculate the offset (most common difference = 9)

offset = calculate_page_offset(pairs)
print(f"Computed offset: {offset}")  # Output: 9

# Apply to original TOC structure

original_toc = [
    {"title": "Executive Summary", "page": 3},
    {"title": "Methodology", "page": 5},
]

corrected_toc = add_page_offset_to_toc_json(original_toc, offset)

# corrected_toc now contains physical_index values (12, 14)

```

This mirrors the internal calls made within `process_toc_with_page_numbers` (lines 631-639 of [`page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index.py)).

## Summary

- **Page offset calculation** reconciles printed TOC page numbers with zero-based physical indices by computing the most common difference between matched pairs.
- The `extract_matching_page_pairs` function filters entries to include only those appearing in both the original TOC and the LLM-derived physical index structure.
- `calculate_page_offset` uses statistical mode selection to tolerate OCR errors and noise, returning the integer offset that aligns the majority of entries.
- `add_page_offset_to_toc_json` applies this offset to transform `page` fields into accurate `physical_index` values for downstream processing.

## Frequently Asked Questions

### What is the purpose of page offset calculation in PageIndex?

The **page offset calculation** bridges the gap between human-readable page numbers printed in a PDF's Table of Contents and the internal zero-based indexing system used by the library. Since TOC page numbers typically start counting from the first content page while the physical index starts after the TOC pages themselves, the offset provides the constant integer needed to align these two reference systems.

### How does PageIndex handle OCR errors or mismatched page numbers?

PageIndex tolerates noise through statistical aggregation in the `calculate_page_offset` function. Rather than averaging differences or using the first match, the algorithm collects all `physical_index - page` differences from matched pairs and selects the **mode** (most frequent value). This approach automatically filters out occasional OCR errors or annotation mismatches that would otherwise skew the offset calculation.

### Which source files contain the page offset calculation logic?

The primary implementation resides in **[`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)**, specifically within the `process_toc_with_page_numbers` method (lines 631-639) and its helper functions: `extract_matching_page_pairs` (lines 371-383), `calculate_page_offset` (lines 386-406), and `add_page_offset_to_toc_json` (lines 408-413). Supporting utilities for index conversion are located in **[`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py)**.

### Can I manually specify a page offset instead of using automatic calculation?

While the `VectifyAI/PageIndex` repository primarily implements automatic offset detection through statistical analysis, the modular design of `add_page_offset_to_toc_json` allows you to manually apply any integer offset to your TOC JSON structure. Simply pass your desired offset value as the second argument to transform printed page numbers into physical indices without invoking the automatic `calculate_page_offset` function.