Page Offset Calculation in PageIndex: Mapping TOC Page Numbers to Physical Indices
PageIndex calculates the page offset by finding the most common difference between printed TOC page numbers and their corresponding physical indices, then applies this integer to align all entries with the document's internal structure.
The VectifyAI/PageIndex repository provides robust PDF indexing capabilities that reconcile printed page numbers with internal document structures. Understanding the page offset calculation is essential for developers working with Table of Contents (TOC) entries that contain pre-existing page numbers, as this process bridges the gap between human-readable pagination and zero-based physical indices.
How Page Offset Calculation Works in page_index.py
The core logic resides in process_toc_with_page_numbers within pageindex/page_index.py. This method executes a three-step statistical alignment process to determine the correct offset between printed page numbers and physical document indices.
Step 1: Extract Matching Page Pairs with extract_matching_page_pairs
First, the system identifies corresponding entries between the original TOC JSON (containing page fields) and the intermediate TOC structure (containing LLM-derived physical_index values). The extract_matching_page_pairs function filters for titles appearing in both structures where the physical index falls after the last TOC page.
matching_pairs = extract_matching_page_pairs(
toc_with_page_number,
toc_with_physical_index,
start_page_index
)
This function operates on lines 371-383 of page_index.py, returning a list of dictionaries with title, page, and physical_index keys.
Step 2: Calculate the Statistical Offset with calculate_page_offset
Next, calculate_page_offset computes the difference between physical indices and printed page numbers for each matched pair:
difference = physical_index - page
The function aggregates all differences, counts their frequencies, and selects the most common value as the page offset. This statistical approach tolerates occasional OCR errors or mismatches by choosing the mode rather than the mean.
offset = calculate_page_offset(matching_pairs)
Located at lines 386-406 in page_index.py, this method ensures robust alignment even when individual entries contain noise.
Step 3: Apply the Offset with add_page_offset_to_toc_json
Finally, add_page_offset_to_toc_json transforms the original TOC structure by replacing each page value with physical_index = page + offset, then removes the obsolete page field.
toc_with_page_number = add_page_offset_to_toc_json(
toc_with_page_number,
offset
)
This function appears at lines 408-413 of page_index.py, producing a corrected TOC where all entries reference accurate physical indices suitable for downstream chunking and retrieval.
Practical Implementation Example
The following example demonstrates the complete page offset calculation pipeline using the actual PageIndex API:
from pageindex.page_index import (
extract_matching_page_pairs,
calculate_page_offset,
add_page_offset_to_toc_json,
)
# Sample data representing matched TOC entries
pairs = [
{"title": "Executive Summary", "page": 3, "physical_index": 12},
{"title": "Methodology", "page": 5, "physical_index": 14},
{"title": "Results", "page": 10, "physical_index": 19},
# Noisy entry (off by 1, will be filtered out by mode)
{"title": "Appendix", "page": 45, "physical_index": 55},
]
# Calculate the offset (most common difference = 9)
offset = calculate_page_offset(pairs)
print(f"Computed offset: {offset}") # Output: 9
# Apply to original TOC structure
original_toc = [
{"title": "Executive Summary", "page": 3},
{"title": "Methodology", "page": 5},
]
corrected_toc = add_page_offset_to_toc_json(original_toc, offset)
# corrected_toc now contains physical_index values (12, 14)
This mirrors the internal calls made within process_toc_with_page_numbers (lines 631-639 of page_index.py).
Summary
- Page offset calculation reconciles printed TOC page numbers with zero-based physical indices by computing the most common difference between matched pairs.
- The
extract_matching_page_pairsfunction filters entries to include only those appearing in both the original TOC and the LLM-derived physical index structure. calculate_page_offsetuses statistical mode selection to tolerate OCR errors and noise, returning the integer offset that aligns the majority of entries.add_page_offset_to_toc_jsonapplies this offset to transformpagefields into accuratephysical_indexvalues for downstream processing.
Frequently Asked Questions
What is the purpose of page offset calculation in PageIndex?
The page offset calculation bridges the gap between human-readable page numbers printed in a PDF's Table of Contents and the internal zero-based indexing system used by the library. Since TOC page numbers typically start counting from the first content page while the physical index starts after the TOC pages themselves, the offset provides the constant integer needed to align these two reference systems.
How does PageIndex handle OCR errors or mismatched page numbers?
PageIndex tolerates noise through statistical aggregation in the calculate_page_offset function. Rather than averaging differences or using the first match, the algorithm collects all physical_index - page differences from matched pairs and selects the mode (most frequent value). This approach automatically filters out occasional OCR errors or annotation mismatches that would otherwise skew the offset calculation.
Which source files contain the page offset calculation logic?
The primary implementation resides in pageindex/page_index.py, specifically within the process_toc_with_page_numbers method (lines 631-639) and its helper functions: extract_matching_page_pairs (lines 371-383), calculate_page_offset (lines 386-406), and add_page_offset_to_toc_json (lines 408-413). Supporting utilities for index conversion are located in pageindex/utils.py.
Can I manually specify a page offset instead of using automatic calculation?
While the VectifyAI/PageIndex repository primarily implements automatic offset detection through statistical analysis, the modular design of add_page_offset_to_toc_json allows you to manually apply any integer offset to your TOC JSON structure. Simply pass your desired offset value as the second argument to transform printed page numbers into physical indices without invoking the automatic calculate_page_offset function.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →