Understanding physical_index Tags in PageIndex: Core Mechanisms for PDF Page Anchoring
physical_index tags are synthetic XML-style markers that wrap PDF page content to provide stable, machine-readable anchors enabling LLM-driven table-of-contents generation and precise page verification.
In the VectifyAI/PageIndex repository, physical_index tags serve as the critical bridge between raw PDF text extraction and structured document understanding. These markers allow the pipeline to maintain explicit page boundaries throughout LLM processing, ensuring that table-of-contents entries map accurately to physical document locations.
What Are physical_index Tags?
physical_index tags are programmatically injected delimiters that take the format <physical_index_X> where X represents the 1-based page number. Unlike natural text markers, these tags are synthetic references generated during the ingestion phase to create explicit page boundaries within a continuous text stream.
The tags function as dual-purpose anchors: they provide visual cues for LLMs to locate sections within tagged text, and numeric references that Python utilities convert to integers for arithmetic operations and list indexing.
How physical_index Tags Drive the PageIndex Pipeline
PDF Ingestion and Tag Injection
During the initial document processing, the get_text_of_pdf_pages_with_labels function in pageindex/utils.py (lines 447-451) wraps each page's content with opening and closing physical_index tags. This transformation converts a list of page strings into a single searchable document while preserving explicit page boundaries.
# utils.get_text_of_pdf_pages_with_labels
def get_text_of_pdf_pages_with_labels(pdf_pages, start_page, end_page):
text = ""
for page_num in range(start_page-1, end_page):
text += f"<physical_index_{page_num+1}>\n{pdf_pages[page_num][0]}\n<physical_index_{page_num+1}>\n"
return text
TOC Creation and Tag References
When generating the table of contents, the toc_index_extractor function in pageindex/page_index.py (lines 245-254) explicitly references these tags in its LLM prompt. The model uses these markers to identify where sections begin within the tagged text stream, enabling accurate mapping of logical document structure to physical page locations.
# page_index.toc_index_extractor (prompt snippet)
tob_extractor_prompt = """
You are given a table of contents in a json format and several pages of a document,
your job is to add the physical_index to the table of contents in the json format.
The provided pages contains tags like <physical_index_X> and <physical_index_X>
to indicate the physical location of the page X.
...
"""
Converting Tags to Numeric Indices
For arithmetic operations and list indexing, the convert_physical_index_to_int function in pageindex/utils.py (lines 545-555) strips the XML-style syntax and extracts the numeric page value. This conversion enables the pipeline to perform offset calculations and retrieve specific pages from the page_list array.
# utils.convert_physical_index_to_int
def convert_physical_index_to_int(data):
if isinstance(data, list):
for i in range(len(data)):
if isinstance(data[i], dict) and 'physical_index' in data[i]:
if isinstance(data[i]['physical_index'], str):
if data[i]['physical_index'].startswith('<physical_index_'):
data[i]['physical_index'] = int(data[i]['physical_index']
.split('_')[-1].rstrip('>').strip())
Page Verification and Lookup
During validation, functions like check_title_appearance in pageindex/page_index.py (lines 13-20) use the numeric physical_index to retrieve raw page content from the page_list array. The tag itself is not stored in the page content; it serves solely as a reference telling the code which index to access.
# page_index.check_title_appearance
async def check_title_appearance(item, page_list, start_index=1, model=None):
if 'physical_index' not in item or item['physical_index'] is None:
return {'answer': 'no', 'page_number': None}
page_number = item['physical_index']
page_text = page_list[page_number-start_index][0] # ← fetch the exact page
…
Similarly, check_title_appearance_in_start_concurrent (lines 88-90) validates that section titles appear at the beginning of their referenced pages by indexing directly into the page_list using the physical_index value.
Key Implementation Files
| File | Role in physical_index Processing |
|---|---|
pageindex/utils.py |
Generates tagged page text via get_text_of_pdf_pages_with_labels, converts tag strings to integers via convert_physical_index_to_int, and performs final tree post-processing. |
pageindex/page_index.py |
Orchestrates the core workflow: prompts LLMs to recognize tags during TOC extraction, validates page mappings via check_title_appearance, and handles correction logic for missing indices. |
run_pageindex.py |
Entry point demonstrating how the pipeline invokes these utilities on input PDF documents. |
Summary
physical_indextags are synthetic XML-style markers injected during PDF ingestion to create explicit page boundaries in plain text.- Tag injection occurs in
utils.pyviaget_text_of_pdf_pages_with_labels, wrapping each page with<physical_index_X>delimiters. - LLM integration leverages these tags as stable anchors during TOC generation, enabling the model to map logical sections to physical page numbers.
- Numeric conversion via
convert_physical_index_to_inttransforms tag strings into integers for array indexing and arithmetic operations. - Validation logic uses the numeric indices to retrieve specific pages from
page_listfor title verification and offset correction.
Frequently Asked Questions
What is the exact format of physical_index tags?
physical_index tags follow the XML-style format <physical_index_X> where X represents the 1-based page number. Each page is wrapped with both opening and closing tags containing the same page number, creating a clear delimiter structure that persists through the LLM processing pipeline.
How does PageIndex handle physical_index tags during TOC validation?
During validation, PageIndex converts tag strings to integers using convert_physical_index_to_int in utils.py, then retrieves the corresponding page content from the page_list array. Functions like check_title_appearance in page_index.py use these numeric indices to verify that section titles actually appear on their referenced pages, enabling automated correction of TOC drift.
Can physical_index tags be converted to integers for arithmetic operations?
Yes, the convert_physical_index_to_int function in pageindex/utils.py (lines 545-555) specifically handles this conversion. It strips the <physical_index_ prefix and > suffix, extracts the numeric component, and casts it to an integer. This enables the pipeline to perform offset calculations, build page ranges, and index directly into Python lists.
Where are physical_index tags generated in the codebase?
The tags are generated in pageindex/utils.py within the get_text_of_pdf_pages_with_labels function (lines 447-451). This function iterates through PDF pages and wraps each page's text content with <physical_index_X> tags, creating a single searchable string that maintains explicit page boundaries for downstream LLM processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →