# Understanding physical_index Tags in PageIndex: Core Mechanisms for PDF Page Anchoring

> Discover how physical_index tags in VectifyAI/PageIndex create stable anchors for LLM table of contents and page verification. Learn core mechanisms for PDF page anchoring.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: internals
- Published: 2026-02-16

---

**`physical_index` tags are synthetic XML-style markers that wrap PDF page content to provide stable, machine-readable anchors enabling LLM-driven table-of-contents generation and precise page verification.**

In the VectifyAI/PageIndex repository, `physical_index` tags serve as the critical bridge between raw PDF text extraction and structured document understanding. These markers allow the pipeline to maintain explicit page boundaries throughout LLM processing, ensuring that table-of-contents entries map accurately to physical document locations.

## What Are physical_index Tags?

`physical_index` tags are programmatically injected delimiters that take the format `<physical_index_X>` where **X** represents the 1-based page number. Unlike natural text markers, these tags are synthetic references generated during the ingestion phase to create explicit page boundaries within a continuous text stream.

The tags function as dual-purpose anchors: they provide **visual cues** for LLMs to locate sections within tagged text, and **numeric references** that Python utilities convert to integers for arithmetic operations and list indexing.

## How physical_index Tags Drive the PageIndex Pipeline

### PDF Ingestion and Tag Injection

During the initial document processing, the `get_text_of_pdf_pages_with_labels` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 447-451) wraps each page's content with opening and closing `physical_index` tags. This transformation converts a list of page strings into a single searchable document while preserving explicit page boundaries.

```python

# utils.get_text_of_pdf_pages_with_labels

def get_text_of_pdf_pages_with_labels(pdf_pages, start_page, end_page):
    text = ""
    for page_num in range(start_page-1, end_page):
        text += f"<physical_index_{page_num+1}>\n{pdf_pages[page_num][0]}\n<physical_index_{page_num+1}>\n"
    return text

```

### TOC Creation and Tag References

When generating the table of contents, the `toc_index_extractor` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) (lines 245-254) explicitly references these tags in its LLM prompt. The model uses these markers to identify where sections begin within the tagged text stream, enabling accurate mapping of logical document structure to physical page locations.

```python

# page_index.toc_index_extractor (prompt snippet)

tob_extractor_prompt = """
You are given a table of contents in a json format and several pages of a document,
your job is to add the physical_index to the table of contents in the json format.

The provided pages contains tags like <physical_index_X> and <physical_index_X>
to indicate the physical location of the page X.
...
"""

```

### Converting Tags to Numeric Indices

For arithmetic operations and list indexing, the `convert_physical_index_to_int` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 545-555) strips the XML-style syntax and extracts the numeric page value. This conversion enables the pipeline to perform offset calculations and retrieve specific pages from the `page_list` array.

```python

# utils.convert_physical_index_to_int

def convert_physical_index_to_int(data):
    if isinstance(data, list):
        for i in range(len(data)):
            if isinstance(data[i], dict) and 'physical_index' in data[i]:
                if isinstance(data[i]['physical_index'], str):
                    if data[i]['physical_index'].startswith('<physical_index_'):
                        data[i]['physical_index'] = int(data[i]['physical_index']
                                                        .split('_')[-1].rstrip('>').strip())

```

### Page Verification and Lookup

During validation, functions like `check_title_appearance` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) (lines 13-20) use the numeric `physical_index` to retrieve raw page content from the `page_list` array. The tag itself is not stored in the page content; it serves solely as a reference telling the code which index to access.

```python

# page_index.check_title_appearance

async def check_title_appearance(item, page_list, start_index=1, model=None):
    if 'physical_index' not in item or item['physical_index'] is None:
        return {'answer': 'no', 'page_number': None}
    page_number = item['physical_index']
    page_text = page_list[page_number-start_index][0]   # ← fetch the exact page

    …

```

Similarly, `check_title_appearance_in_start_concurrent` (lines 88-90) validates that section titles appear at the beginning of their referenced pages by indexing directly into the `page_list` using the `physical_index` value.

## Key Implementation Files

| File | Role in physical_index Processing |
|------|-----------------------------------|
| **[`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py)** | Generates tagged page text via `get_text_of_pdf_pages_with_labels`, converts tag strings to integers via `convert_physical_index_to_int`, and performs final tree post-processing. |
| **[`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)** | Orchestrates the core workflow: prompts LLMs to recognize tags during TOC extraction, validates page mappings via `check_title_appearance`, and handles correction logic for missing indices. |
| **[`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py)** | Entry point demonstrating how the pipeline invokes these utilities on input PDF documents. |

## Summary

- **`physical_index` tags** are synthetic XML-style markers injected during PDF ingestion to create explicit page boundaries in plain text.
- **Tag injection** occurs in [`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py) via `get_text_of_pdf_pages_with_labels`, wrapping each page with `<physical_index_X>` delimiters.
- **LLM integration** leverages these tags as stable anchors during TOC generation, enabling the model to map logical sections to physical page numbers.
- **Numeric conversion** via `convert_physical_index_to_int` transforms tag strings into integers for array indexing and arithmetic operations.
- **Validation logic** uses the numeric indices to retrieve specific pages from `page_list` for title verification and offset correction.

## Frequently Asked Questions

### What is the exact format of physical_index tags?

`physical_index` tags follow the XML-style format `<physical_index_X>` where **X** represents the 1-based page number. Each page is wrapped with both opening and closing tags containing the same page number, creating a clear delimiter structure that persists through the LLM processing pipeline.

### How does PageIndex handle physical_index tags during TOC validation?

During validation, PageIndex converts tag strings to integers using `convert_physical_index_to_int` in [`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py), then retrieves the corresponding page content from the `page_list` array. Functions like `check_title_appearance` in [`page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index.py) use these numeric indices to verify that section titles actually appear on their referenced pages, enabling automated correction of TOC drift.

### Can physical_index tags be converted to integers for arithmetic operations?

Yes, the `convert_physical_index_to_int` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines 545-555) specifically handles this conversion. It strips the `<physical_index_` prefix and `>` suffix, extracts the numeric component, and casts it to an integer. This enables the pipeline to perform offset calculations, build page ranges, and index directly into Python lists.

### Where are physical_index tags generated in the codebase?

The tags are generated in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) within the `get_text_of_pdf_pages_with_labels` function (lines 447-451). This function iterates through PDF pages and wraps each page's text content with `<physical_index_X>` tags, creating a single searchable string that maintains explicit page boundaries for downstream LLM processing.