# Common Failure Modes in PageIndex and How to Debug Them

> Learn about common PageIndex failure modes like LLM truncation and misaligned indices. Debug effectively by inspecting raw LLM outputs and logs in VectifyAI/PageIndex.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: debugging-guide
- Published: 2026-02-16

---

**TLDR:** PageIndex fails most often when LLM responses are truncated, physical page indices are misaligned, or JSON extraction fails, all of which can be diagnosed by inspecting raw LLM outputs and validation logs in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py).

PageIndex is an open-source hierarchical indexing tool developed by VectifyAI that transforms PDFs and Markdown files into searchable table-of-contents (TOC) structures using LLM reasoning. Because the pipeline orchestrates multiple asynchronous LLM calls, token-limited chunking, and custom page-number parsing, understanding **common failure modes in PageIndex** is essential for building reliable document retrieval systems.

## Top 10 Failure Modes in PageIndex

### 1. TOC Not Detected

**Where it happens:** `toc_detector_single_page` in **[`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py)** (lines [13‑25](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L13-L25)).

**Why it happens:** The prompt sent to ChatGPT does not return `"toc_detected": "yes"` or the LLM stops early (`finish_reason = "length"`).

**Quick debug check:** Print the raw LLM response (`response`) before `extract_json`. Verify the prompt string contains the full page text.

### 2. Page Indices Missing in TOC

**Where it happens:** `toc_extractor` → `toc_index_extractor` (lines [40‑66](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L40-L66)).

**Why it happens:** The LLM fails to insert `<physical_index_X>` tags, or the extractor mis‑parses them.

**Quick debug check:** After `toc_index_extractor`, dump the returned JSON and look for items without `physical_index`.

### 3. Physical Index Out‑of‑Range

**Where it happens:** `validate_and_truncate_physical_indices` (lines [14‑31](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L14-L31)).

**Why it happens:** Some sections claim page > `len(page_list)`. This can happen when the PDF parser drops pages or when the TOC spans a corrupted document.

**Quick debug check:** Log `page_list_length` and the first item with `physical_index > max_allowed_page`.

### 4. Incorrect Page‑Offset Calculation

**Where it happens:** `calculate_page_offset` (lines [38‑50](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L38-L50)).

**Why it happens:** If the “matching pairs” between the raw TOC and the physical‑index TOC are wrong, the most‑common offset can be off by several pages, breaking downstream `start_index / end_index`.

**Quick debug check:** Print `matching_pairs` and the resulting `offset`. Check that the `physical_index` and `page` fields truly refer to the same section title.

### 5. Token‑Limit Truncation

**Where it happens:** `post_processing` (lines [60‑71](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py#L60-L71)).

**Why it happens:** When `max_token_num_each_node` is too low, a node may be split incorrectly, leading to missing `appear_start` flags or empty children.

**Quick debug check:** After `post_processing`, verify that every node has `start_index` and `end_index` and that `node['text']` is not empty.

### 6. Async‑Concurrency Race or Exception

**Where it happens:** `check_title_appearance_in_start_concurrent` (lines [74‑102](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L74-L102)).

**Why it happens:** If any individual LLM call raises an exception, the whole concurrent gather can return an exception object instead of a result.

**Quick debug check:** Wrap the `await asyncio.gather` call with `return_exceptions=True` **and** log each exception; the code already does this, but you can add `logger.error(str(e))` inside the loop.

### 7. Markdown Header Parsing Errors

**Where it happens:** `extract_nodes_from_markdown` in **[`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py)** (lines [32‑57](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py#L32-L57)).

**Why it happens:** Incorrect header levels (e.g., missing `#`) cause the hierarchy to be flat or mis‑nested.

**Quick debug check:** Run the function on a test file and print `node_list`. Verify each node’s `level`.

### 8. JSON Extraction Failure

**Where it happens:** `extract_json` in **[`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py)** (lines [25‑45](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py#L25-L45)).

**Why it happens:** The LLM may prepend extra text or forget the ```json``` fence; the helper then returns `{}` silently.

**Quick debug check:** After each LLM call, `print(json_content)`; if it’s empty, inspect `response` for stray explanations.

### 9. Missing Environment Variable

**Where it happens:** `CHATGPT_API_KEY` loading in **[`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py)** (line [20](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py#L20)).

**Why it happens:** Without a valid key, all OpenAI calls return `"Error"` and the pipeline continues with empty strings.

**Quick debug check:** Run `print(os.getenv("CHATGPT_API_KEY"))` at startup; ensure `.env` is present.

### 10. Large‑Node Recursion Overflow

**Where it happens:** `process_large_node_recursively` (lines [93‑119](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L93-L119)).

**Why it happens:** When a node contains many pages and the token budget is exceeded, the function recurses on every child. If `max_page_num_each_node` is set too low, recursion depth can hit Python limits.

**Quick debug check:** Enable `logger.debug` before the recursive call; watch the number of nodes created. Increase `max_page_num_each_node` or raise `sys.setrecursionlimit`.

## Step-by-Step Debugging Workflow

1. **Enable a logger** – `logger = JsonLogger(doc)` is already created in `page_index_main`. Pass `logger` down to every helper (`check_title_appearance_in_start_concurrent`, `fix_incorrect_toc`, etc.) and inspect the generated JSON log files under `./logs`.

2. **Capture raw LLM output** – modify `ChatGPT_API` and `ChatGPT_API_async` to `return response.choices[0].message.content, response.choices[0].finish_reason` and log both. This isolates “finished vs truncated”.

3. **Validate page tokenization** – run `get_page_tokens(pdf_path)` manually and verify every `page_text` length. If `token_length` is `0`, the PDF parser missed the page (common with scanned PDFs). Switch to the PyMuPDF parser (`--pdf_parser pymupdf`).

4. **Assert TOC‑page alignment** – after `find_toc_pages`, dump `toc_page_list`. If empty, the TOC detector may be too strict; adjust the prompt or increase `opt.toc_check_page_num`.

5. **Unit‑test individual functions** – the repository ships with a `tests/` folder containing expected JSON structures. Compare your intermediate structures (`toc_with_page_number`, `matching_pairs`, etc.) against the fixtures.

6. **Fallback path** – if verification (`verify_toc`) yields < 0.6 accuracy, the code automatically falls back to “process_no_toc”. You can force the “with page numbers” path by calling `meta_processor` with `mode='process_toc_with_page_numbers'` to see where it fails.

## Practical Code Examples for Debugging

### Run PageIndex with a Debug Logger

```python
from pageindex.page_index import page_index_main
from pageindex.utils import JsonLogger

doc_path = "tests/pdfs/2023-annual-report.pdf"
logger = JsonLogger(doc_path)          # creates ./logs/<pdf>_TIMESTAMP.json

result = page_index_main(doc_path, opt=None)  # opt can be built with ConfigLoader if needed

print(result)   # final tree

```

*Inspect `./logs/2023-annual-report_*.json` for every intermediate step (TOC detection, offset calc, etc.).*

### Print Raw LLM Response for a Failing TOC Detection

```python
from pageindex.utils import ChatGPT_API_with_finish_reason

model = "gpt-4o-2024-11-20"
page_text = "<your page text here>"
response, reason = ChatGPT_API_with_finish_reason(model, f"Detect TOC: {page_text}")
print("Response:", response)
print("Finish reason:", reason)   # 'finished' vs 'max_output_reached'

```

### Verify Physical Indices Are Inside Document Bounds

```python
from pageindex.utils import count_tokens
from pageindex.page_index import validate_and_truncate_physical_indices

toc = [...]                     # JSON from toc_extractor

pages = get_page_tokens("my.pdf")
validated = validate_and_truncate_physical_indices(toc, len(pages))
for item in validated:
    if item.get("physical_index") is None:
        print("Out‑of‑range:", item["title"])

```

### Force the TOC‑with‑Page‑Numbers Path to Isolate Offset Bugs

```python
from pageindex.page_index import meta_processor

# Manually load the PDF → page_list (list of (text, token_len))

page_list = get_page_tokens("my.pdf")
opt = ConfigLoader().load({"model": "gpt-4o-2024-11-20",
                           "toc_check_page_num": 30})

# Assuming you already have `toc_content` & `toc_page_list` from a prior run:

tree = await meta_processor(page_list,
                            mode="process_toc_with_page_numbers",
                            toc_content=toc_content,
                            toc_page_list=toc_page_list,
                            start_index=1,
                            opt=opt,
                            logger=JsonLogger("my.pdf"))
print(tree)

```

### Unit‑Test the JSON Extractor on Malformed LLM Output

```python
from pageindex.utils import extract_json

bad_response = "Here is the answer: {\"thinking\": \"...\", \"toc_detected\": \"yes\"}"
print(extract_json(bad_response))   # should return {'thinking':..., 'toc_detected':'yes'}

```

## Key Source Files for Troubleshooting

| File | Role | Direct Link |
|------|------|-------------|
| [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) | Core pipeline: TOC detection, index extraction, validation, verification, fixing, and recursive node processing. | [page_index.py](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) |
| [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) | Helper utilities – token counting, OpenAI API wrappers, JSON extraction, PDF parsing, logging, and tree‑building helpers. | [utils.py](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) |
| [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py) | Markdown‑specific extractor (header parsing, thinning, summary generation). | [page_index_md.py](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py) |
| [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py) | CLI entry‑point that wires argument parsing to `page_index_main`. | [run_pageindex.py](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py) |
| `tests/` (e.g., [`tests/results/q1-fy25-earnings_structure.json`](https://github.com/VectifyAI/PageIndex/blob/main/tests/results/q1-fy25-earnings_structure.json)) | Reference outputs used for sanity‑checking the pipeline. | [tests folder](https://github.com/VectifyAI/PageIndex/tree/main/tests) |
| [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml) | Default configuration (model, page limits, toggles). | [config.yaml](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml) |

## Summary

- **TOC detection failures** usually stem from truncated LLM outputs or incorrect `finish_reason` values in `toc_detector_single_page`.
- **Physical index errors** (missing or out-of-range) occur in `validate_and_truncate_physical_indices` when the PDF parser drops pages or the LLM hallucinates page numbers.
- **Offset calculation bugs** in `calculate_page_offset` break downstream navigation when matching pairs between raw and indexed TOCs are misaligned.
- **Token truncation** in `post_processing` creates empty nodes or missing flags when `max_token_num_each_node` is set too aggressively.
- **Async exceptions** in `check_title_appearance_in_start_concurrent` can silently return exception objects instead of results unless `return_exceptions=True` is properly handled.
- **Environment and parsing errors** (missing `CHATGPT_API_KEY`, malformed JSON, or Markdown header issues) are caught by validating inputs at the entry points in [`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py) and [`page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/page_index_md.py).

## Frequently Asked Questions

### Why does PageIndex return an empty TOC even when my PDF has a clear table of contents?

This typically occurs in `toc_detector_single_page` when the LLM returns `"toc_detected": "no"` or the response is truncated (`finish_reason = "length"`). To debug, print the raw LLM response before `extract_json` processes it, and verify that the prompt contains the full page text. Increasing `opt.toc_check_page_num` may also help if the TOC appears later in the document.

### How do I fix "physical_index out of range" errors when processing large PDFs?

These errors originate in `validate_and_truncate_physical_indices` (lines [14‑31](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L14-L31)) when the LLM assigns page numbers exceeding the actual page count. This usually indicates that the PDF parser dropped pages (common with scanned PDFs). Switch to the PyMuPDF parser using `--pdf_parser pymupdf`, or manually verify with `get_page_tokens(pdf_path)` that every page has a non-zero token length.

### What causes incorrect page offsets that break section navigation?

Incorrect offsets stem from `calculate_page_offset` (lines [38‑50](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L38-L50)) when the "matching pairs" between the raw TOC and the physical-index TOC are misaligned. This happens if the LLM hallucinates section titles or if `toc_index_extractor` fails to tag pages correctly. Debug by printing `matching_pairs` and the resulting `offset` to ensure that `physical_index` and `page` fields refer to the same section title.

### How can I prevent async exceptions from crashing the entire pipeline?

The `check_title_appearance_in_start_concurrent` function (lines [74‑102](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py#L74-L102)) uses `asyncio.gather` to run multiple LLM calls concurrently. If one call raises an exception, the gather returns an exception object instead of a result. Ensure you are using `return_exceptions=True` and add explicit logging inside the exception handler with `logger.error(str(e))` to identify which specific call failed.

### Why does the JSON extractor return empty dictionaries?

The `extract_json` utility in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) (lines [25‑45](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py#L25-L45)) fails when the LLM prepends explanatory text or omits the ```json``` fence. The function then returns `{}` silently. To debug, print `json_content` immediately after the LLM call; if empty, inspect the raw `response` for stray explanations and consider tightening the prompt to forbid preamble text.