Common Failure Modes in PageIndex and How to Debug Them

TLDR: PageIndex fails most often when LLM responses are truncated, physical page indices are misaligned, or JSON extraction fails, all of which can be diagnosed by inspecting raw LLM outputs and validation logs in pageindex/page_index.py.

PageIndex is an open-source hierarchical indexing tool developed by VectifyAI that transforms PDFs and Markdown files into searchable table-of-contents (TOC) structures using LLM reasoning. Because the pipeline orchestrates multiple asynchronous LLM calls, token-limited chunking, and custom page-number parsing, understanding common failure modes in PageIndex is essential for building reliable document retrieval systems.

Top 10 Failure Modes in PageIndex

1. TOC Not Detected

Where it happens: toc_detector_single_page in pageindex/page_index.py (lines 13‑25).

Why it happens: The prompt sent to ChatGPT does not return "toc_detected": "yes" or the LLM stops early (finish_reason = "length").

Quick debug check: Print the raw LLM response (response) before extract_json. Verify the prompt string contains the full page text.

2. Page Indices Missing in TOC

Where it happens: toc_extractor → toc_index_extractor (lines 40‑66).

Why it happens: The LLM fails to insert <physical_index_X> tags, or the extractor mis‑parses them.

Quick debug check: After toc_index_extractor, dump the returned JSON and look for items without physical_index.

3. Physical Index Out‑of‑Range

Where it happens: validate_and_truncate_physical_indices (lines 14‑31).

Why it happens: Some sections claim page > len(page_list). This can happen when the PDF parser drops pages or when the TOC spans a corrupted document.

Quick debug check: Log page_list_length and the first item with physical_index > max_allowed_page.

4. Incorrect Page‑Offset Calculation

Where it happens: calculate_page_offset (lines 38‑50).

Why it happens: If the “matching pairs” between the raw TOC and the physical‑index TOC are wrong, the most‑common offset can be off by several pages, breaking downstream start_index / end_index.

Quick debug check: Print matching_pairs and the resulting offset. Check that the physical_index and page fields truly refer to the same section title.

5. Token‑Limit Truncation

Where it happens: post_processing (lines 60‑71).

Why it happens: When max_token_num_each_node is too low, a node may be split incorrectly, leading to missing appear_start flags or empty children.

Quick debug check: After post_processing, verify that every node has start_index and end_index and that node['text'] is not empty.

6. Async‑Concurrency Race or Exception

Where it happens: check_title_appearance_in_start_concurrent (lines 74‑102).

Why it happens: If any individual LLM call raises an exception, the whole concurrent gather can return an exception object instead of a result.

Quick debug check: Wrap the await asyncio.gather call with return_exceptions=True and log each exception; the code already does this, but you can add logger.error(str(e)) inside the loop.

7. Markdown Header Parsing Errors

Where it happens: extract_nodes_from_markdown in pageindex/page_index_md.py (lines 32‑57).

Why it happens: Incorrect header levels (e.g., missing #) cause the hierarchy to be flat or mis‑nested.

Quick debug check: Run the function on a test file and print node_list. Verify each node’s level.

8. JSON Extraction Failure

Where it happens: extract_json in pageindex/utils.py (lines 25‑45).

Why it happens: The LLM may prepend extra text or forget the json fence; the helper then returns {} silently.

Quick debug check: After each LLM call, print(json_content); if it’s empty, inspect response for stray explanations.

9. Missing Environment Variable

Where it happens: CHATGPT_API_KEY loading in utils.py (line 20).

Why it happens: Without a valid key, all OpenAI calls return "Error" and the pipeline continues with empty strings.

Quick debug check: Run print(os.getenv("CHATGPT_API_KEY")) at startup; ensure .env is present.

10. Large‑Node Recursion Overflow

Where it happens: process_large_node_recursively (lines 93‑119).

Why it happens: When a node contains many pages and the token budget is exceeded, the function recurses on every child. If max_page_num_each_node is set too low, recursion depth can hit Python limits.

Quick debug check: Enable logger.debug before the recursive call; watch the number of nodes created. Increase max_page_num_each_node or raise sys.setrecursionlimit.

Step-by-Step Debugging Workflow

  1. Enable a logger – logger = JsonLogger(doc) is already created in page_index_main. Pass logger down to every helper (check_title_appearance_in_start_concurrent, fix_incorrect_toc, etc.) and inspect the generated JSON log files under ./logs.

  2. Capture raw LLM output – modify ChatGPT_API and ChatGPT_API_async to return response.choices[0].message.content, response.choices[0].finish_reason and log both. This isolates “finished vs truncated”.

  3. Validate page tokenization – run get_page_tokens(pdf_path) manually and verify every page_text length. If token_length is 0, the PDF parser missed the page (common with scanned PDFs). Switch to the PyMuPDF parser (--pdf_parser pymupdf).

  4. Assert TOC‑page alignment – after find_toc_pages, dump toc_page_list. If empty, the TOC detector may be too strict; adjust the prompt or increase opt.toc_check_page_num.

  5. Unit‑test individual functions – the repository ships with a tests/ folder containing expected JSON structures. Compare your intermediate structures (toc_with_page_number, matching_pairs, etc.) against the fixtures.

  6. Fallback path – if verification (verify_toc) yields < 0.6 accuracy, the code automatically falls back to “process_no_toc”. You can force the “with page numbers” path by calling meta_processor with mode='process_toc_with_page_numbers' to see where it fails.

Practical Code Examples for Debugging

Run PageIndex with a Debug Logger

from pageindex.page_index import page_index_main
from pageindex.utils import JsonLogger

doc_path = "tests/pdfs/2023-annual-report.pdf"
logger = JsonLogger(doc_path)          # creates ./logs/<pdf>_TIMESTAMP.json

result = page_index_main(doc_path, opt=None)  # opt can be built with ConfigLoader if needed

print(result)   # final tree

Inspect ./logs/2023-annual-report_*.json for every intermediate step (TOC detection, offset calc, etc.).

from pageindex.utils import ChatGPT_API_with_finish_reason

model = "gpt-4o-2024-11-20"
page_text = "<your page text here>"
response, reason = ChatGPT_API_with_finish_reason(model, f"Detect TOC: {page_text}")
print("Response:", response)
print("Finish reason:", reason)   # 'finished' vs 'max_output_reached'

Verify Physical Indices Are Inside Document Bounds

from pageindex.utils import count_tokens
from pageindex.page_index import validate_and_truncate_physical_indices

toc = [...]                     # JSON from toc_extractor

pages = get_page_tokens("my.pdf")
validated = validate_and_truncate_physical_indices(toc, len(pages))
for item in validated:
    if item.get("physical_index") is None:
        print("Out‑of‑range:", item["title"])

Force the TOC‑with‑Page‑Numbers Path to Isolate Offset Bugs

from pageindex.page_index import meta_processor

# Manually load the PDF → page_list (list of (text, token_len))

page_list = get_page_tokens("my.pdf")
opt = ConfigLoader().load({"model": "gpt-4o-2024-11-20",
                           "toc_check_page_num": 30})

# Assuming you already have `toc_content` & `toc_page_list` from a prior run:

tree = await meta_processor(page_list,
                            mode="process_toc_with_page_numbers",
                            toc_content=toc_content,
                            toc_page_list=toc_page_list,
                            start_index=1,
                            opt=opt,
                            logger=JsonLogger("my.pdf"))
print(tree)

Unit‑Test the JSON Extractor on Malformed LLM Output

from pageindex.utils import extract_json

bad_response = "Here is the answer: {\"thinking\": \"...\", \"toc_detected\": \"yes\"}"
print(extract_json(bad_response))   # should return {'thinking':..., 'toc_detected':'yes'}

Key Source Files for Troubleshooting

File Role Direct Link
pageindex/page_index.py Core pipeline: TOC detection, index extraction, validation, verification, fixing, and recursive node processing. page_index.py
pageindex/utils.py Helper utilities – token counting, OpenAI API wrappers, JSON extraction, PDF parsing, logging, and tree‑building helpers. utils.py
pageindex/page_index_md.py Markdown‑specific extractor (header parsing, thinning, summary generation). page_index_md.py
run_pageindex.py CLI entry‑point that wires argument parsing to page_index_main. run_pageindex.py
tests/ (e.g., tests/results/q1-fy25-earnings_structure.json) Reference outputs used for sanity‑checking the pipeline. tests folder
pageindex/config.yaml Default configuration (model, page limits, toggles). config.yaml

Summary

  • TOC detection failures usually stem from truncated LLM outputs or incorrect finish_reason values in toc_detector_single_page.
  • Physical index errors (missing or out-of-range) occur in validate_and_truncate_physical_indices when the PDF parser drops pages or the LLM hallucinates page numbers.
  • Offset calculation bugs in calculate_page_offset break downstream navigation when matching pairs between raw and indexed TOCs are misaligned.
  • Token truncation in post_processing creates empty nodes or missing flags when max_token_num_each_node is set too aggressively.
  • Async exceptions in check_title_appearance_in_start_concurrent can silently return exception objects instead of results unless return_exceptions=True is properly handled.
  • Environment and parsing errors (missing CHATGPT_API_KEY, malformed JSON, or Markdown header issues) are caught by validating inputs at the entry points in utils.py and page_index_md.py.

Frequently Asked Questions

Why does PageIndex return an empty TOC even when my PDF has a clear table of contents?

This typically occurs in toc_detector_single_page when the LLM returns "toc_detected": "no" or the response is truncated (finish_reason = "length"). To debug, print the raw LLM response before extract_json processes it, and verify that the prompt contains the full page text. Increasing opt.toc_check_page_num may also help if the TOC appears later in the document.

How do I fix "physical_index out of range" errors when processing large PDFs?

These errors originate in validate_and_truncate_physical_indices (lines 14‑31) when the LLM assigns page numbers exceeding the actual page count. This usually indicates that the PDF parser dropped pages (common with scanned PDFs). Switch to the PyMuPDF parser using --pdf_parser pymupdf, or manually verify with get_page_tokens(pdf_path) that every page has a non-zero token length.

What causes incorrect page offsets that break section navigation?

Incorrect offsets stem from calculate_page_offset (lines 38‑50) when the "matching pairs" between the raw TOC and the physical-index TOC are misaligned. This happens if the LLM hallucinates section titles or if toc_index_extractor fails to tag pages correctly. Debug by printing matching_pairs and the resulting offset to ensure that physical_index and page fields refer to the same section title.

How can I prevent async exceptions from crashing the entire pipeline?

The check_title_appearance_in_start_concurrent function (lines 74‑102) uses asyncio.gather to run multiple LLM calls concurrently. If one call raises an exception, the gather returns an exception object instead of a result. Ensure you are using return_exceptions=True and add explicit logging inside the exception handler with logger.error(str(e)) to identify which specific call failed.

Why does the JSON extractor return empty dictionaries?

The extract_json utility in pageindex/utils.py (lines 25‑45) fails when the LLM prepends explanatory text or omits the json fence. The function then returns {} silently. To debug, print json_content immediately after the LLM call; if empty, inspect the raw response for stray explanations and consider tightening the prompt to forbid preamble text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →