Common Failure Modes in PageIndex and How to Debug Them
TLDR: PageIndex fails most often when LLM responses are truncated, physical page indices are misaligned, or JSON extraction fails, all of which can be diagnosed by inspecting raw LLM outputs and validation logs in pageindex/page_index.py.
PageIndex is an open-source hierarchical indexing tool developed by VectifyAI that transforms PDFs and Markdown files into searchable table-of-contents (TOC) structures using LLM reasoning. Because the pipeline orchestrates multiple asynchronous LLM calls, token-limited chunking, and custom page-number parsing, understanding common failure modes in PageIndex is essential for building reliable document retrieval systems.
Top 10 Failure Modes in PageIndex
1. TOC Not Detected
Where it happens: toc_detector_single_page in pageindex/page_index.py (lines 13‑25).
Why it happens: The prompt sent to ChatGPT does not return "toc_detected": "yes" or the LLM stops early (finish_reason = "length").
Quick debug check: Print the raw LLM response (response) before extract_json. Verify the prompt string contains the full page text.
2. Page Indices Missing in TOC
Where it happens: toc_extractor → toc_index_extractor (lines 40‑66).
Why it happens: The LLM fails to insert <physical_index_X> tags, or the extractor mis‑parses them.
Quick debug check: After toc_index_extractor, dump the returned JSON and look for items without physical_index.
3. Physical Index Out‑of‑Range
Where it happens: validate_and_truncate_physical_indices (lines 14‑31).
Why it happens: Some sections claim page > len(page_list). This can happen when the PDF parser drops pages or when the TOC spans a corrupted document.
Quick debug check: Log page_list_length and the first item with physical_index > max_allowed_page.
4. Incorrect Page‑Offset Calculation
Where it happens: calculate_page_offset (lines 38‑50).
Why it happens: If the “matching pairs” between the raw TOC and the physical‑index TOC are wrong, the most‑common offset can be off by several pages, breaking downstream start_index / end_index.
Quick debug check: Print matching_pairs and the resulting offset. Check that the physical_index and page fields truly refer to the same section title.
5. Token‑Limit Truncation
Where it happens: post_processing (lines 60‑71).
Why it happens: When max_token_num_each_node is too low, a node may be split incorrectly, leading to missing appear_start flags or empty children.
Quick debug check: After post_processing, verify that every node has start_index and end_index and that node['text'] is not empty.
6. Async‑Concurrency Race or Exception
Where it happens: check_title_appearance_in_start_concurrent (lines 74‑102).
Why it happens: If any individual LLM call raises an exception, the whole concurrent gather can return an exception object instead of a result.
Quick debug check: Wrap the await asyncio.gather call with return_exceptions=True and log each exception; the code already does this, but you can add logger.error(str(e)) inside the loop.
7. Markdown Header Parsing Errors
Where it happens: extract_nodes_from_markdown in pageindex/page_index_md.py (lines 32‑57).
Why it happens: Incorrect header levels (e.g., missing #) cause the hierarchy to be flat or mis‑nested.
Quick debug check: Run the function on a test file and print node_list. Verify each node’s level.
8. JSON Extraction Failure
Where it happens: extract_json in pageindex/utils.py (lines 25‑45).
Why it happens: The LLM may prepend extra text or forget the json fence; the helper then returns {} silently.
Quick debug check: After each LLM call, print(json_content); if it’s empty, inspect response for stray explanations.
9. Missing Environment Variable
Where it happens: CHATGPT_API_KEY loading in utils.py (line 20).
Why it happens: Without a valid key, all OpenAI calls return "Error" and the pipeline continues with empty strings.
Quick debug check: Run print(os.getenv("CHATGPT_API_KEY")) at startup; ensure .env is present.
10. Large‑Node Recursion Overflow
Where it happens: process_large_node_recursively (lines 93‑119).
Why it happens: When a node contains many pages and the token budget is exceeded, the function recurses on every child. If max_page_num_each_node is set too low, recursion depth can hit Python limits.
Quick debug check: Enable logger.debug before the recursive call; watch the number of nodes created. Increase max_page_num_each_node or raise sys.setrecursionlimit.
Step-by-Step Debugging Workflow
-
Enable a logger –
logger = JsonLogger(doc)is already created inpage_index_main. Passloggerdown to every helper (check_title_appearance_in_start_concurrent,fix_incorrect_toc, etc.) and inspect the generated JSON log files under./logs. -
Capture raw LLM output – modify
ChatGPT_APIandChatGPT_API_asynctoreturn response.choices[0].message.content, response.choices[0].finish_reasonand log both. This isolates “finished vs truncated”. -
Validate page tokenization – run
get_page_tokens(pdf_path)manually and verify everypage_textlength. Iftoken_lengthis0, the PDF parser missed the page (common with scanned PDFs). Switch to the PyMuPDF parser (--pdf_parser pymupdf). -
Assert TOC‑page alignment – after
find_toc_pages, dumptoc_page_list. If empty, the TOC detector may be too strict; adjust the prompt or increaseopt.toc_check_page_num. -
Unit‑test individual functions – the repository ships with a
tests/folder containing expected JSON structures. Compare your intermediate structures (toc_with_page_number,matching_pairs, etc.) against the fixtures. -
Fallback path – if verification (
verify_toc) yields < 0.6 accuracy, the code automatically falls back to “process_no_toc”. You can force the “with page numbers” path by callingmeta_processorwithmode='process_toc_with_page_numbers'to see where it fails.
Practical Code Examples for Debugging
Run PageIndex with a Debug Logger
from pageindex.page_index import page_index_main
from pageindex.utils import JsonLogger
doc_path = "tests/pdfs/2023-annual-report.pdf"
logger = JsonLogger(doc_path) # creates ./logs/<pdf>_TIMESTAMP.json
result = page_index_main(doc_path, opt=None) # opt can be built with ConfigLoader if needed
print(result) # final tree
Inspect ./logs/2023-annual-report_*.json for every intermediate step (TOC detection, offset calc, etc.).
Print Raw LLM Response for a Failing TOC Detection
from pageindex.utils import ChatGPT_API_with_finish_reason
model = "gpt-4o-2024-11-20"
page_text = "<your page text here>"
response, reason = ChatGPT_API_with_finish_reason(model, f"Detect TOC: {page_text}")
print("Response:", response)
print("Finish reason:", reason) # 'finished' vs 'max_output_reached'
Verify Physical Indices Are Inside Document Bounds
from pageindex.utils import count_tokens
from pageindex.page_index import validate_and_truncate_physical_indices
toc = [...] # JSON from toc_extractor
pages = get_page_tokens("my.pdf")
validated = validate_and_truncate_physical_indices(toc, len(pages))
for item in validated:
if item.get("physical_index") is None:
print("Out‑of‑range:", item["title"])
Force the TOC‑with‑Page‑Numbers Path to Isolate Offset Bugs
from pageindex.page_index import meta_processor
# Manually load the PDF → page_list (list of (text, token_len))
page_list = get_page_tokens("my.pdf")
opt = ConfigLoader().load({"model": "gpt-4o-2024-11-20",
"toc_check_page_num": 30})
# Assuming you already have `toc_content` & `toc_page_list` from a prior run:
tree = await meta_processor(page_list,
mode="process_toc_with_page_numbers",
toc_content=toc_content,
toc_page_list=toc_page_list,
start_index=1,
opt=opt,
logger=JsonLogger("my.pdf"))
print(tree)
Unit‑Test the JSON Extractor on Malformed LLM Output
from pageindex.utils import extract_json
bad_response = "Here is the answer: {\"thinking\": \"...\", \"toc_detected\": \"yes\"}"
print(extract_json(bad_response)) # should return {'thinking':..., 'toc_detected':'yes'}
Key Source Files for Troubleshooting
| File | Role | Direct Link |
|---|---|---|
pageindex/page_index.py |
Core pipeline: TOC detection, index extraction, validation, verification, fixing, and recursive node processing. | page_index.py |
pageindex/utils.py |
Helper utilities – token counting, OpenAI API wrappers, JSON extraction, PDF parsing, logging, and tree‑building helpers. | utils.py |
pageindex/page_index_md.py |
Markdown‑specific extractor (header parsing, thinning, summary generation). | page_index_md.py |
run_pageindex.py |
CLI entry‑point that wires argument parsing to page_index_main. |
run_pageindex.py |
tests/ (e.g., tests/results/q1-fy25-earnings_structure.json) |
Reference outputs used for sanity‑checking the pipeline. | tests folder |
pageindex/config.yaml |
Default configuration (model, page limits, toggles). | config.yaml |
Summary
- TOC detection failures usually stem from truncated LLM outputs or incorrect
finish_reasonvalues intoc_detector_single_page. - Physical index errors (missing or out-of-range) occur in
validate_and_truncate_physical_indiceswhen the PDF parser drops pages or the LLM hallucinates page numbers. - Offset calculation bugs in
calculate_page_offsetbreak downstream navigation when matching pairs between raw and indexed TOCs are misaligned. - Token truncation in
post_processingcreates empty nodes or missing flags whenmax_token_num_each_nodeis set too aggressively. - Async exceptions in
check_title_appearance_in_start_concurrentcan silently return exception objects instead of results unlessreturn_exceptions=Trueis properly handled. - Environment and parsing errors (missing
CHATGPT_API_KEY, malformed JSON, or Markdown header issues) are caught by validating inputs at the entry points inutils.pyandpage_index_md.py.
Frequently Asked Questions
Why does PageIndex return an empty TOC even when my PDF has a clear table of contents?
This typically occurs in toc_detector_single_page when the LLM returns "toc_detected": "no" or the response is truncated (finish_reason = "length"). To debug, print the raw LLM response before extract_json processes it, and verify that the prompt contains the full page text. Increasing opt.toc_check_page_num may also help if the TOC appears later in the document.
How do I fix "physical_index out of range" errors when processing large PDFs?
These errors originate in validate_and_truncate_physical_indices (lines 14‑31) when the LLM assigns page numbers exceeding the actual page count. This usually indicates that the PDF parser dropped pages (common with scanned PDFs). Switch to the PyMuPDF parser using --pdf_parser pymupdf, or manually verify with get_page_tokens(pdf_path) that every page has a non-zero token length.
What causes incorrect page offsets that break section navigation?
Incorrect offsets stem from calculate_page_offset (lines 38‑50) when the "matching pairs" between the raw TOC and the physical-index TOC are misaligned. This happens if the LLM hallucinates section titles or if toc_index_extractor fails to tag pages correctly. Debug by printing matching_pairs and the resulting offset to ensure that physical_index and page fields refer to the same section title.
How can I prevent async exceptions from crashing the entire pipeline?
The check_title_appearance_in_start_concurrent function (lines 74‑102) uses asyncio.gather to run multiple LLM calls concurrently. If one call raises an exception, the gather returns an exception object instead of a result. Ensure you are using return_exceptions=True and add explicit logging inside the exception handler with logger.error(str(e)) to identify which specific call failed.
Why does the JSON extractor return empty dictionaries?
The extract_json utility in pageindex/utils.py (lines 25‑45) fails when the LLM prepends explanatory text or omits the json fence. The function then returns {} silently. To debug, print json_content immediately after the LLM call; if empty, inspect the raw response for stray explanations and consider tightening the prompt to forbid preamble text.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →