How to Add Custom Node Summaries or Document Descriptions to PageIndex
Set if_add_node_summary='yes' and if_add_doc_description='yes' in the configuration to enable automatic LLM generation, or disable these flags and inject your own text directly into the returned structure.
PageIndex (VectifyAI/PageIndex) is an open-source library that converts PDFs and Markdown files into hierarchical tree structures. When you need to add custom node summaries or document descriptions to PageIndex, the library offers both automatic generation via LLM prompts and manual injection methods that bypass AI calls entirely.
Enabling Automatic Summaries in PDF and Markdown Workflows
PageIndex generates summaries through dedicated utility functions triggered by configuration flags. The implementation differs slightly between PDF and Markdown processing pipelines.
PDF Processing with page_index
For PDF documents, the page_index function in pageindex/page_index.py checks the if_add_node_summary option before calling generate_summaries_for_structure from pageindex/utils.py.
from pageindex import page_index
# Enable both node summaries and document description
result = page_index(
doc="annual_report.pdf",
model="gpt-4o-2024-11-20",
if_add_node_summary="yes", # Triggers generate_node_summary per node
if_add_doc_description="yes", # Triggers generate_doc_description
)
print(result["doc_description"])
for node in result["structure"]:
print(f"{node['title']} → {node.get('summary', 'N/A')}")
You can also enable these features via the CLI entry point in run_pageindex.py:
python run_pageindex.py \
--pdf_path annual_report.pdf \
--if-add-node-summary yes \
--if-add-doc-description yes
Markdown Processing with md_to_tree
For Markdown files, the md_to_tree function in pageindex/page_index_md.py handles summary generation through generate_summaries_for_structure_md.
import asyncio
from pageindex.page_index_md import md_to_tree
tree = asyncio.run(
md_to_tree(
md_path="documentation.md",
if_add_node_summary="yes",
if_add_doc_description="yes",
summary_token_threshold=250, # Only summarize sections > 250 tokens
model="gpt-4o-2024-11-20",
)
)
The summary_token_threshold parameter prevents the LLM from processing sections that are already concise, optimizing API costs and latency.
Customizing LLM Prompts for Node Summaries
The default prompts reside as f-strings in pageindex/utils.py. You can override these by subclassing or replacing the generation functions.
Node summary prompt location: Lines 5-12 in utils.py:
prompt = f"""You are given a part of a document, your task is to generate a description …
Partial Document Text: {node['text']}
Directly return the description, do not include any other text."""
Document description prompt location: Lines 50-57 in utils.py:
prompt = f"""Your are an expert in generating descriptions for a document.
You are given a structure of a document. Your task is to generate a one-sentence description …
Document Structure: {structure}
Directly return the description, do not include any other text."""
To use a custom prompt without modifying the source, create a wrapper function:
# custom_utils.py
from pageindex.api import ChatGPT_API_async
async def generate_node_summary(node, model=None):
"""Custom summary generator using a bullet-point prompt."""
prompt = f"""Summarise the following section as exactly 3 bullet points.
Section text:
{node['text']}
Output only the bullet points, one per line."""
return await ChatGPT_API_async(model, prompt)
Then import this custom version before calling page_index:
from custom_utils import generate_node_summary # Overrides the import in pageindex
from pageindex import page_index
result = page_index("doc.pdf", if_add_node_summary="yes")
Injecting Pre-Computed Summaries Without LLM Calls
To bypass automatic generation entirely, set the flags to "no" and manually inject your summaries into the structure dictionary.
def inject_custom_summary(structure, node_id, summary_text):
"""Recursively locate node by ID and inject custom summary."""
for node in structure:
if node.get("node_id") == node_id:
node["summary"] = summary_text
return True
if node.get("nodes"):
if inject_custom_summary(node["nodes"], node_id, summary_text):
return True
return False
# Generate structure without LLM summaries
toc = page_index("report.pdf", if_add_node_summary="no")
# Add your own summaries
inject_custom_summary(toc["structure"], "1.2.3", "Custom algorithm explanation here.")
toc["doc_description"] = "Handcrafted document overview."
This approach is ideal when you have domain-specific summarization models or human-curated descriptions that outperform generic LLM outputs.
Configuration Options Reference
| Parameter | Type | Location | Description |
|---|---|---|---|
if_add_node_summary |
string ("yes"/"no") |
config.yaml, function args |
Enables per-section summary generation via generate_node_summary |
if_add_doc_description |
string ("yes"/"no") |
config.yaml, function args |
Enables one-sentence document description via generate_doc_description |
summary_token_threshold |
integer | Function args | Minimum token count required before a section is summarized (Markdown only) |
model |
string | Function args | LLM model identifier (e.g., gpt-4o-2024-11-20) used for generation |
Default values are defined in pageindex/config.yaml and can be overridden per call.
Summary
- Set
if_add_node_summary='yes'to generate per-section summaries automatically using the LLM defined inpageindex/utils.py. - Set
if_add_doc_description='yes'to obtain a concise one-sentence overview of the entire document. - Customize generation by overriding
generate_node_summaryorgenerate_doc_descriptionwith your own prompt logic. - Disable automatic generation entirely and inject pre-computed summaries directly into the structure dictionary for full manual control.
Frequently Asked Questions
What is the difference between a node summary and a document description?
A node summary is a brief description of an individual section or leaf node within the document hierarchy, generated by generate_node_summary in pageindex/utils.py. A document description is a single-sentence overview of the entire document, produced by generate_doc_description. The former helps navigate large documents section-by-section, while the latter provides immediate context about the file as a whole.
Can I use a different LLM model for summaries than for the main processing?
Yes. The model parameter passed to page_index or md_to_tree is forwarded to the summary generation functions. If you need different models for different tasks, you can subclass the generation functions in pageindex/utils.py and pass a specific model identifier to ChatGPT_API_async or ChatGPT_API within your custom implementation.
How do I completely disable automatic generation and use my own summaries?
Set both if_add_node_summary="no" and if_add_doc_description="no" when calling page_index or md_to_tree. The resulting structure will contain empty summary fields that you can populate manually. Use a recursive function like the inject_custom_summary example shown above to traverse the tree by node_id and assign your pre-computed strings to the summary keys.
Where are the default prompts defined in the source code?
The default prompts are hard-coded as f-strings in pageindex/utils.py. The node summary prompt appears at lines 5-12 inside the generate_node_summary function, while the document description prompt appears at lines 50-57 inside generate_doc_description. These strings are passed directly to the LLM API wrapper and can be edited in place or overridden via custom function imports.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →