How to Add Custom Node Summaries or Document Descriptions to PageIndex

Set if_add_node_summary='yes' and if_add_doc_description='yes' in the configuration to enable automatic LLM generation, or disable these flags and inject your own text directly into the returned structure.

PageIndex (VectifyAI/PageIndex) is an open-source library that converts PDFs and Markdown files into hierarchical tree structures. When you need to add custom node summaries or document descriptions to PageIndex, the library offers both automatic generation via LLM prompts and manual injection methods that bypass AI calls entirely.

Enabling Automatic Summaries in PDF and Markdown Workflows

PageIndex generates summaries through dedicated utility functions triggered by configuration flags. The implementation differs slightly between PDF and Markdown processing pipelines.

PDF Processing with page_index

For PDF documents, the page_index function in pageindex/page_index.py checks the if_add_node_summary option before calling generate_summaries_for_structure from pageindex/utils.py.

from pageindex import page_index

# Enable both node summaries and document description

result = page_index(
    doc="annual_report.pdf",
    model="gpt-4o-2024-11-20",
    if_add_node_summary="yes",      # Triggers generate_node_summary per node

    if_add_doc_description="yes",   # Triggers generate_doc_description

)

print(result["doc_description"])
for node in result["structure"]:
    print(f"{node['title']} → {node.get('summary', 'N/A')}")

You can also enable these features via the CLI entry point in run_pageindex.py:

python run_pageindex.py \
  --pdf_path annual_report.pdf \
  --if-add-node-summary yes \
  --if-add-doc-description yes

Markdown Processing with md_to_tree

For Markdown files, the md_to_tree function in pageindex/page_index_md.py handles summary generation through generate_summaries_for_structure_md.

import asyncio
from pageindex.page_index_md import md_to_tree

tree = asyncio.run(
    md_to_tree(
        md_path="documentation.md",
        if_add_node_summary="yes",
        if_add_doc_description="yes",
        summary_token_threshold=250,  # Only summarize sections > 250 tokens

        model="gpt-4o-2024-11-20",
    )
)

The summary_token_threshold parameter prevents the LLM from processing sections that are already concise, optimizing API costs and latency.

Customizing LLM Prompts for Node Summaries

The default prompts reside as f-strings in pageindex/utils.py. You can override these by subclassing or replacing the generation functions.

Node summary prompt location: Lines 5-12 in utils.py:

prompt = f"""You are given a part of a document, your task is to generate a description … 
Partial Document Text: {node['text']}
Directly return the description, do not include any other text."""

Document description prompt location: Lines 50-57 in utils.py:

prompt = f"""Your are an expert in generating descriptions for a document.
You are given a structure of a document. Your task is to generate a one-sentence description …
Document Structure: {structure}
Directly return the description, do not include any other text."""

To use a custom prompt without modifying the source, create a wrapper function:


# custom_utils.py

from pageindex.api import ChatGPT_API_async

async def generate_node_summary(node, model=None):
    """Custom summary generator using a bullet-point prompt."""
    prompt = f"""Summarise the following section as exactly 3 bullet points.
Section text:
{node['text']}
Output only the bullet points, one per line."""
    
    return await ChatGPT_API_async(model, prompt)

Then import this custom version before calling page_index:

from custom_utils import generate_node_summary  # Overrides the import in pageindex

from pageindex import page_index

result = page_index("doc.pdf", if_add_node_summary="yes")

Injecting Pre-Computed Summaries Without LLM Calls

To bypass automatic generation entirely, set the flags to "no" and manually inject your summaries into the structure dictionary.

def inject_custom_summary(structure, node_id, summary_text):
    """Recursively locate node by ID and inject custom summary."""
    for node in structure:
        if node.get("node_id") == node_id:
            node["summary"] = summary_text
            return True
        if node.get("nodes"):
            if inject_custom_summary(node["nodes"], node_id, summary_text):
                return True
    return False

# Generate structure without LLM summaries

toc = page_index("report.pdf", if_add_node_summary="no")

# Add your own summaries

inject_custom_summary(toc["structure"], "1.2.3", "Custom algorithm explanation here.")
toc["doc_description"] = "Handcrafted document overview."

This approach is ideal when you have domain-specific summarization models or human-curated descriptions that outperform generic LLM outputs.

Configuration Options Reference

Parameter Type Location Description
if_add_node_summary string ("yes"/"no") config.yaml, function args Enables per-section summary generation via generate_node_summary
if_add_doc_description string ("yes"/"no") config.yaml, function args Enables one-sentence document description via generate_doc_description
summary_token_threshold integer Function args Minimum token count required before a section is summarized (Markdown only)
model string Function args LLM model identifier (e.g., gpt-4o-2024-11-20) used for generation

Default values are defined in pageindex/config.yaml and can be overridden per call.

Summary

  • Set if_add_node_summary='yes' to generate per-section summaries automatically using the LLM defined in pageindex/utils.py.
  • Set if_add_doc_description='yes' to obtain a concise one-sentence overview of the entire document.
  • Customize generation by overriding generate_node_summary or generate_doc_description with your own prompt logic.
  • Disable automatic generation entirely and inject pre-computed summaries directly into the structure dictionary for full manual control.

Frequently Asked Questions

What is the difference between a node summary and a document description?

A node summary is a brief description of an individual section or leaf node within the document hierarchy, generated by generate_node_summary in pageindex/utils.py. A document description is a single-sentence overview of the entire document, produced by generate_doc_description. The former helps navigate large documents section-by-section, while the latter provides immediate context about the file as a whole.

Can I use a different LLM model for summaries than for the main processing?

Yes. The model parameter passed to page_index or md_to_tree is forwarded to the summary generation functions. If you need different models for different tasks, you can subclass the generation functions in pageindex/utils.py and pass a specific model identifier to ChatGPT_API_async or ChatGPT_API within your custom implementation.

How do I completely disable automatic generation and use my own summaries?

Set both if_add_node_summary="no" and if_add_doc_description="no" when calling page_index or md_to_tree. The resulting structure will contain empty summary fields that you can populate manually. Use a recursive function like the inject_custom_summary example shown above to traverse the tree by node_id and assign your pre-computed strings to the summary keys.

Where are the default prompts defined in the source code?

The default prompts are hard-coded as f-strings in pageindex/utils.py. The node summary prompt appears at lines 5-12 inside the generate_node_summary function, while the document description prompt appears at lines 50-57 inside generate_doc_description. These strings are passed directly to the LLM API wrapper and can be edited in place or overridden via custom function imports.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →