# How to Add Custom Node Summaries or Document Descriptions to PageIndex

> Learn how to add custom node summaries or document descriptions to PageIndex. Easily configure LLM generation or inject your own text for detailed content insights.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: how-to-guide
- Published: 2026-02-16

---

**Set `if_add_node_summary='yes'` and `if_add_doc_description='yes'` in the configuration to enable automatic LLM generation, or disable these flags and inject your own text directly into the returned structure.**

PageIndex (VectifyAI/PageIndex) is an open-source library that converts PDFs and Markdown files into hierarchical tree structures. When you need to add custom node summaries or document descriptions to PageIndex, the library offers both automatic generation via LLM prompts and manual injection methods that bypass AI calls entirely.

## Enabling Automatic Summaries in PDF and Markdown Workflows

PageIndex generates summaries through dedicated utility functions triggered by configuration flags. The implementation differs slightly between PDF and Markdown processing pipelines.

### PDF Processing with page_index

For PDF documents, the `page_index` function in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) checks the `if_add_node_summary` option before calling `generate_summaries_for_structure` from [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py).

```python
from pageindex import page_index

# Enable both node summaries and document description

result = page_index(
    doc="annual_report.pdf",
    model="gpt-4o-2024-11-20",
    if_add_node_summary="yes",      # Triggers generate_node_summary per node

    if_add_doc_description="yes",   # Triggers generate_doc_description

)

print(result["doc_description"])
for node in result["structure"]:
    print(f"{node['title']} → {node.get('summary', 'N/A')}")

```

You can also enable these features via the CLI entry point in [`run_pageindex.py`](https://github.com/VectifyAI/PageIndex/blob/main/run_pageindex.py):

```bash
python run_pageindex.py \
  --pdf_path annual_report.pdf \
  --if-add-node-summary yes \
  --if-add-doc-description yes

```

### Markdown Processing with md_to_tree

For Markdown files, the `md_to_tree` function in [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py) handles summary generation through `generate_summaries_for_structure_md`.

```python
import asyncio
from pageindex.page_index_md import md_to_tree

tree = asyncio.run(
    md_to_tree(
        md_path="documentation.md",
        if_add_node_summary="yes",
        if_add_doc_description="yes",
        summary_token_threshold=250,  # Only summarize sections > 250 tokens

        model="gpt-4o-2024-11-20",
    )
)

```

The `summary_token_threshold` parameter prevents the LLM from processing sections that are already concise, optimizing API costs and latency.

## Customizing LLM Prompts for Node Summaries

The default prompts reside as f-strings in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py). You can override these by subclassing or replacing the generation functions.

**Node summary prompt location:** Lines 5-12 in [`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py):

```python
prompt = f"""You are given a part of a document, your task is to generate a description … 
Partial Document Text: {node['text']}
Directly return the description, do not include any other text."""

```

**Document description prompt location:** Lines 50-57 in [`utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/utils.py):

```python
prompt = f"""Your are an expert in generating descriptions for a document.
You are given a structure of a document. Your task is to generate a one-sentence description …
Document Structure: {structure}
Directly return the description, do not include any other text."""

```

To use a custom prompt without modifying the source, create a wrapper function:

```python

# custom_utils.py

from pageindex.api import ChatGPT_API_async

async def generate_node_summary(node, model=None):
    """Custom summary generator using a bullet-point prompt."""
    prompt = f"""Summarise the following section as exactly 3 bullet points.
Section text:
{node['text']}
Output only the bullet points, one per line."""
    
    return await ChatGPT_API_async(model, prompt)

```

Then import this custom version before calling `page_index`:

```python
from custom_utils import generate_node_summary  # Overrides the import in pageindex

from pageindex import page_index

result = page_index("doc.pdf", if_add_node_summary="yes")

```

## Injecting Pre-Computed Summaries Without LLM Calls

To bypass automatic generation entirely, set the flags to `"no"` and manually inject your summaries into the structure dictionary.

```python
def inject_custom_summary(structure, node_id, summary_text):
    """Recursively locate node by ID and inject custom summary."""
    for node in structure:
        if node.get("node_id") == node_id:
            node["summary"] = summary_text
            return True
        if node.get("nodes"):
            if inject_custom_summary(node["nodes"], node_id, summary_text):
                return True
    return False

# Generate structure without LLM summaries

toc = page_index("report.pdf", if_add_node_summary="no")

# Add your own summaries

inject_custom_summary(toc["structure"], "1.2.3", "Custom algorithm explanation here.")
toc["doc_description"] = "Handcrafted document overview."

```

This approach is ideal when you have domain-specific summarization models or human-curated descriptions that outperform generic LLM outputs.

## Configuration Options Reference

| Parameter | Type | Location | Description |
|-----------|------|----------|-------------|
| `if_add_node_summary` | string (`"yes"`/`"no"`) | [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml), function args | Enables per-section summary generation via `generate_node_summary` |
| `if_add_doc_description` | string (`"yes"`/`"no"`) | [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml), function args | Enables one-sentence document description via `generate_doc_description` |
| `summary_token_threshold` | integer | Function args | Minimum token count required before a section is summarized (Markdown only) |
| `model` | string | Function args | LLM model identifier (e.g., `gpt-4o-2024-11-20`) used for generation |

Default values are defined in [`pageindex/config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/config.yaml) and can be overridden per call.

## Summary

- Set `if_add_node_summary='yes'` to generate per-section summaries automatically using the LLM defined in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py).
- Set `if_add_doc_description='yes'` to obtain a concise one-sentence overview of the entire document.
- Customize generation by overriding `generate_node_summary` or `generate_doc_description` with your own prompt logic.
- Disable automatic generation entirely and inject pre-computed summaries directly into the structure dictionary for full manual control.

## Frequently Asked Questions

### What is the difference between a node summary and a document description?

A **node summary** is a brief description of an individual section or leaf node within the document hierarchy, generated by `generate_node_summary` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py). A **document description** is a single-sentence overview of the entire document, produced by `generate_doc_description`. The former helps navigate large documents section-by-section, while the latter provides immediate context about the file as a whole.

### Can I use a different LLM model for summaries than for the main processing?

Yes. The `model` parameter passed to `page_index` or `md_to_tree` is forwarded to the summary generation functions. If you need different models for different tasks, you can subclass the generation functions in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) and pass a specific model identifier to `ChatGPT_API_async` or `ChatGPT_API` within your custom implementation.

### How do I completely disable automatic generation and use my own summaries?

Set both `if_add_node_summary="no"` and `if_add_doc_description="no"` when calling `page_index` or `md_to_tree`. The resulting structure will contain empty summary fields that you can populate manually. Use a recursive function like the `inject_custom_summary` example shown above to traverse the tree by `node_id` and assign your pre-computed strings to the `summary` keys.

### Where are the default prompts defined in the source code?

The default prompts are hard-coded as f-strings in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py). The node summary prompt appears at lines 5-12 inside the `generate_node_summary` function, while the document description prompt appears at lines 50-57 inside `generate_doc_description`. These strings are passed directly to the LLM API wrapper and can be edited in place or overridden via custom function imports.