# How Node Text Extraction and Summarization Work in PageIndex

> Discover how PageIndex performs node text extraction and summarization by slicing page content and generating LLM summaries. Learn about the PageIndex repository.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: internals
- Published: 2026-02-16

---

**PageIndex extracts node text by slicing concatenated page content based on hierarchical boundaries, then generates LLM summaries for nodes exceeding a configurable token threshold.**

PageIndex (VectifyAI/PageIndex) transforms PDF documents into intelligent, hierarchical structures by combining precise **node text extraction and summarization** techniques. The pipeline tokenizes input pages, detects the table of contents to build a node tree, extracts raw text for each node, and conditionally summarizes oversized sections using OpenAI models.

## The Pipeline Architecture

The end-to-end flow orchestrates multiple specialized components across three core files:

1. **Tokenization** – `get_page_tokens` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) reads PDF pages using PyPDF2 or PyMuPDF, recording raw text and token counts via tiktoken.
2. **TOC Detection** – LLM prompts (`toc_detector_single_page`, `toc_transformer`) in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) locate the table of contents and map entries to physical page indices.
3. **Tree Construction** – `post_processing` converts flat TOC items into a nested hierarchy where each node stores `start_index` and `end_index` page boundaries.
4. **Text Extraction** – `add_node_text` slices the page list using node boundaries and concatenates page texts into `node['text']`.
5. **Summarization** – `get_node_summary` checks token counts against `summary_token_threshold`, calling `generate_node_summary` for oversized nodes.

## Extracting Raw Node Text

The **node text extraction** process operates in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) through the `add_node_text` function. After the hierarchical tree is built, this utility traverses every node and performs boundary-based text slicing.

For each node, the system:
- Retrieves the `start_index` and `end_index` from the node's metadata
- Slices the pre-tokenized page list to capture all pages within the section boundaries
- Concatenates the `page_text` fields into a single string
- Stores the result in `node['text']`

This approach ensures that parent nodes contain the complete text of their entire subtree, while leaf nodes contain only their specific section content.

## Summarization Logic and Thresholds

**Node summarization** prevents context window overflow while preserving semantic meaning. The implementation in [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py) and [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) uses token-aware conditional processing.

The `get_node_summary` function:
1. Calculates the token count of `node['text']` using `count_tokens`
2. Compares the count against `summary_token_threshold` (default 200 tokens in demo configurations)
3. For nodes exceeding the threshold, invokes `generate_node_summary` to create a concise LLM-generated description
4. Stores summaries in `node['summary']` for leaf nodes or `node['prefix_summary']` for internal nodes

Short nodes retain their original text unchanged, eliminating unnecessary API calls for content that already fits within context limits.

## Configuration and Customization

Control the **node text extraction and summarization** behavior through [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml) or runtime parameters:

```yaml
if_add_node_text: "yes"          # Enable raw text extraction

if_add_node_summary: "yes"       # Enable LLM summarization

summary_token_threshold: 200     # Token limit before summarization triggers

model: "gpt-4o-2024-11-20"     # OpenAI model for summary generation

```

Override defaults programmatically when calling the high-level API:

```python
from pageindex import page_index

result = page_index(
    "document.pdf",
    if_add_node_text="yes",
    if_add_node_summary="yes",
    summary_token_threshold=300,
    model="gpt-4o-2024-11-20"
)

```

## Practical Implementation Examples

### Extract Text and Summaries from a PDF

```python
from pageindex import page_index

# Process a financial report with full extraction

result = page_index("annual-report.pdf")

# Access the first section's content

first_node = result["structure"][0]
print(f"Section: {first_node['title']}")
print(f"Text length: {len(first_node['text'])} chars")
print(f"Summary: {first_node.get('summary', 'N/A')}")

```

### Process Markdown with Custom Thresholds

```python
import asyncio
from pageindex.page_index_md import md_to_tree

async def process_markdown():
    tree = await md_to_tree(
        md_path="documentation.md",
        if_add_node_summary="yes",
        summary_token_threshold=150,
        model="gpt-4o-2024-11-20"
    )
    
    # Inspect generated summaries

    for node in tree["structure"]:
        if "summary" in node:
            print(f"{node['title']}: {node['summary'][:100]}...")

asyncio.run(process_markdown())

```

## Summary

- **PageIndex** extracts node text by slicing pre-tokenized pages using hierarchical boundary indices (`start_index`/`end_index`) stored in each node.
- The `add_node_text` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) concatenates page texts within node boundaries to populate `node['text']`.
- **Summarization** occurs conditionally when `get_node_summary` detects token counts exceeding `summary_token_threshold`, triggering `generate_node_summary` to create concise LLM descriptions.
- Leaf nodes store summaries in `node['summary']` while internal nodes use `node['prefix_summary']` to provide context without duplicating full child content.
- Configuration through [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml) or runtime parameters controls whether text extraction and summarization execute, along with model selection and token thresholds.

## Frequently Asked Questions

### What determines whether a node gets summarized or keeps its full text?

The `summary_token_threshold` configuration parameter controls this decision. When `get_node_summary` calculates that a node's text exceeds this token count (default 200 in demo configurations), it triggers `generate_node_summary` to create an LLM-generated summary. Nodes with text below the threshold retain their original content unchanged, avoiding unnecessary API calls.

### How does PageIndex handle text extraction for nested hierarchical nodes?

PageIndex uses boundary-based slicing where each node stores `start_index` and `end_index` referencing the global page list. The `add_node_text` function concatenates all page texts within these indices, meaning parent nodes automatically include the complete text of their entire subtree, while leaf nodes contain only their specific section. This approach ensures semantic continuity across hierarchical boundaries.

### Can I use PageIndex summarization with Markdown files instead of PDFs?

Yes, the [`pageindex/page_index_md.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index_md.py) module provides `md_to_tree` and `generate_summaries_for_structure_md` functions that mirror the PDF pipeline. These functions parse Markdown headings to build the hierarchical tree, then apply the same token-counting logic and `generate_node_summary` calls to produce summaries for large sections, storing results in `node['summary']` or `node['prefix_summary']` depending on node type.

### What LLM models does PageIndex support for summary generation?

PageIndex uses OpenAI's chat models through the `generate_node_summary` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py). The default configuration specifies `gpt-4o-2024-11-20`, but you can override this via the `model` parameter in [`config.yaml`](https://github.com/VectifyAI/PageIndex/blob/main/config.yaml) or runtime options. The implementation uses the OpenAI API's chat completions endpoint, making it compatible with any OpenAI model supporting the chat interface, including GPT-4, GPT-4o, and GPT-3.5-turbo variants.