How Node Text Extraction and Summarization Work in PageIndex

PageIndex extracts node text by slicing concatenated page content based on hierarchical boundaries, then generates LLM summaries for nodes exceeding a configurable token threshold.

PageIndex (VectifyAI/PageIndex) transforms PDF documents into intelligent, hierarchical structures by combining precise node text extraction and summarization techniques. The pipeline tokenizes input pages, detects the table of contents to build a node tree, extracts raw text for each node, and conditionally summarizes oversized sections using OpenAI models.

The Pipeline Architecture

The end-to-end flow orchestrates multiple specialized components across three core files:

  1. Tokenization – get_page_tokens in pageindex/utils.py reads PDF pages using PyPDF2 or PyMuPDF, recording raw text and token counts via tiktoken.
  2. TOC Detection – LLM prompts (toc_detector_single_page, toc_transformer) in pageindex/page_index.py locate the table of contents and map entries to physical page indices.
  3. Tree Construction – post_processing converts flat TOC items into a nested hierarchy where each node stores start_index and end_index page boundaries.
  4. Text Extraction – add_node_text slices the page list using node boundaries and concatenates page texts into node['text'].
  5. Summarization – get_node_summary checks token counts against summary_token_threshold, calling generate_node_summary for oversized nodes.

Extracting Raw Node Text

The node text extraction process operates in pageindex/utils.py through the add_node_text function. After the hierarchical tree is built, this utility traverses every node and performs boundary-based text slicing.

For each node, the system:

  • Retrieves the start_index and end_index from the node's metadata
  • Slices the pre-tokenized page list to capture all pages within the section boundaries
  • Concatenates the page_text fields into a single string
  • Stores the result in node['text']

This approach ensures that parent nodes contain the complete text of their entire subtree, while leaf nodes contain only their specific section content.

Summarization Logic and Thresholds

Node summarization prevents context window overflow while preserving semantic meaning. The implementation in pageindex/page_index_md.py and pageindex/utils.py uses token-aware conditional processing.

The get_node_summary function:

  1. Calculates the token count of node['text'] using count_tokens
  2. Compares the count against summary_token_threshold (default 200 tokens in demo configurations)
  3. For nodes exceeding the threshold, invokes generate_node_summary to create a concise LLM-generated description
  4. Stores summaries in node['summary'] for leaf nodes or node['prefix_summary'] for internal nodes

Short nodes retain their original text unchanged, eliminating unnecessary API calls for content that already fits within context limits.

Configuration and Customization

Control the node text extraction and summarization behavior through config.yaml or runtime parameters:

if_add_node_text: "yes"          # Enable raw text extraction

if_add_node_summary: "yes"       # Enable LLM summarization

summary_token_threshold: 200     # Token limit before summarization triggers

model: "gpt-4o-2024-11-20"     # OpenAI model for summary generation

Override defaults programmatically when calling the high-level API:

from pageindex import page_index

result = page_index(
    "document.pdf",
    if_add_node_text="yes",
    if_add_node_summary="yes",
    summary_token_threshold=300,
    model="gpt-4o-2024-11-20"
)

Practical Implementation Examples

Extract Text and Summaries from a PDF

from pageindex import page_index

# Process a financial report with full extraction

result = page_index("annual-report.pdf")

# Access the first section's content

first_node = result["structure"][0]
print(f"Section: {first_node['title']}")
print(f"Text length: {len(first_node['text'])} chars")
print(f"Summary: {first_node.get('summary', 'N/A')}")

Process Markdown with Custom Thresholds

import asyncio
from pageindex.page_index_md import md_to_tree

async def process_markdown():
    tree = await md_to_tree(
        md_path="documentation.md",
        if_add_node_summary="yes",
        summary_token_threshold=150,
        model="gpt-4o-2024-11-20"
    )
    
    # Inspect generated summaries

    for node in tree["structure"]:
        if "summary" in node:
            print(f"{node['title']}: {node['summary'][:100]}...")

asyncio.run(process_markdown())

Summary

  • PageIndex extracts node text by slicing pre-tokenized pages using hierarchical boundary indices (start_index/end_index) stored in each node.
  • The add_node_text function in pageindex/utils.py concatenates page texts within node boundaries to populate node['text'].
  • Summarization occurs conditionally when get_node_summary detects token counts exceeding summary_token_threshold, triggering generate_node_summary to create concise LLM descriptions.
  • Leaf nodes store summaries in node['summary'] while internal nodes use node['prefix_summary'] to provide context without duplicating full child content.
  • Configuration through config.yaml or runtime parameters controls whether text extraction and summarization execute, along with model selection and token thresholds.

Frequently Asked Questions

What determines whether a node gets summarized or keeps its full text?

The summary_token_threshold configuration parameter controls this decision. When get_node_summary calculates that a node's text exceeds this token count (default 200 in demo configurations), it triggers generate_node_summary to create an LLM-generated summary. Nodes with text below the threshold retain their original content unchanged, avoiding unnecessary API calls.

How does PageIndex handle text extraction for nested hierarchical nodes?

PageIndex uses boundary-based slicing where each node stores start_index and end_index referencing the global page list. The add_node_text function concatenates all page texts within these indices, meaning parent nodes automatically include the complete text of their entire subtree, while leaf nodes contain only their specific section. This approach ensures semantic continuity across hierarchical boundaries.

Can I use PageIndex summarization with Markdown files instead of PDFs?

Yes, the pageindex/page_index_md.py module provides md_to_tree and generate_summaries_for_structure_md functions that mirror the PDF pipeline. These functions parse Markdown headings to build the hierarchical tree, then apply the same token-counting logic and generate_node_summary calls to produce summaries for large sections, storing results in node['summary'] or node['prefix_summary'] depending on node type.

What LLM models does PageIndex support for summary generation?

PageIndex uses OpenAI's chat models through the generate_node_summary function in pageindex/utils.py. The default configuration specifies gpt-4o-2024-11-20, but you can override this via the model parameter in config.yaml or runtime options. The implementation uses the OpenAI API's chat completions endpoint, making it compatible with any OpenAI model supporting the chat interface, including GPT-4, GPT-4o, and GPT-3.5-turbo variants.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →