How Node Text Extraction and Summarization Work in PageIndex
PageIndex extracts node text by slicing concatenated page content based on hierarchical boundaries, then generates LLM summaries for nodes exceeding a configurable token threshold.
PageIndex (VectifyAI/PageIndex) transforms PDF documents into intelligent, hierarchical structures by combining precise node text extraction and summarization techniques. The pipeline tokenizes input pages, detects the table of contents to build a node tree, extracts raw text for each node, and conditionally summarizes oversized sections using OpenAI models.
The Pipeline Architecture
The end-to-end flow orchestrates multiple specialized components across three core files:
- Tokenization –
get_page_tokensinpageindex/utils.pyreads PDF pages using PyPDF2 or PyMuPDF, recording raw text and token counts via tiktoken. - TOC Detection – LLM prompts (
toc_detector_single_page,toc_transformer) inpageindex/page_index.pylocate the table of contents and map entries to physical page indices. - Tree Construction –
post_processingconverts flat TOC items into a nested hierarchy where each node storesstart_indexandend_indexpage boundaries. - Text Extraction –
add_node_textslices the page list using node boundaries and concatenates page texts intonode['text']. - Summarization –
get_node_summarychecks token counts againstsummary_token_threshold, callinggenerate_node_summaryfor oversized nodes.
Extracting Raw Node Text
The node text extraction process operates in pageindex/utils.py through the add_node_text function. After the hierarchical tree is built, this utility traverses every node and performs boundary-based text slicing.
For each node, the system:
- Retrieves the
start_indexandend_indexfrom the node's metadata - Slices the pre-tokenized page list to capture all pages within the section boundaries
- Concatenates the
page_textfields into a single string - Stores the result in
node['text']
This approach ensures that parent nodes contain the complete text of their entire subtree, while leaf nodes contain only their specific section content.
Summarization Logic and Thresholds
Node summarization prevents context window overflow while preserving semantic meaning. The implementation in pageindex/page_index_md.py and pageindex/utils.py uses token-aware conditional processing.
The get_node_summary function:
- Calculates the token count of
node['text']usingcount_tokens - Compares the count against
summary_token_threshold(default 200 tokens in demo configurations) - For nodes exceeding the threshold, invokes
generate_node_summaryto create a concise LLM-generated description - Stores summaries in
node['summary']for leaf nodes ornode['prefix_summary']for internal nodes
Short nodes retain their original text unchanged, eliminating unnecessary API calls for content that already fits within context limits.
Configuration and Customization
Control the node text extraction and summarization behavior through config.yaml or runtime parameters:
if_add_node_text: "yes" # Enable raw text extraction
if_add_node_summary: "yes" # Enable LLM summarization
summary_token_threshold: 200 # Token limit before summarization triggers
model: "gpt-4o-2024-11-20" # OpenAI model for summary generation
Override defaults programmatically when calling the high-level API:
from pageindex import page_index
result = page_index(
"document.pdf",
if_add_node_text="yes",
if_add_node_summary="yes",
summary_token_threshold=300,
model="gpt-4o-2024-11-20"
)
Practical Implementation Examples
Extract Text and Summaries from a PDF
from pageindex import page_index
# Process a financial report with full extraction
result = page_index("annual-report.pdf")
# Access the first section's content
first_node = result["structure"][0]
print(f"Section: {first_node['title']}")
print(f"Text length: {len(first_node['text'])} chars")
print(f"Summary: {first_node.get('summary', 'N/A')}")
Process Markdown with Custom Thresholds
import asyncio
from pageindex.page_index_md import md_to_tree
async def process_markdown():
tree = await md_to_tree(
md_path="documentation.md",
if_add_node_summary="yes",
summary_token_threshold=150,
model="gpt-4o-2024-11-20"
)
# Inspect generated summaries
for node in tree["structure"]:
if "summary" in node:
print(f"{node['title']}: {node['summary'][:100]}...")
asyncio.run(process_markdown())
Summary
- PageIndex extracts node text by slicing pre-tokenized pages using hierarchical boundary indices (
start_index/end_index) stored in each node. - The
add_node_textfunction inpageindex/utils.pyconcatenates page texts within node boundaries to populatenode['text']. - Summarization occurs conditionally when
get_node_summarydetects token counts exceedingsummary_token_threshold, triggeringgenerate_node_summaryto create concise LLM descriptions. - Leaf nodes store summaries in
node['summary']while internal nodes usenode['prefix_summary']to provide context without duplicating full child content. - Configuration through
config.yamlor runtime parameters controls whether text extraction and summarization execute, along with model selection and token thresholds.
Frequently Asked Questions
What determines whether a node gets summarized or keeps its full text?
The summary_token_threshold configuration parameter controls this decision. When get_node_summary calculates that a node's text exceeds this token count (default 200 in demo configurations), it triggers generate_node_summary to create an LLM-generated summary. Nodes with text below the threshold retain their original content unchanged, avoiding unnecessary API calls.
How does PageIndex handle text extraction for nested hierarchical nodes?
PageIndex uses boundary-based slicing where each node stores start_index and end_index referencing the global page list. The add_node_text function concatenates all page texts within these indices, meaning parent nodes automatically include the complete text of their entire subtree, while leaf nodes contain only their specific section. This approach ensures semantic continuity across hierarchical boundaries.
Can I use PageIndex summarization with Markdown files instead of PDFs?
Yes, the pageindex/page_index_md.py module provides md_to_tree and generate_summaries_for_structure_md functions that mirror the PDF pipeline. These functions parse Markdown headings to build the hierarchical tree, then apply the same token-counting logic and generate_node_summary calls to produce summaries for large sections, storing results in node['summary'] or node['prefix_summary'] depending on node type.
What LLM models does PageIndex support for summary generation?
PageIndex uses OpenAI's chat models through the generate_node_summary function in pageindex/utils.py. The default configuration specifies gpt-4o-2024-11-20, but you can override this via the model parameter in config.yaml or runtime options. The implementation uses the OpenAI API's chat completions endpoint, making it compatible with any OpenAI model supporting the chat interface, including GPT-4, GPT-4o, and GPT-3.5-turbo variants.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →