Performance Implications of Choosing Different Models in PageIndex: A Technical Analysis

Choosing different LLM models in PageIndex directly impacts tokenization accuracy, API latency, operational costs, and TOC extraction quality through model-specific encoders and context window constraints.

PageIndex, an open-source hierarchical table-of-contents (TOC) generator for PDFs from VectifyAI, relies on repeated calls to OpenAI's Chat Completion API to analyze document structure. The specific model you configure—whether gpt-4o-2024-11-20, gpt-4o-mini, or gpt-3.5-turbo—determines how the system tokenizes text, manages context windows, and balances speed against extraction accuracy.

How Model Selection Affects Tokenization and Pagination

PageIndex uses tiktoken to count tokens before sending prompts to the API. In pageindex/utils.py, the count_tokens function calls tiktoken.encoding_for_model(model) to obtain the correct encoder for your specific model.


# From pageindex/utils.py lines 22-27

def count_tokens(text, model):
    try:
        encoding = tiktoken.encoding_for_model(model)
    except KeyError:
        encoding = tiktoken.get_encoding("cl100k_base")
    return len(encoding.encode(text))

Different models use different tokenizers. For example, gpt-4o uses the o200k_base encoder, while older models use cl100k_base. This affects the max_token_num_each_node calculation in pageindex/page_index.py. If you switch models without adjusting this parameter, you risk either under-utilizing available context window or exceeding the model's limits, causing prompt truncation.

Latency and Throughput Considerations

PageIndex generates TOCs by issuing numerous short prompts for title checking, TOC extraction, and summary generation. The cumulative latency varies significantly based on model choice:

  • gpt-4o-2024-11-20: 2-4 seconds per call, highest accuracy
  • gpt-4o-mini: 1-2 seconds per call, balanced performance
  • gpt-3.5-turbo: <1 second per call, fastest but less reliable

In pageindex/utils.py, all LLM calls (ChatGPT_API, ChatGPT_API_async, ChatGPT_API_with_finish_reason) forward the model argument directly to the OpenAI client:


# From pageindex/utils.py lines 29-48

def ChatGPT_API_async(prompt, model, temperature=0.0, ...):
    # ... setup ...

    response = client.chat.completions.create(
        model=model,  # Model passed here

        messages=messages,
        # ...

    )
    return response.choices[0].message.content

For processing large PDFs with hundreds of pages, choosing gpt-4o-mini can reduce total processing time from minutes to seconds compared to gpt-4o, while gpt-3.5-turbo offers the fastest throughput for batch operations where perfect accuracy is less critical.

Cost Analysis and Token Economics

PageIndex processes every page twice: once to count tokens using count_tokens, and once to query the model. The total cost scales with both input and output token counts, making model selection a primary cost driver.

The model configuration originates in pageindex/config.yaml:


# From pageindex/config.yaml line 1

model: "gpt-4o-2024-11-20"

This value propagates through page_index_main in pageindex/page_index.py (lines 58-66) to every API call. When you specify a cheaper model like gpt-3.5-turbo, you reduce per-token costs by approximately 90% compared to gpt-4o, though you may sacrifice extraction accuracy on complex documents with ambiguous headings.

Accuracy and Context Window Constraints

More capable models better understand ambiguous headings, fuzzy matches, and long prompts. PageIndex uses the check_title_appearance function in pageindex/page_index.py (lines 13-40) to perform fuzzy matching for title validation:


# Conceptual flow from pageindex/page_index.py

def check_title_appearance(title, page_text, model):
    prompt = f"Does the title '{title}' appear in the following text..."
    # Calls ChatGPT_API_async with the specified model

    return ChatGPT_API_async(prompt, model=model)

Weaker models may return false negatives during this check, causing the final TOC to miss sections. Additionally, the default token limit for your chosen model determines the maximum prompt size. PageIndex splits large documents in page_list_to_group_text (lines 18-50 of pageindex/page_index.py) based on max_token_num_each_node (default 20,000). If you switch to a model with a smaller context window (e.g., 8k), you must lower max_token_num_each_node or increase the number of groups to prevent prompt truncation.

Practical Configuration Examples

Running PageIndex with the Default Model

from pageindex.page_index import page_index

# Uses config.yaml model = "gpt-4o-2024-11-20"

toc = page_index("annual_report.pdf")
print(toc["structure"])

Switching to a Cost-Effective Model

from pageindex.page_index import page_index

# Override for faster, cheaper processing

toc = page_index(
    "simple_document.pdf",
    model="gpt-4o-mini",
    max_token_num_each_node=16000  # Adjust for model context window

)

Optimizing for Speed with GPT-3.5-Turbo

from pageindex.page_index import page_index

# Fastest option for batch processing

toc = page_index(
    "batch_001.pdf",
    model="gpt-3.5-turbo",
    max_token_num_each_node=4000  # Critical: respect 4k context window

)

Measuring Performance Impact

import time
from pageindex.page_index import page_index
from pageindex.utils import count_tokens

pdf_path = "technical_manual.pdf"

# Benchmark GPT-4o

start = time.time()
result_4o = page_index(pdf_path, model="gpt-4o-2024-11-20")
time_4o = time.time() - start

# Benchmark GPT-4o-mini

start = time.time()
result_mini = page_index(pdf_path, model="gpt-4o-mini")
time_mini = time.time() - start

print(f"GPT-4o time: {time_4o:.2f}s")
print(f"GPT-4o-mini time: {time_mini:.2f}s")
print(f"Speed improvement: {(time_4o/time_mini - 1)*100:.0f}%")

Summary

  • Tokenization varies by model: PageIndex uses tiktoken.encoding_for_model() in pageindex/utils.py to ensure accurate token counts, but different models encode the same text with different token lengths, affecting pagination logic.

  • Latency scales with model capability: Each LLM call in pageindex/utils.py forwards your chosen model to the OpenAI client; expect 2-4 seconds per call for gpt-4o, 1-2 seconds for gpt-4o-mini, and sub-second for gpt-3.5-turbo.

  • Cost doubles through token counting: PageIndex processes each page twice (once for count_tokens, once for generation), making the per-token price differential between models significant for large documents.

  • Accuracy depends on model strength: The check_title_appearance function in pageindex/page_index.py relies on the model's ability to perform fuzzy matching; weaker models increase false negatives in TOC extraction.

  • Context windows require configuration: When using models with smaller context windows (e.g., 8k), you must adjust max_token_num_each_node in page_list_to_group_text (lines 18-50 of pageindex/page_index.py) to prevent prompt truncation.

Frequently Asked Questions

Does PageIndex support models other than OpenAI GPT?

Currently, PageIndex is designed specifically for OpenAI's Chat Completion API. The ChatGPT_API and ChatGPT_API_async functions in pageindex/utils.py directly instantiate the OpenAI client and pass the model string to it. While the architecture could theoretically support other providers through adapter patterns, the current implementation in pageindex/utils.py (lines 29-48) hardcodes OpenAI-specific client initialization and response handling.

How do I switch from GPT-4o to GPT-4o-mini in PageIndex?

You can override the default model by passing the model parameter directly to the page_index function or by modifying pageindex/config.yaml. For programmatic switching, use page_index("document.pdf", model="gpt-4o-mini"). The page_index_main function in pageindex/page_index.py (lines 58-66) merges your configuration with defaults and propagates the model name to all subsequent API calls in utils.py.

Why does my TOC extraction fail with certain models?

TOC extraction failures typically occur when using models with small context windows (like older GPT-3.5 variants with 4k limits) without adjusting max_token_num_each_node. The page_list_to_group_text function in pageindex/page_index.py (lines 18-50) splits documents based on this parameter; if the value exceeds the model's context window, prompts get truncated, causing the check_title_appearance function to miss sections or return malformed JSON. Additionally, weaker models may fail at the fuzzy matching required for title verification.

What is the optimal max_token_num_each_node for GPT-3.5-Turbo?

For gpt-3.5-turbo with its 4,096 token context window, set max_token_num_each_node to approximately 2,000-3,000 tokens to leave room for the system prompt and response tokens. In pageindex/page_index.py, the page_list_to_group_text function uses this parameter to group pages before sending them to the LLM. Setting this too high (e.g., the default 20,000) will cause prompt truncation and failed extractions, while setting it too low increases API call count and cost.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →