Text-Based vs Vision-Based RAG Implementations in PageIndex: Architecture and Code Guide

PageIndex supports both text-based RAG using extracted PDF text with standard LLMs and vision-based RAG using raw page images with multimodal VLMs, sharing the same hierarchical tree index while differing only in content extraction and generation components.

PageIndex by VectifyAI is an open-source framework for building retrieval-augmented generation (RAG) systems over complex PDF documents. Understanding the differences between text-based and vision-based RAG implementations in PageIndex is essential for selecting the right pipeline—whether you need fast token-efficient retrieval on text-heavy documents or comprehensive visual understanding of scanned PDFs with complex layouts.

Shared Core Architecture

Both implementations rely on PageIndex’s hierarchical tree structure built from the document’s table of contents. In pageindex/page_index.py, the meta_processor function orchestrates TOC detection, tree construction, and page offset handling, producing a JSON tree where each node carries a physical_index (page number) linking it to the source material.

The key distinction lies in what content gets associated with these nodes: text-based pipelines store extracted strings, while vision-based pipelines reference image files, yet both use identical tree-reasoning logic for retrieval.

Text-Based RAG Pipeline

Input Processing with OCR-Free Extraction

Text-based RAG begins with extracting selectable text directly from PDF pages without optical character recognition. The utils.get_page_tokens function in pageindex/utils.py handles this extraction using libraries like PyPDF2 or PyMuPDF, then counts tokens using tiktoken to manage context windows.

Tree Construction with Text Nodes

When building the index via page_index() in pageindex/page_index.py, setting if_add_node_text="yes" populates each tree node with the extracted text content. The physical_index remains the critical link between the hierarchical structure and the actual page numbers.

LLM Retrieval and Generation

Retrieval uses a text-only LLM such as gpt-4o-2024-11-20. The model receives the tree structure (titles only) and selects relevant node IDs. For generation, the system retrieves the text snippets associated with those nodes and prompts the LLM with the question plus the concatenated text content.

from pageindex import page_index

# Build the tree and get a full structure with text nodes

result = page_index(
    doc="my_report.pdf",               # PDF path (or BytesIO)

    model="gpt-4o-2024-11-20",
    if_add_node_text="yes",            # keep extracted text for each node

)

print(result["structure"])   # hierarchical JSON with `text` fields

Vision-Based RAG Pipeline

Image Extraction with PyMuPDF

Vision-based RAG bypasses text extraction entirely. Instead, the extract_pdf_page_images helper function in cookbook/vision_RAG_pageindex.ipynb uses fitz (PyMuPDF) to render high-resolution JPEGs of each page at 2x zoom for clarity.

def extract_images(pdf_path, out_dir="pdf_images"):
    os.makedirs(out_dir, exist_ok=True)
    doc = fitz.open(pdf_path)
    images = {}
    for i, page in enumerate(doc):
        pix = page.get_pixmap(matrix=fitz.Matrix(2, 2))
        img_path = f"{out_dir}/page_{i+1}.jpg"
        pix.save(img_path)
        images[i+1] = img_path
    return images

Tree Construction Without Text Content

The tree is built using the same pageindex/page_index.py logic, but nodes contain only titles and physical_index references—no extracted text. The content remains in the image files, referenced by page number.

Multimodal Retrieval and Generation

Retrieval still uses a text-only LLM to select node IDs from the tree structure. However, generation employs a multimodal VLM such as gpt-4.1 or Qwen-VL. The get_page_images_for_nodes function maps the selected node IDs to their corresponding JPEG files using the physical_index, and the VLM receives both the question and the actual page images.


# Retrieve candidate nodes (text-only LLM)

search_prompt = f"""
You are given a question and a tree of section titles.
Find node IDs that likely contain the answer.

Question: {query}
Tree: {json.dumps(utils.remove_fields(tree, ["text"]))}
"""
candidate_ids = await call_vlm(search_prompt)

# Gather image paths for those nodes

retrieved_imgs = utils.get_page_images_for_nodes(
    candidate_ids, 
    utils.create_node_mapping(tree), 
    page_images
)

# Generate answer with multimodal VLM

answer_prompt = f"Answer the question using ONLY the supplied images.\n\nQuestion: {query}"
answer = await call_vlm(answer_prompt, image_paths=retrieved_imgs)

Critical Differences Between Text and Vision Pipelines

Understanding the differences between text-based and vision-based RAG implementations in PageIndex requires examining four key architectural dimensions:

Input Processing

  • Text-based: Uses utils.get_page_tokens in pageindex/utils.py to extract selectable text and count tokens with tiktoken.
  • Vision-based: Uses fitz (PyMuPDF) in cookbook/vision_RAG_pageindex.ipynb to render pages as high-resolution JPEGs without OCR.

Tokenization Strategy

  • Text-based: Explicit token counting via get_page_tokens to manage LLM context windows and prevent overflow.
  • Vision-based: No token counting during retrieval; the VLM’s internal vision encoder processes image content directly.

Generation Model Requirements

  • Text-based: Compatible with text-only LLMs like gpt-4o-2024-11-20.
  • Vision-based: Requires multimodal VLMs like gpt-4.1 or Qwen-VL capable of processing image inputs.

Content Fidelity

  • Text-based: Dependent on embedded PDF text quality; loses visual layout, figures, and complex table structures.
  • Vision-based: Preserves complete visual context including charts, diagrams, and formatting, but requires more compute resources.

Summary

  • PageIndex unifies both RAG approaches under a single hierarchical tree structure built by pageindex/page_index.py, using physical_index to link nodes to page numbers.
  • Text-based RAG extracts strings via utils.get_page_tokens, counts tokens with tiktoken, and uses standard LLMs for generation.
  • Vision-based RAG extracts images via PyMuPDF, skips token counting, and uses multimodal VLMs like gpt-4.1 to interpret visual content.
  • Both implementations share identical tree-reasoning retrieval logic but differ in content extraction and the final generation component.

Frequently Asked Questions

Can I switch between text and vision RAG without rebuilding the tree index?

No, while the hierarchical tree structure built by pageindex/page_index.py is shared between both approaches, the content associated with each node differs fundamentally. Text-based RAG requires if_add_node_text="yes" to populate nodes with extracted strings, while vision-based RAG relies on external image files mapped via physical_index. You must reprocess the document with the appropriate content extraction method when switching pipelines.

Which pipeline performs better on scanned documents?

Vision-based RAG significantly outperforms text-based approaches on scanned documents or PDFs with complex layouts. Since the vision pipeline in cookbook/vision_RAG_pageindex.ipynb processes raw page images without OCR, it preserves visual elements like charts, tables, and diagrams that text extraction often corrupts or omits. However, this comes at higher computational cost compared to the token-efficient text pipeline.

How does token counting differ between text and vision RAG?

In text-based RAG, utils.get_page_tokens in pageindex/utils.py explicitly counts tokens using tiktoken to manage context windows and prevent LLM overflow. Vision-based RAG eliminates explicit token counting during retrieval; instead, the multimodal VLM processes images through its internal vision encoder, handling content size constraints automatically. This architectural difference makes vision RAG simpler to implement but potentially more expensive at inference time.

What multimodal models are compatible with PageIndex vision RAG?

PageIndex’s vision-based RAG implementation in cookbook/vision_RAG_pageindex.ipynb supports any multimodal VLM capable of processing image inputs alongside text prompts. The reference implementation uses gpt-4.1 and Qwen-VL, but the pattern works with other vision-language models like gpt-4o-vision or Claude 3 Opus. The critical requirement is that the model must accept base64-encoded images or image URLs in its API payload, as demonstrated in the notebook’s call_vlm function.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →