# Text-Based vs Vision-Based RAG Implementations in PageIndex: Architecture and Code Guide

> Explore text-based vs vision-based RAG in VectifyAI PageIndex. Learn about architectural differences, code implementation, and discover how to leverage both standard LLMs and multimodal VLMs efficiently.

- Repository: [Vectify AI/PageIndex](https://github.com/vectifyai/pageindex)
- Tags: architecture
- Published: 2026-02-16

---

**PageIndex supports both text-based RAG using extracted PDF text with standard LLMs and vision-based RAG using raw page images with multimodal VLMs, sharing the same hierarchical tree index while differing only in content extraction and generation components.**

PageIndex by VectifyAI is an open-source framework for building retrieval-augmented generation (RAG) systems over complex PDF documents. Understanding the differences between text-based and vision-based RAG implementations in PageIndex is essential for selecting the right pipeline—whether you need fast token-efficient retrieval on text-heavy documents or comprehensive visual understanding of scanned PDFs with complex layouts.

## Shared Core Architecture

Both implementations rely on PageIndex’s hierarchical tree structure built from the document’s table of contents. In [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), the `meta_processor` function orchestrates TOC detection, tree construction, and page offset handling, producing a JSON tree where each node carries a **physical_index** (page number) linking it to the source material.

The key distinction lies in what content gets associated with these nodes: text-based pipelines store extracted strings, while vision-based pipelines reference image files, yet both use identical tree-reasoning logic for retrieval.

## Text-Based RAG Pipeline

### Input Processing with OCR-Free Extraction

Text-based RAG begins with extracting selectable text directly from PDF pages without optical character recognition. The `utils.get_page_tokens` function in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) handles this extraction using libraries like PyPDF2 or PyMuPDF, then counts tokens using `tiktoken` to manage context windows.

### Tree Construction with Text Nodes

When building the index via `page_index()` in [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), setting `if_add_node_text="yes"` populates each tree node with the extracted text content. The **physical_index** remains the critical link between the hierarchical structure and the actual page numbers.

### LLM Retrieval and Generation

Retrieval uses a **text-only LLM** such as `gpt-4o-2024-11-20`. The model receives the tree structure (titles only) and selects relevant node IDs. For generation, the system retrieves the text snippets associated with those nodes and prompts the LLM with the question plus the concatenated text content.

```python
from pageindex import page_index

# Build the tree and get a full structure with text nodes

result = page_index(
    doc="my_report.pdf",               # PDF path (or BytesIO)

    model="gpt-4o-2024-11-20",
    if_add_node_text="yes",            # keep extracted text for each node

)

print(result["structure"])   # hierarchical JSON with `text` fields

```

## Vision-Based RAG Pipeline

### Image Extraction with PyMuPDF

Vision-based RAG bypasses text extraction entirely. Instead, the `extract_pdf_page_images` helper function in `cookbook/vision_RAG_pageindex.ipynb` uses `fitz` (PyMuPDF) to render high-resolution JPEGs of each page at 2x zoom for clarity.

```python
def extract_images(pdf_path, out_dir="pdf_images"):
    os.makedirs(out_dir, exist_ok=True)
    doc = fitz.open(pdf_path)
    images = {}
    for i, page in enumerate(doc):
        pix = page.get_pixmap(matrix=fitz.Matrix(2, 2))
        img_path = f"{out_dir}/page_{i+1}.jpg"
        pix.save(img_path)
        images[i+1] = img_path
    return images

```

### Tree Construction Without Text Content

The tree is built using the same [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) logic, but nodes contain only titles and **physical_index** references—no extracted text. The content remains in the image files, referenced by page number.

### Multimodal Retrieval and Generation

Retrieval still uses a text-only LLM to select node IDs from the tree structure. However, generation employs a **multimodal VLM** such as `gpt-4.1` or `Qwen-VL`. The `get_page_images_for_nodes` function maps the selected node IDs to their corresponding JPEG files using the **physical_index**, and the VLM receives both the question and the actual page images.

```python

# Retrieve candidate nodes (text-only LLM)

search_prompt = f"""
You are given a question and a tree of section titles.
Find node IDs that likely contain the answer.

Question: {query}
Tree: {json.dumps(utils.remove_fields(tree, ["text"]))}
"""
candidate_ids = await call_vlm(search_prompt)

# Gather image paths for those nodes

retrieved_imgs = utils.get_page_images_for_nodes(
    candidate_ids, 
    utils.create_node_mapping(tree), 
    page_images
)

# Generate answer with multimodal VLM

answer_prompt = f"Answer the question using ONLY the supplied images.\n\nQuestion: {query}"
answer = await call_vlm(answer_prompt, image_paths=retrieved_imgs)

```

## Critical Differences Between Text and Vision Pipelines

Understanding the differences between text-based and vision-based RAG implementations in PageIndex requires examining four key architectural dimensions:

**Input Processing**
- **Text-based**: Uses `utils.get_page_tokens` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) to extract selectable text and count tokens with `tiktoken`.
- **Vision-based**: Uses `fitz` (PyMuPDF) in `cookbook/vision_RAG_pageindex.ipynb` to render pages as high-resolution JPEGs without OCR.

**Tokenization Strategy**
- **Text-based**: Explicit token counting via `get_page_tokens` to manage LLM context windows and prevent overflow.
- **Vision-based**: No token counting during retrieval; the VLM’s internal vision encoder processes image content directly.

**Generation Model Requirements**
- **Text-based**: Compatible with text-only LLMs like `gpt-4o-2024-11-20`.
- **Vision-based**: Requires multimodal VLMs like `gpt-4.1` or `Qwen-VL` capable of processing image inputs.

**Content Fidelity**
- **Text-based**: Dependent on embedded PDF text quality; loses visual layout, figures, and complex table structures.
- **Vision-based**: Preserves complete visual context including charts, diagrams, and formatting, but requires more compute resources.

## Summary

- PageIndex unifies both RAG approaches under a single hierarchical tree structure built by [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py), using **physical_index** to link nodes to page numbers.
- **Text-based RAG** extracts strings via `utils.get_page_tokens`, counts tokens with `tiktoken`, and uses standard LLMs for generation.
- **Vision-based RAG** extracts images via PyMuPDF, skips token counting, and uses multimodal VLMs like `gpt-4.1` to interpret visual content.
- Both implementations share identical tree-reasoning retrieval logic but differ in content extraction and the final generation component.

## Frequently Asked Questions

### Can I switch between text and vision RAG without rebuilding the tree index?

No, while the hierarchical tree structure built by [`pageindex/page_index.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/page_index.py) is shared between both approaches, the content associated with each node differs fundamentally. Text-based RAG requires `if_add_node_text="yes"` to populate nodes with extracted strings, while vision-based RAG relies on external image files mapped via **physical_index**. You must reprocess the document with the appropriate content extraction method when switching pipelines.

### Which pipeline performs better on scanned documents?

Vision-based RAG significantly outperforms text-based approaches on scanned documents or PDFs with complex layouts. Since the vision pipeline in `cookbook/vision_RAG_pageindex.ipynb` processes raw page images without OCR, it preserves visual elements like charts, tables, and diagrams that text extraction often corrupts or omits. However, this comes at higher computational cost compared to the token-efficient text pipeline.

### How does token counting differ between text and vision RAG?

In text-based RAG, `utils.get_page_tokens` in [`pageindex/utils.py`](https://github.com/VectifyAI/PageIndex/blob/main/pageindex/utils.py) explicitly counts tokens using `tiktoken` to manage context windows and prevent LLM overflow. Vision-based RAG eliminates explicit token counting during retrieval; instead, the multimodal VLM processes images through its internal vision encoder, handling content size constraints automatically. This architectural difference makes vision RAG simpler to implement but potentially more expensive at inference time.

### What multimodal models are compatible with PageIndex vision RAG?

PageIndex’s vision-based RAG implementation in `cookbook/vision_RAG_pageindex.ipynb` supports any multimodal VLM capable of processing image inputs alongside text prompts. The reference implementation uses `gpt-4.1` and `Qwen-VL`, but the pattern works with other vision-language models like `gpt-4o-vision` or `Claude 3 Opus`. The critical requirement is that the model must accept base64-encoded images or image URLs in its API payload, as demonstrated in the notebook’s `call_vlm` function.