# How to Implement Multi-Modal RAG: Two Production-Ready Architectures Explained

> Learn how to implement multi-modal RAG by exploring two production-ready architectures. Understand combining text and visual data for enhanced document comprehension and retrieval.

- Repository: [NirDiamant/RAG_Techniques](https://github.com/nirdiamant/rag_techniques)
- Tags: architecture
- Published: 2026-02-19

---

**Multi-modal RAG combines text and visual data by either converting images to captions for text-based retrieval or using vision-language models like ColPali to embed pages directly, enabling rich document understanding beyond pure text.**

The `NirDiamant/RAG_Techniques` repository provides complete, runnable implementations of both approaches. Whether you need to query charts in financial reports or diagrams in research papers, these architectures demonstrate how to index and retrieve multi-modal content using modern embedding models and LLMs.

## Architecture 1: Caption-Based Multi-Modal RAG

The caption-based approach treats images as text by generating descriptive summaries. This method works with any standard vector database and text-based retriever, making it compatible with existing RAG infrastructure.

### Extracting Text and Images from PDFs

The pipeline begins with `PyMuPDF` (`fitz`) to parse PDF documents. In `all_rag_techniques/multi_model_rag_with_captioning.ipynb`, the extraction logic iterates through pages, capturing raw text via `page.get_text()` and binary image data via `page.get_images(full=True)`.

```python
import fitz
from PIL import Image
import io

text_data, image_data = [], []

with fitz.open("attention_is_all_you_need.pdf") as pdf:
    for page_num, page in enumerate(pdf):
        # Extract text

        text_data.append({
            "response": page.get_text().strip(),
            "name": page_num + 1
        })
        
        # Extract images

        for img_idx, img in enumerate(page.get_images(full=True)):
            base_image = pdf.extract_image(img[0])["image"]
            image_data.append({
                "image_bytes": base_image,
                "page": page_num + 1,
                "index": img_idx
            })

```

### Generating Captions with Gemini

Each extracted image is sent to **Gemini 1.5-flash** via the `google-generativeai` SDK. The model generates concise captions optimized for retrieval, effectively translating visual information into searchable text.

```python
import google.generativeai as genai

genai.configure(api_key=os.getenv("GOOGLE_API_KEY"))
model = genai.GenerativeModel("gemini-1.5-flash")

# Caption generation

caption = model.generate_content([
    Image.open(io.BytesIO(image_bytes)),
    "Summarise this image for retrieval."
]).text

```

### Embedding and Indexing with Cohere and Chroma

Both raw text chunks and generated captions are processed through **Cohere Embeddings** (`embed-english-v3.0`) and stored in **Chroma**. The `RecursiveCharacterTextSplitter` creates 400-token chunks with 50-token overlap to maintain context.

```python
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_cohere import CohereEmbeddings
from langchain_community.vectorstores import Chroma
from langchain_core.documents import Document

# Prepare documents

docs = [Document(page_content=d["response"], metadata={"name": d["name"]}) 
        for d in text_data + img_captions]

# Chunking

splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
    chunk_size=400, 
    chunk_overlap=50
)
splits = splitter.split_documents(docs)

# Embedding and storage

embedding = CohereEmbeddings(model="embed-english-v3.0")
vectorstore = Chroma.from_documents(
    splits, 
    collection_name="multi_model_rag", 
    embedding=embedding
)

```

### Retrieval and Answer Generation

The retrieval phase uses `vectorstore.as_retriever(search_kwargs={'k': 1})` to fetch the most relevant chunk. A **Cohere Chat** model (`command-r-plus`) synthesizes the final answer from the retrieved context.

```python
from langchain_cohere import ChatCohere
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser

retriever = vectorstore.as_retriever(search_kwargs={"k": 1})
retrieved_docs = retriever.invoke("What is the BLEU score of the Transformer (base model)?")

prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a concise QA assistant."),
    ("human", "Retrieved: <docs>{documents}</docs>\nQuestion: <question>{question}</question>")
])

rag_chain = prompt | ChatCohere(model="command-r-plus") | StrOutputParser()
answer = rag_chain.invoke({
    "documents": retrieved_docs[0].page_content, 
    "question": "What is the BLEU score of the Transformer (base model)?"
})

```

## Architecture 2: ColPali-Based Multi-Modal RAG

The **ColPali** approach eliminates caption generation by using a vision-language model that jointly embeds text and image pixels into a single vector space. This method, implemented in `all_rag_techniques/multi_model_rag_with_colpali.ipynb`, retrieves entire PDF pages as images and answers questions based on visual evidence.

### Understanding ColPali's Joint Embedding

**ColPali-v1.2** (accessed via the `byaldi` library) processes PDF pages as unified multimodal objects. Unlike the caption-based method, it preserves the spatial relationship between text and visual elements, making it superior for charts, tables, and diagrams.

### Indexing PDFs with RAGMultiModalModel

The `RAGMultiModalModel.from_pretrained()` method loads the ColPali weights, while `RAG.index()` processes the PDF into a searchable collection stored under `.byaldi/`.

```python
from byaldi import RAGMultiModalModel

# Load pretrained ColPali model

RAG = RAGMultiModalModel.from_pretrained("vidore/colpali-v1.2", verbose=1)

# Index PDF - creates .byaldi/ directory with embeddings

RAG.index(
    input_path="./docs/attention_is_all_you_need.pdf",
    index_name="attention_is_all_you_need",
    store_collection_with_index=True,
    overwrite=True,
)

```

### Searching and Decoding Visual Results

Queries return page-level results containing **Base64-encoded images**. The `results[0].base64` attribute contains the visual data, which is decoded using Python's `base64` module.

```python
import base64
from PIL import Image
import io

# Search query

query = "What is the BLEU score of the Transformer (base model)?"
results = RAG.search(query, k=1)

# Decode retrieved image

image_bytes = base64.b64decode(results[0].base64)
image = Image.open(io.BytesIO(image_bytes))

# Save or display

image.save("retrieved_page.jpg")

```

### Generating Answers from Retrieved Images

The decoded image is passed to **Gemini 1.5-flash** alongside the original query. The LLM performs visual question answering on the retrieved page content.

```python
import google.generativeai as genai
import os

genai.configure(api_key=os.getenv("GOOGLE_API_KEY"))
gemini = genai.GenerativeModel("gemini-1.5-flash")

# Generate answer from image + query

response = gemini.generate_content([image, query])
print(response.text)

```

## Complete Code Examples

### Caption-Based Implementation

```python
!pip install langchain langchain-community pillow pymupdf google-generativeai cohere

import fitz, io, os
from PIL import Image
import google.generativeai as genai
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import Chroma
from langchain_cohere import CohereEmbeddings, ChatCohere
from langchain_core.documents import Document
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser

# Configuration

genai.configure(api_key=os.getenv("GOOGLE_API_KEY"))
pdf_path = "attention_is_all_you_need.pdf"

# 1️⃣ Extract text and images

text_data, img_captions = [], []
with fitz.open(pdf_path) as pdf:
    for page_num, page in enumerate(pdf):
        text_data.append({
            "response": page.get_text().strip(),
            "name": page_num + 1
        })
        
        model = genai.GenerativeModel("gemini-1.5-flash")
        for img_idx, img in enumerate(page.get_images(full=True)):
            base_image = pdf.extract_image(img[0])["image"]
            caption = model.generate_content([
                Image.open(io.BytesIO(base_image)),
                "Summarise this image for retrieval."
            ]).text
            img_captions.append({
                "response": caption,
                "name": f"p{page_num+1}_i{img_idx}"
            })

# 2️⃣ Chunk and index

all_docs = [Document(page_content=d["response"], metadata={"name": d["name"]}) 
            for d in text_data + img_captions]

splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
    chunk_size=400, 
    chunk_overlap=50
)
splits = splitter.split_documents(all_docs)

embedding = CohereEmbeddings(model="embed-english-v3.0")
vectorstore = Chroma.from_documents(
    splits, 
    collection_name="multi_model_rag", 
    embedding=embedding
)

# 3️⃣ Retrieve and generate

retriever = vectorstore.as_retriever(search_kwargs={"k": 1})
query = "What is the BLEU score of the Transformer (base model)?"
retrieved = retriever.invoke(query)

prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a concise QA assistant."),
    ("human", "Retrieved: <docs>{documents}</docs>\nQuestion: <question>{question}</question>")
])

rag_chain = prompt | ChatCohere(model="command-r-plus") | StrOutputParser()
answer = rag_chain.invoke({
    "documents": retrieved[0].page_content, 
    "question": query
})
print(answer)

```

### ColPali-Based Implementation

```python
!pip install byaldi pillow google-generativeai

import os, base64, io
from byaldi import RAGMultiModalModel
from PIL import Image
import google.generativeai as genai

# Configure Gemini

genai.configure(api_key=os.getenv("GOOGLE_API_KEY"))
gemini = genai.GenerativeModel("gemini-1.5-flash")

# Initialise ColPali model

rag = RAGMultiModalModel.from_pretrained("vidore/colpali-v1.2", verbose=1)

# Index PDF (run once per document)

rag.index(
    input_path="attention_is_all_you_need.pdf",
    index_name="attention_is_all_you_need",
    store_collection_with_index=True,
    overwrite=True,
)

# Search

query = "What is the BLEU score of the Transformer (base model)?"
results = rag.search(query, k=1)

# Decode retrieved image

image_bytes = base64.b64decode(results[0].base64)
image_path = "retrieved_page.jpg"
with open(image_path, "wb") as f:
    f.write(image_bytes)

# Generate answer from image

image = Image.open(image_path)
response = gemini.generate_content([image, query])
print(response.text)

```

## Key Files in the Repository

The `NirDiamant/RAG_Techniques` repository contains the following critical resources for implementing multi-modal RAG:

- **`all_rag_techniques/multi_model_rag_with_captioning.ipynb`** – Complete implementation of the caption-based pipeline using PyMuPDF, Gemini, and Cohere embeddings.

- **`all_rag_techniques/multi_model_rag_with_colpali.ipynb`** – End-to-end ColPali implementation using the `byaldi` library for joint text-image embedding.

- **[`helper_functions.py`](https://github.com/NirDiamant/RAG_Techniques/blob/main/helper_functions.py)** – Shared utilities for environment configuration and document processing used across both notebooks.

- **[`README.md`](https://github.com/NirDiamant/RAG_Techniques/blob/main/README.md)** – Comprehensive index of all RAG techniques available in the repository, including multimodal approaches.

## Summary

- **Multi-modal RAG** enables retrieval from documents containing charts, diagrams, and images, not just text.

- The **caption-based approach** (`multi_model_rag_with_captioning.ipynb`) uses Gemini 1.5-flash to describe images, then indexes these captions alongside text using Cohere embeddings in Chroma.

- The **ColPali approach** (`multi_model_rag_with_colpali.ipynb`) uses the `byaldi` library to embed entire PDF pages as multimodal vectors, retrieving Base64-encoded images that are decoded and analyzed by Gemini.

- Both methods support production deployment, with caption-based offering broader compatibility with existing text-based vector stores, while ColPali preserves richer visual-textual relationships.

## Frequently Asked Questions

### What is the difference between caption-based and ColPali multi-modal RAG?

Caption-based multi-modal RAG converts images to text descriptions using vision-language models like Gemini, then indexes these captions as standard text chunks. ColPali-based RAG uses a specialized multimodal encoder to create joint embeddings of text and image pixels, retrieving entire pages as images without intermediate caption generation. Caption-based works with any text vector store, while ColPali requires compatible multimodal indexing infrastructure.

### Which multi-modal RAG approach is better for tables and charts?

ColPali generally performs better for tables and charts because it preserves the spatial relationships between text and visual elements in its joint embedding space. The caption-based approach may lose fine-grained structural details when converting complex tables to text descriptions. However, if your infrastructure only supports text embeddings, caption-based RAG with detailed prompting for table structure extraction remains a viable alternative.

### What are the hardware requirements for running ColPali-based RAG?

ColPali requires sufficient GPU memory to load the `vidore/colpali-v1.2` vision-language model for embedding generation. While the `byaldi` library handles efficient indexing, initial PDF processing involves encoding full-resolution page images. For production deployment, plan for GPU resources similar to those required for other vision transformers (typically 8GB+ VRAM), though CPU-only inference is possible with significant latency trade-offs.

### Can I use other LLMs besides Gemini and Cohere for multi-modal RAG?

Yes, both architectures support model substitution. For caption-based RAG, you can replace Gemini 1.5-flash with OpenAI's GPT-4V, Anthropic's Claude 3, or local multimodal models like LLaVA. Similarly, Cohere embeddings can be swapped for OpenAI's `text-embedding-3-large` or open-source alternatives like BGE or E5. For ColPali-based RAG, while the embedding model is fixed (ColPali), the final answer generation can use any LLM capable of processing images, including local models via Ollama or vLLM.