How to Implement Multi-Modal RAG: Two Production-Ready Architectures Explained

Multi-modal RAG combines text and visual data by either converting images to captions for text-based retrieval or using vision-language models like ColPali to embed pages directly, enabling rich document understanding beyond pure text.

The NirDiamant/RAG_Techniques repository provides complete, runnable implementations of both approaches. Whether you need to query charts in financial reports or diagrams in research papers, these architectures demonstrate how to index and retrieve multi-modal content using modern embedding models and LLMs.

Architecture 1: Caption-Based Multi-Modal RAG

The caption-based approach treats images as text by generating descriptive summaries. This method works with any standard vector database and text-based retriever, making it compatible with existing RAG infrastructure.

Extracting Text and Images from PDFs

The pipeline begins with PyMuPDF (fitz) to parse PDF documents. In all_rag_techniques/multi_model_rag_with_captioning.ipynb, the extraction logic iterates through pages, capturing raw text via page.get_text() and binary image data via page.get_images(full=True).

import fitz
from PIL import Image
import io

text_data, image_data = [], []

with fitz.open("attention_is_all_you_need.pdf") as pdf:
    for page_num, page in enumerate(pdf):
        # Extract text

        text_data.append({
            "response": page.get_text().strip(),
            "name": page_num + 1
        })
        
        # Extract images

        for img_idx, img in enumerate(page.get_images(full=True)):
            base_image = pdf.extract_image(img[0])["image"]
            image_data.append({
                "image_bytes": base_image,
                "page": page_num + 1,
                "index": img_idx
            })

Generating Captions with Gemini

Each extracted image is sent to Gemini 1.5-flash via the google-generativeai SDK. The model generates concise captions optimized for retrieval, effectively translating visual information into searchable text.

import google.generativeai as genai

genai.configure(api_key=os.getenv("GOOGLE_API_KEY"))
model = genai.GenerativeModel("gemini-1.5-flash")

# Caption generation

caption = model.generate_content([
    Image.open(io.BytesIO(image_bytes)),
    "Summarise this image for retrieval."
]).text

Embedding and Indexing with Cohere and Chroma

Both raw text chunks and generated captions are processed through Cohere Embeddings (embed-english-v3.0) and stored in Chroma. The RecursiveCharacterTextSplitter creates 400-token chunks with 50-token overlap to maintain context.

from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_cohere import CohereEmbeddings
from langchain_community.vectorstores import Chroma
from langchain_core.documents import Document

# Prepare documents

docs = [Document(page_content=d["response"], metadata={"name": d["name"]}) 
        for d in text_data + img_captions]

# Chunking

splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
    chunk_size=400, 
    chunk_overlap=50
)
splits = splitter.split_documents(docs)

# Embedding and storage

embedding = CohereEmbeddings(model="embed-english-v3.0")
vectorstore = Chroma.from_documents(
    splits, 
    collection_name="multi_model_rag", 
    embedding=embedding
)

Retrieval and Answer Generation

The retrieval phase uses vectorstore.as_retriever(search_kwargs={'k': 1}) to fetch the most relevant chunk. A Cohere Chat model (command-r-plus) synthesizes the final answer from the retrieved context.

from langchain_cohere import ChatCohere
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser

retriever = vectorstore.as_retriever(search_kwargs={"k": 1})
retrieved_docs = retriever.invoke("What is the BLEU score of the Transformer (base model)?")

prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a concise QA assistant."),
    ("human", "Retrieved: <docs>{documents}</docs>\nQuestion: <question>{question}</question>")
])

rag_chain = prompt | ChatCohere(model="command-r-plus") | StrOutputParser()
answer = rag_chain.invoke({
    "documents": retrieved_docs[0].page_content, 
    "question": "What is the BLEU score of the Transformer (base model)?"
})

Architecture 2: ColPali-Based Multi-Modal RAG

The ColPali approach eliminates caption generation by using a vision-language model that jointly embeds text and image pixels into a single vector space. This method, implemented in all_rag_techniques/multi_model_rag_with_colpali.ipynb, retrieves entire PDF pages as images and answers questions based on visual evidence.

Understanding ColPali's Joint Embedding

ColPali-v1.2 (accessed via the byaldi library) processes PDF pages as unified multimodal objects. Unlike the caption-based method, it preserves the spatial relationship between text and visual elements, making it superior for charts, tables, and diagrams.

Indexing PDFs with RAGMultiModalModel

The RAGMultiModalModel.from_pretrained() method loads the ColPali weights, while RAG.index() processes the PDF into a searchable collection stored under .byaldi/.

from byaldi import RAGMultiModalModel

# Load pretrained ColPali model

RAG = RAGMultiModalModel.from_pretrained("vidore/colpali-v1.2", verbose=1)

# Index PDF - creates .byaldi/ directory with embeddings

RAG.index(
    input_path="./docs/attention_is_all_you_need.pdf",
    index_name="attention_is_all_you_need",
    store_collection_with_index=True,
    overwrite=True,
)

Searching and Decoding Visual Results

Queries return page-level results containing Base64-encoded images. The results[0].base64 attribute contains the visual data, which is decoded using Python's base64 module.

import base64
from PIL import Image
import io

# Search query

query = "What is the BLEU score of the Transformer (base model)?"
results = RAG.search(query, k=1)

# Decode retrieved image

image_bytes = base64.b64decode(results[0].base64)
image = Image.open(io.BytesIO(image_bytes))

# Save or display

image.save("retrieved_page.jpg")

Generating Answers from Retrieved Images

The decoded image is passed to Gemini 1.5-flash alongside the original query. The LLM performs visual question answering on the retrieved page content.

import google.generativeai as genai
import os

genai.configure(api_key=os.getenv("GOOGLE_API_KEY"))
gemini = genai.GenerativeModel("gemini-1.5-flash")

# Generate answer from image + query

response = gemini.generate_content([image, query])
print(response.text)

Complete Code Examples

Caption-Based Implementation

!pip install langchain langchain-community pillow pymupdf google-generativeai cohere

import fitz, io, os
from PIL import Image
import google.generativeai as genai
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import Chroma
from langchain_cohere import CohereEmbeddings, ChatCohere
from langchain_core.documents import Document
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser

# Configuration

genai.configure(api_key=os.getenv("GOOGLE_API_KEY"))
pdf_path = "attention_is_all_you_need.pdf"

# 1️⃣ Extract text and images

text_data, img_captions = [], []
with fitz.open(pdf_path) as pdf:
    for page_num, page in enumerate(pdf):
        text_data.append({
            "response": page.get_text().strip(),
            "name": page_num + 1
        })
        
        model = genai.GenerativeModel("gemini-1.5-flash")
        for img_idx, img in enumerate(page.get_images(full=True)):
            base_image = pdf.extract_image(img[0])["image"]
            caption = model.generate_content([
                Image.open(io.BytesIO(base_image)),
                "Summarise this image for retrieval."
            ]).text
            img_captions.append({
                "response": caption,
                "name": f"p{page_num+1}_i{img_idx}"
            })

# 2️⃣ Chunk and index

all_docs = [Document(page_content=d["response"], metadata={"name": d["name"]}) 
            for d in text_data + img_captions]

splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
    chunk_size=400, 
    chunk_overlap=50
)
splits = splitter.split_documents(all_docs)

embedding = CohereEmbeddings(model="embed-english-v3.0")
vectorstore = Chroma.from_documents(
    splits, 
    collection_name="multi_model_rag", 
    embedding=embedding
)

# 3️⃣ Retrieve and generate

retriever = vectorstore.as_retriever(search_kwargs={"k": 1})
query = "What is the BLEU score of the Transformer (base model)?"
retrieved = retriever.invoke(query)

prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a concise QA assistant."),
    ("human", "Retrieved: <docs>{documents}</docs>\nQuestion: <question>{question}</question>")
])

rag_chain = prompt | ChatCohere(model="command-r-plus") | StrOutputParser()
answer = rag_chain.invoke({
    "documents": retrieved[0].page_content, 
    "question": query
})
print(answer)

ColPali-Based Implementation

!pip install byaldi pillow google-generativeai

import os, base64, io
from byaldi import RAGMultiModalModel
from PIL import Image
import google.generativeai as genai

# Configure Gemini

genai.configure(api_key=os.getenv("GOOGLE_API_KEY"))
gemini = genai.GenerativeModel("gemini-1.5-flash")

# Initialise ColPali model

rag = RAGMultiModalModel.from_pretrained("vidore/colpali-v1.2", verbose=1)

# Index PDF (run once per document)

rag.index(
    input_path="attention_is_all_you_need.pdf",
    index_name="attention_is_all_you_need",
    store_collection_with_index=True,
    overwrite=True,
)

# Search

query = "What is the BLEU score of the Transformer (base model)?"
results = rag.search(query, k=1)

# Decode retrieved image

image_bytes = base64.b64decode(results[0].base64)
image_path = "retrieved_page.jpg"
with open(image_path, "wb") as f:
    f.write(image_bytes)

# Generate answer from image

image = Image.open(image_path)
response = gemini.generate_content([image, query])
print(response.text)

Key Files in the Repository

The NirDiamant/RAG_Techniques repository contains the following critical resources for implementing multi-modal RAG:

  • all_rag_techniques/multi_model_rag_with_captioning.ipynb – Complete implementation of the caption-based pipeline using PyMuPDF, Gemini, and Cohere embeddings.

  • all_rag_techniques/multi_model_rag_with_colpali.ipynb – End-to-end ColPali implementation using the byaldi library for joint text-image embedding.

  • helper_functions.py – Shared utilities for environment configuration and document processing used across both notebooks.

  • README.md – Comprehensive index of all RAG techniques available in the repository, including multimodal approaches.

Summary

  • Multi-modal RAG enables retrieval from documents containing charts, diagrams, and images, not just text.

  • The caption-based approach (multi_model_rag_with_captioning.ipynb) uses Gemini 1.5-flash to describe images, then indexes these captions alongside text using Cohere embeddings in Chroma.

  • The ColPali approach (multi_model_rag_with_colpali.ipynb) uses the byaldi library to embed entire PDF pages as multimodal vectors, retrieving Base64-encoded images that are decoded and analyzed by Gemini.

  • Both methods support production deployment, with caption-based offering broader compatibility with existing text-based vector stores, while ColPali preserves richer visual-textual relationships.

Frequently Asked Questions

What is the difference between caption-based and ColPali multi-modal RAG?

Caption-based multi-modal RAG converts images to text descriptions using vision-language models like Gemini, then indexes these captions as standard text chunks. ColPali-based RAG uses a specialized multimodal encoder to create joint embeddings of text and image pixels, retrieving entire pages as images without intermediate caption generation. Caption-based works with any text vector store, while ColPali requires compatible multimodal indexing infrastructure.

Which multi-modal RAG approach is better for tables and charts?

ColPali generally performs better for tables and charts because it preserves the spatial relationships between text and visual elements in its joint embedding space. The caption-based approach may lose fine-grained structural details when converting complex tables to text descriptions. However, if your infrastructure only supports text embeddings, caption-based RAG with detailed prompting for table structure extraction remains a viable alternative.

What are the hardware requirements for running ColPali-based RAG?

ColPali requires sufficient GPU memory to load the vidore/colpali-v1.2 vision-language model for embedding generation. While the byaldi library handles efficient indexing, initial PDF processing involves encoding full-resolution page images. For production deployment, plan for GPU resources similar to those required for other vision transformers (typically 8GB+ VRAM), though CPU-only inference is possible with significant latency trade-offs.

Can I use other LLMs besides Gemini and Cohere for multi-modal RAG?

Yes, both architectures support model substitution. For caption-based RAG, you can replace Gemini 1.5-flash with OpenAI's GPT-4V, Anthropic's Claude 3, or local multimodal models like LLaVA. Similarly, Cohere embeddings can be swapped for OpenAI's text-embedding-3-large or open-source alternatives like BGE or E5. For ColPali-based RAG, while the embedding model is fixed (ColPali), the final answer generation can use any LLM capable of processing images, including local models via Ollama or vLLM.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →