How to Use Chroma for RAG in Python: A Complete Implementation Guide

Chroma is a fast, persistent vector store that integrates with LangChain to power Retrieval-Augmented Generation (RAG) pipelines by storing document embeddings and retrieving relevant context for LLM queries.

The Shubhamsaboo/awesome-llm-apps repository provides production-ready examples demonstrating how to implement Chroma-backed RAG systems in Python. Chroma enables local, persistent storage of vector embeddings, making it ideal for applications that require fast similarity search without external database dependencies.

What Is Chroma and Why Use It for RAG?

Chroma is an open-source embedding database designed specifically for AI applications. In a RAG pipeline, Chroma serves as the vector store layer that persists document chunks as dense vectors. When a user submits a query, Chroma performs similarity search to retrieve the most relevant context, which is then fed into an LLM to generate grounded answers.

Key advantages of using Chroma for RAG include:

  • Persistent storage via persist_directory for data survival across sessions
  • Collection-based organization for managing multiple document sets
  • Native LangChain integration through both langchain-chroma and langchain-community packages
  • Local execution without requiring cloud vector database services

Project Structure and Key Files

The repository contains several reference implementations that demonstrate Chroma integration patterns:

File Implementation Pattern
rag_tutorials/rag_chain/app.py Full Streamlit application using langchain-chroma with Google Gemini embeddings and PDF ingestion
rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py Local-only pipeline using Ollama embeddings, web page loading, and Chroma persistence
advanced_llm_apps/chat_with_X_tutorials/chat_with_github/chat_github_llama3.py Embedchain abstraction layer configuring Chroma as the vector database backend

Step-by-Step Implementation

1. Install Dependencies

Install the required packages for Chroma and LangChain integration:

pip install langchain langchain-chroma langchain-community pypdf streamlit

For Google Gemini support, add:

pip install langchain-google-genai

For local Ollama embeddings:

pip install langchain-ollama

2. Initialize the Embedding Model

Choose an embedding model based on your deployment requirements.

Google Gemini embeddings (as implemented in rag_tutorials/rag_chain/app.py):

from langchain_google_genai import GoogleGenerativeAIEmbeddings

embedding_model = GoogleGenerativeAIEmbeddings(model="models/embedding-001")

Ollama local embeddings (as implemented in llama3.1_local_rag.py):

from langchain_ollama import OllamaEmbeddings

embeddings = OllamaEmbeddings(
    model="llama3.1",
    base_url="http://127.0.0.1:11434"
)

3. Create a Persistent Chroma Vector Store

Initialize Chroma with a collection name and persistence directory.

Using langchain-chroma (modern approach):

from langchain_chroma import Chroma

db = Chroma(
    collection_name="my_collection",
    embedding_function=embedding_model,
    persist_directory="./my_chroma_db"
)

Using langchain-community (alternative approach):

from langchain_community.vectorstores import Chroma

vectorstore = Chroma.from_documents(
    documents=doc_chunks,
    embedding=embeddings,
    persist_directory="./my_chroma_db"
)

The persist_directory parameter ensures vector data survives between Python sessions. The collection_name parameter organizes related documents into logical groups.

4. Load and Split Documents

Process source documents into chunks suitable for embedding.

PDF processing (from rag_chain/app.py):

from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters.sentence_transformers import SentenceTransformersTokenTextSplitter

loader = PyPDFLoader("my_report.pdf")
raw_docs = loader.load()

splitter = SentenceTransformersTokenTextSplitter(
    model_name="sentence-transformers/all-mpnet-base-v2",
    chunk_size=100,
    chunk_overlap=50
)

doc_chunks = splitter.create_documents(
    [d.page_content for d in raw_docs],
    [d.metadata for d in raw_docs]
)

Web page processing (from llama3.1_local_rag.py):

from langchain_community.document_loaders import WebBaseLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter

loader = WebBaseLoader("https://example.com/article")
docs = loader.load()

splitter = RecursiveCharacterTextSplitter(
    chunk_size=500,
    chunk_overlap=10
)
doc_chunks = splitter.split_documents(docs)

5. Add Documents to Chroma

Persist the processed chunks to the vector store:

db.add_documents(doc_chunks)

For the community wrapper:

vectorstore.add_documents(doc_chunks)

6. Build the Retriever

Configure similarity search to fetch relevant context:

retriever = db.as_retriever(
    search_type="similarity",
    search_kwargs={"k": 5}
)

The k parameter controls the number of retrieved documents. Adjust based on your context window requirements.

7. Assemble the RAG Chain

Combine retrieval, formatting, and LLM generation using LangChain's pipe syntax:

from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough

def format_docs(docs):
    return "\n\n".join(doc.page_content for doc in docs)

prompt_template = ChatPromptTemplate.from_template(
    """
    You are an AI assistant. Answer the question using ONLY the following context:

    {context}

    Question: {question}
    """
)

chat_model = ChatGoogleGenerativeAI(
    model="gemini-1.5-pro",
    temperature=1,
)

rag_chain = {
    "context": retriever | format_docs,
    "question": RunnablePassthrough()
} | prompt_template | chat_model | StrOutputParser()

8. Invoke the Pipeline

Execute the RAG chain with a user query:

response = rag_chain.invoke("What are the AI applications in drug discovery?")
print(response)

Complete Minimal Example

Here is a condensed, runnable implementation combining all steps:

import streamlit as st
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from langchain_chroma import Chroma
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters.sentence_transformers import SentenceTransformersTokenTextSplitter
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_core.output_parsers import StrOutputParser
from langchain_google_genai import ChatGoogleGenerativeAI

# 1️⃣ Embeddings

embeddings = GoogleGenerativeAIEmbeddings(model="models/embedding-001")

# 2️⃣ Vector store

db = Chroma(
    collection_name="demo",
    embedding_function=embeddings,
    persist_directory="./demo_db"
)

# 3️⃣ Load & split a PDF

loader = PyPDFLoader("sample.pdf")
raw = loader.load()
splitter = SentenceTransformersTokenTextSplitter(
    model_name="sentence-transformers/all-mpnet-base-v2",
    chunk_size=120,
    chunk_overlap=30
)
chunks = splitter.create_documents(
    [d.page_content for d in raw],
    [d.metadata for d in raw]
)

# 4️⃣ Add to DB (first run only)

if not db._client.get_collection("demo"):
    db.add_documents(chunks)

# 5️⃣ Retriever

retriever = db.as_retriever(search_kwargs={"k": 4})

# 6️⃣ Prompt + LLM

prompt = ChatPromptTemplate.from_template(
    """Answer using only the context below.

    {context}

    Question: {question}"""
)
llm = ChatGoogleGenerativeAI(
    model="gemini-1.5-flash",
    api_key=st.secrets["GEMINI_API_KEY"],
    temperature=0.7
)

rag = {
    "context": retriever | (lambda docs: "\n\n".join(d.page_content for d in docs)),
    "question": RunnablePassthrough()
} | prompt | llm | StrOutputParser()

# 7️⃣ UI

st.title("📚 Chroma‑Powered RAG Demo")
query = st.text_input("Ask a question")
if query:
    st.write(rag.invoke(query))

Alternative: Using Embedchain with Chroma

For higher-level abstraction, the repository also demonstrates using Embedchain to handle Chroma configuration automatically. In advanced_llm_apps/chat_with_X_tutorials/chat_with_github/chat_github_llama3.py, Chroma is configured as the vector database backend:

from embedchain import App

app = App.from_config({
    "vectordb": {
        "provider": "chroma",
        "config": {
            "collection_name": "github_repo",
            "dir": "./chroma_db"
        }
    },
    "embedder": {
        "provider": "ollama",
        "config": {"model": "llama3.1"}
    }
})

This approach abstracts the embedding and storage logic while still leveraging Chroma's persistent storage capabilities.

Summary

  • Chroma provides persistent, local vector storage for RAG pipelines through the langchain-chroma or langchain-community packages.
  • Persistence is configured via the persist_directory parameter, ensuring embeddings survive between Python sessions.
  • Integration follows a standard pattern: initialize embeddings → create Chroma store → load/split documents → add to store → build retriever → assemble RAG chain.
  • Flexibility allows swapping embedding models (Google Gemini, Ollama) and LLM backends while maintaining the same Chroma storage layer.
  • Embedchain offers a higher-level abstraction that configures Chroma automatically for simpler implementations.

Frequently Asked Questions

What is the difference between LangChain-Chroma and LangChain-Community Chroma?

LangChain-Chroma (langchain-chroma) is the modern, dedicated package that provides the Chroma class with updated async support and cleaner API patterns. LangChain-Community (langchain_community.vectorstores.Chroma) is the legacy wrapper that maintains backward compatibility. Both support the same core functionality including persist_directory and collection_name, but the dedicated package is recommended for new projects according to the repository examples.

How do I ensure Chroma data persists between Python sessions?

To make your Chroma vector store persistent, you must specify the persist_directory parameter when initializing the database. For example: Chroma(persist_directory="./my_chroma_db", collection_name="my_collection", embedding_function=embeddings). This saves the vector index to disk, allowing you to reload the same data in subsequent runs without re-embedding documents.

Can I use Chroma with local LLMs like Ollama?

Yes, Chroma works seamlessly with local LLMs through Ollama integration. As shown in rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py, you can use OllamaEmbeddings to generate vectors locally and store them in Chroma, then retrieve context to feed into a local Llama 3.1 model. This creates a fully offline RAG pipeline.

What embedding models work best with Chroma for RAG?

The repository demonstrates two primary approaches. Google Generative AI embeddings (models/embedding-001) provide high-quality vectors for cloud-based applications requiring advanced semantic understanding. Ollama embeddings (such as llama3.1 or nomic-embed-text) offer privacy-preserving local alternatives. Both integrate with Chroma through the embedding_function parameter, and the choice depends on whether you prioritize cloud performance or data privacy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →