How to Use Chroma for RAG in Python: A Complete Implementation Guide
Chroma is a fast, persistent vector store that integrates with LangChain to power Retrieval-Augmented Generation (RAG) pipelines by storing document embeddings and retrieving relevant context for LLM queries.
The Shubhamsaboo/awesome-llm-apps repository provides production-ready examples demonstrating how to implement Chroma-backed RAG systems in Python. Chroma enables local, persistent storage of vector embeddings, making it ideal for applications that require fast similarity search without external database dependencies.
What Is Chroma and Why Use It for RAG?
Chroma is an open-source embedding database designed specifically for AI applications. In a RAG pipeline, Chroma serves as the vector store layer that persists document chunks as dense vectors. When a user submits a query, Chroma performs similarity search to retrieve the most relevant context, which is then fed into an LLM to generate grounded answers.
Key advantages of using Chroma for RAG include:
- Persistent storage via
persist_directoryfor data survival across sessions - Collection-based organization for managing multiple document sets
- Native LangChain integration through both
langchain-chromaandlangchain-communitypackages - Local execution without requiring cloud vector database services
Project Structure and Key Files
The repository contains several reference implementations that demonstrate Chroma integration patterns:
| File | Implementation Pattern |
|---|---|
rag_tutorials/rag_chain/app.py |
Full Streamlit application using langchain-chroma with Google Gemini embeddings and PDF ingestion |
rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py |
Local-only pipeline using Ollama embeddings, web page loading, and Chroma persistence |
advanced_llm_apps/chat_with_X_tutorials/chat_with_github/chat_github_llama3.py |
Embedchain abstraction layer configuring Chroma as the vector database backend |
Step-by-Step Implementation
1. Install Dependencies
Install the required packages for Chroma and LangChain integration:
pip install langchain langchain-chroma langchain-community pypdf streamlit
For Google Gemini support, add:
pip install langchain-google-genai
For local Ollama embeddings:
pip install langchain-ollama
2. Initialize the Embedding Model
Choose an embedding model based on your deployment requirements.
Google Gemini embeddings (as implemented in rag_tutorials/rag_chain/app.py):
from langchain_google_genai import GoogleGenerativeAIEmbeddings
embedding_model = GoogleGenerativeAIEmbeddings(model="models/embedding-001")
Ollama local embeddings (as implemented in llama3.1_local_rag.py):
from langchain_ollama import OllamaEmbeddings
embeddings = OllamaEmbeddings(
model="llama3.1",
base_url="http://127.0.0.1:11434"
)
3. Create a Persistent Chroma Vector Store
Initialize Chroma with a collection name and persistence directory.
Using langchain-chroma (modern approach):
from langchain_chroma import Chroma
db = Chroma(
collection_name="my_collection",
embedding_function=embedding_model,
persist_directory="./my_chroma_db"
)
Using langchain-community (alternative approach):
from langchain_community.vectorstores import Chroma
vectorstore = Chroma.from_documents(
documents=doc_chunks,
embedding=embeddings,
persist_directory="./my_chroma_db"
)
The persist_directory parameter ensures vector data survives between Python sessions. The collection_name parameter organizes related documents into logical groups.
4. Load and Split Documents
Process source documents into chunks suitable for embedding.
PDF processing (from rag_chain/app.py):
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters.sentence_transformers import SentenceTransformersTokenTextSplitter
loader = PyPDFLoader("my_report.pdf")
raw_docs = loader.load()
splitter = SentenceTransformersTokenTextSplitter(
model_name="sentence-transformers/all-mpnet-base-v2",
chunk_size=100,
chunk_overlap=50
)
doc_chunks = splitter.create_documents(
[d.page_content for d in raw_docs],
[d.metadata for d in raw_docs]
)
Web page processing (from llama3.1_local_rag.py):
from langchain_community.document_loaders import WebBaseLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
loader = WebBaseLoader("https://example.com/article")
docs = loader.load()
splitter = RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=10
)
doc_chunks = splitter.split_documents(docs)
5. Add Documents to Chroma
Persist the processed chunks to the vector store:
db.add_documents(doc_chunks)
For the community wrapper:
vectorstore.add_documents(doc_chunks)
6. Build the Retriever
Configure similarity search to fetch relevant context:
retriever = db.as_retriever(
search_type="similarity",
search_kwargs={"k": 5}
)
The k parameter controls the number of retrieved documents. Adjust based on your context window requirements.
7. Assemble the RAG Chain
Combine retrieval, formatting, and LLM generation using LangChain's pipe syntax:
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough
def format_docs(docs):
return "\n\n".join(doc.page_content for doc in docs)
prompt_template = ChatPromptTemplate.from_template(
"""
You are an AI assistant. Answer the question using ONLY the following context:
{context}
Question: {question}
"""
)
chat_model = ChatGoogleGenerativeAI(
model="gemini-1.5-pro",
temperature=1,
)
rag_chain = {
"context": retriever | format_docs,
"question": RunnablePassthrough()
} | prompt_template | chat_model | StrOutputParser()
8. Invoke the Pipeline
Execute the RAG chain with a user query:
response = rag_chain.invoke("What are the AI applications in drug discovery?")
print(response)
Complete Minimal Example
Here is a condensed, runnable implementation combining all steps:
import streamlit as st
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from langchain_chroma import Chroma
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters.sentence_transformers import SentenceTransformersTokenTextSplitter
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_core.output_parsers import StrOutputParser
from langchain_google_genai import ChatGoogleGenerativeAI
# 1️⃣ Embeddings
embeddings = GoogleGenerativeAIEmbeddings(model="models/embedding-001")
# 2️⃣ Vector store
db = Chroma(
collection_name="demo",
embedding_function=embeddings,
persist_directory="./demo_db"
)
# 3️⃣ Load & split a PDF
loader = PyPDFLoader("sample.pdf")
raw = loader.load()
splitter = SentenceTransformersTokenTextSplitter(
model_name="sentence-transformers/all-mpnet-base-v2",
chunk_size=120,
chunk_overlap=30
)
chunks = splitter.create_documents(
[d.page_content for d in raw],
[d.metadata for d in raw]
)
# 4️⃣ Add to DB (first run only)
if not db._client.get_collection("demo"):
db.add_documents(chunks)
# 5️⃣ Retriever
retriever = db.as_retriever(search_kwargs={"k": 4})
# 6️⃣ Prompt + LLM
prompt = ChatPromptTemplate.from_template(
"""Answer using only the context below.
{context}
Question: {question}"""
)
llm = ChatGoogleGenerativeAI(
model="gemini-1.5-flash",
api_key=st.secrets["GEMINI_API_KEY"],
temperature=0.7
)
rag = {
"context": retriever | (lambda docs: "\n\n".join(d.page_content for d in docs)),
"question": RunnablePassthrough()
} | prompt | llm | StrOutputParser()
# 7️⃣ UI
st.title("📚 Chroma‑Powered RAG Demo")
query = st.text_input("Ask a question")
if query:
st.write(rag.invoke(query))
Alternative: Using Embedchain with Chroma
For higher-level abstraction, the repository also demonstrates using Embedchain to handle Chroma configuration automatically. In advanced_llm_apps/chat_with_X_tutorials/chat_with_github/chat_github_llama3.py, Chroma is configured as the vector database backend:
from embedchain import App
app = App.from_config({
"vectordb": {
"provider": "chroma",
"config": {
"collection_name": "github_repo",
"dir": "./chroma_db"
}
},
"embedder": {
"provider": "ollama",
"config": {"model": "llama3.1"}
}
})
This approach abstracts the embedding and storage logic while still leveraging Chroma's persistent storage capabilities.
Summary
- Chroma provides persistent, local vector storage for RAG pipelines through the
langchain-chromaorlangchain-communitypackages. - Persistence is configured via the
persist_directoryparameter, ensuring embeddings survive between Python sessions. - Integration follows a standard pattern: initialize embeddings → create Chroma store → load/split documents → add to store → build retriever → assemble RAG chain.
- Flexibility allows swapping embedding models (Google Gemini, Ollama) and LLM backends while maintaining the same Chroma storage layer.
- Embedchain offers a higher-level abstraction that configures Chroma automatically for simpler implementations.
Frequently Asked Questions
What is the difference between LangChain-Chroma and LangChain-Community Chroma?
LangChain-Chroma (langchain-chroma) is the modern, dedicated package that provides the Chroma class with updated async support and cleaner API patterns. LangChain-Community (langchain_community.vectorstores.Chroma) is the legacy wrapper that maintains backward compatibility. Both support the same core functionality including persist_directory and collection_name, but the dedicated package is recommended for new projects according to the repository examples.
How do I ensure Chroma data persists between Python sessions?
To make your Chroma vector store persistent, you must specify the persist_directory parameter when initializing the database. For example: Chroma(persist_directory="./my_chroma_db", collection_name="my_collection", embedding_function=embeddings). This saves the vector index to disk, allowing you to reload the same data in subsequent runs without re-embedding documents.
Can I use Chroma with local LLMs like Ollama?
Yes, Chroma works seamlessly with local LLMs through Ollama integration. As shown in rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py, you can use OllamaEmbeddings to generate vectors locally and store them in Chroma, then retrieve context to feed into a local Llama 3.1 model. This creates a fully offline RAG pipeline.
What embedding models work best with Chroma for RAG?
The repository demonstrates two primary approaches. Google Generative AI embeddings (models/embedding-001) provide high-quality vectors for cloud-based applications requiring advanced semantic understanding. Ollama embeddings (such as llama3.1 or nomic-embed-text) offer privacy-preserving local alternatives. Both integrate with Chroma through the embedding_function parameter, and the choice depends on whether you prioritize cloud performance or data privacy.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →