# How to Use Chroma for RAG in Python: A Complete Implementation Guide

> Learn how to use Chroma for RAG in Python with this comprehensive guide. Integrate Chroma with LangChain to build powerful retrieval-augmented generation applications.

- Repository: [Shubham Saboo/awesome-llm-apps](https://github.com/shubhamsaboo/awesome-llm-apps)
- Tags: how-to-guide
- Published: 2026-02-16

---

**Chroma is a fast, persistent vector store that integrates with LangChain to power Retrieval-Augmented Generation (RAG) pipelines by storing document embeddings and retrieving relevant context for LLM queries.**

The `Shubhamsaboo/awesome-llm-apps` repository provides production-ready examples demonstrating how to implement Chroma-backed RAG systems in Python. Chroma enables local, persistent storage of vector embeddings, making it ideal for applications that require fast similarity search without external database dependencies.

## What Is Chroma and Why Use It for RAG?

Chroma is an open-source embedding database designed specifically for AI applications. In a RAG pipeline, Chroma serves as the **vector store** layer that persists document chunks as dense vectors. When a user submits a query, Chroma performs similarity search to retrieve the most relevant context, which is then fed into an LLM to generate grounded answers.

Key advantages of using Chroma for RAG include:

- **Persistent storage** via `persist_directory` for data survival across sessions
- **Collection-based organization** for managing multiple document sets
- **Native LangChain integration** through both `langchain-chroma` and `langchain-community` packages
- **Local execution** without requiring cloud vector database services

## Project Structure and Key Files

The repository contains several reference implementations that demonstrate Chroma integration patterns:

| File | Implementation Pattern |
|------|------------------------|
| [`rag_tutorials/rag_chain/app.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/rag_chain/app.py) | Full Streamlit application using `langchain-chroma` with Google Gemini embeddings and PDF ingestion |
| [`rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py) | Local-only pipeline using Ollama embeddings, web page loading, and Chroma persistence |
| [`advanced_llm_apps/chat_with_X_tutorials/chat_with_github/chat_github_llama3.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/advanced_llm_apps/chat_with_X_tutorials/chat_with_github/chat_github_llama3.py) | Embedchain abstraction layer configuring Chroma as the vector database backend |

## Step-by-Step Implementation

### 1. Install Dependencies

Install the required packages for Chroma and LangChain integration:

```bash
pip install langchain langchain-chroma langchain-community pypdf streamlit

```

For Google Gemini support, add:

```bash
pip install langchain-google-genai

```

For local Ollama embeddings:

```bash
pip install langchain-ollama

```

### 2. Initialize the Embedding Model

Choose an embedding model based on your deployment requirements.

**Google Gemini embeddings** (as implemented in [`rag_tutorials/rag_chain/app.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/rag_chain/app.py)):

```python
from langchain_google_genai import GoogleGenerativeAIEmbeddings

embedding_model = GoogleGenerativeAIEmbeddings(model="models/embedding-001")

```

**Ollama local embeddings** (as implemented in [`llama3.1_local_rag.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/llama3.1_local_rag.py)):

```python
from langchain_ollama import OllamaEmbeddings

embeddings = OllamaEmbeddings(
    model="llama3.1",
    base_url="http://127.0.0.1:11434"
)

```

### 3. Create a Persistent Chroma Vector Store

Initialize Chroma with a collection name and persistence directory.

Using `langchain-chroma` (modern approach):

```python
from langchain_chroma import Chroma

db = Chroma(
    collection_name="my_collection",
    embedding_function=embedding_model,
    persist_directory="./my_chroma_db"
)

```

Using `langchain-community` (alternative approach):

```python
from langchain_community.vectorstores import Chroma

vectorstore = Chroma.from_documents(
    documents=doc_chunks,
    embedding=embeddings,
    persist_directory="./my_chroma_db"
)

```

The `persist_directory` parameter ensures vector data survives between Python sessions. The `collection_name` parameter organizes related documents into logical groups.

### 4. Load and Split Documents

Process source documents into chunks suitable for embedding.

**PDF processing** (from [`rag_chain/app.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_chain/app.py)):

```python
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters.sentence_transformers import SentenceTransformersTokenTextSplitter

loader = PyPDFLoader("my_report.pdf")
raw_docs = loader.load()

splitter = SentenceTransformersTokenTextSplitter(
    model_name="sentence-transformers/all-mpnet-base-v2",
    chunk_size=100,
    chunk_overlap=50
)

doc_chunks = splitter.create_documents(
    [d.page_content for d in raw_docs],
    [d.metadata for d in raw_docs]
)

```

**Web page processing** (from [`llama3.1_local_rag.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/llama3.1_local_rag.py)):

```python
from langchain_community.document_loaders import WebBaseLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter

loader = WebBaseLoader("https://example.com/article")
docs = loader.load()

splitter = RecursiveCharacterTextSplitter(
    chunk_size=500,
    chunk_overlap=10
)
doc_chunks = splitter.split_documents(docs)

```

### 5. Add Documents to Chroma

Persist the processed chunks to the vector store:

```python
db.add_documents(doc_chunks)

```

For the community wrapper:

```python
vectorstore.add_documents(doc_chunks)

```

### 6. Build the Retriever

Configure similarity search to fetch relevant context:

```python
retriever = db.as_retriever(
    search_type="similarity",
    search_kwargs={"k": 5}
)

```

The `k` parameter controls the number of retrieved documents. Adjust based on your context window requirements.

### 7. Assemble the RAG Chain

Combine retrieval, formatting, and LLM generation using LangChain's pipe syntax:

```python
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough

def format_docs(docs):
    return "\n\n".join(doc.page_content for doc in docs)

prompt_template = ChatPromptTemplate.from_template(
    """
    You are an AI assistant. Answer the question using ONLY the following context:

    {context}

    Question: {question}
    """
)

chat_model = ChatGoogleGenerativeAI(
    model="gemini-1.5-pro",
    temperature=1,
)

rag_chain = {
    "context": retriever | format_docs,
    "question": RunnablePassthrough()
} | prompt_template | chat_model | StrOutputParser()

```

### 8. Invoke the Pipeline

Execute the RAG chain with a user query:

```python
response = rag_chain.invoke("What are the AI applications in drug discovery?")
print(response)

```

## Complete Minimal Example

Here is a condensed, runnable implementation combining all steps:

```python
import streamlit as st
from langchain_google_genai import GoogleGenerativeAIEmbeddings
from langchain_chroma import Chroma
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters.sentence_transformers import SentenceTransformersTokenTextSplitter
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_core.output_parsers import StrOutputParser
from langchain_google_genai import ChatGoogleGenerativeAI

# 1️⃣ Embeddings

embeddings = GoogleGenerativeAIEmbeddings(model="models/embedding-001")

# 2️⃣ Vector store

db = Chroma(
    collection_name="demo",
    embedding_function=embeddings,
    persist_directory="./demo_db"
)

# 3️⃣ Load & split a PDF

loader = PyPDFLoader("sample.pdf")
raw = loader.load()
splitter = SentenceTransformersTokenTextSplitter(
    model_name="sentence-transformers/all-mpnet-base-v2",
    chunk_size=120,
    chunk_overlap=30
)
chunks = splitter.create_documents(
    [d.page_content for d in raw],
    [d.metadata for d in raw]
)

# 4️⃣ Add to DB (first run only)

if not db._client.get_collection("demo"):
    db.add_documents(chunks)

# 5️⃣ Retriever

retriever = db.as_retriever(search_kwargs={"k": 4})

# 6️⃣ Prompt + LLM

prompt = ChatPromptTemplate.from_template(
    """Answer using only the context below.

    {context}

    Question: {question}"""
)
llm = ChatGoogleGenerativeAI(
    model="gemini-1.5-flash",
    api_key=st.secrets["GEMINI_API_KEY"],
    temperature=0.7
)

rag = {
    "context": retriever | (lambda docs: "\n\n".join(d.page_content for d in docs)),
    "question": RunnablePassthrough()
} | prompt | llm | StrOutputParser()

# 7️⃣ UI

st.title("📚 Chroma‑Powered RAG Demo")
query = st.text_input("Ask a question")
if query:
    st.write(rag.invoke(query))

```

## Alternative: Using Embedchain with Chroma

For higher-level abstraction, the repository also demonstrates using Embedchain to handle Chroma configuration automatically. In [`advanced_llm_apps/chat_with_X_tutorials/chat_with_github/chat_github_llama3.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/advanced_llm_apps/chat_with_X_tutorials/chat_with_github/chat_github_llama3.py), Chroma is configured as the vector database backend:

```python
from embedchain import App

app = App.from_config({
    "vectordb": {
        "provider": "chroma",
        "config": {
            "collection_name": "github_repo",
            "dir": "./chroma_db"
        }
    },
    "embedder": {
        "provider": "ollama",
        "config": {"model": "llama3.1"}
    }
})

```

This approach abstracts the embedding and storage logic while still leveraging Chroma's persistent storage capabilities.

## Summary

- **Chroma** provides persistent, local vector storage for RAG pipelines through the `langchain-chroma` or `langchain-community` packages.
- **Persistence** is configured via the `persist_directory` parameter, ensuring embeddings survive between Python sessions.
- **Integration** follows a standard pattern: initialize embeddings → create Chroma store → load/split documents → add to store → build retriever → assemble RAG chain.
- **Flexibility** allows swapping embedding models (Google Gemini, Ollama) and LLM backends while maintaining the same Chroma storage layer.
- **Embedchain** offers a higher-level abstraction that configures Chroma automatically for simpler implementations.

## Frequently Asked Questions

### What is the difference between LangChain-Chroma and LangChain-Community Chroma?

**LangChain-Chroma** (`langchain-chroma`) is the modern, dedicated package that provides the `Chroma` class with updated async support and cleaner API patterns. **LangChain-Community** (`langchain_community.vectorstores.Chroma`) is the legacy wrapper that maintains backward compatibility. Both support the same core functionality including `persist_directory` and `collection_name`, but the dedicated package is recommended for new projects according to the repository examples.

### How do I ensure Chroma data persists between Python sessions?

To make your Chroma vector store persistent, you must specify the `persist_directory` parameter when initializing the database. For example: `Chroma(persist_directory="./my_chroma_db", collection_name="my_collection", embedding_function=embeddings)`. This saves the vector index to disk, allowing you to reload the same data in subsequent runs without re-embedding documents.

### Can I use Chroma with local LLMs like Ollama?

Yes, Chroma works seamlessly with local LLMs through Ollama integration. As shown in [`rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py), you can use `OllamaEmbeddings` to generate vectors locally and store them in Chroma, then retrieve context to feed into a local Llama 3.1 model. This creates a fully offline RAG pipeline.

### What embedding models work best with Chroma for RAG?

The repository demonstrates two primary approaches. **Google Generative AI embeddings** (`models/embedding-001`) provide high-quality vectors for cloud-based applications requiring advanced semantic understanding. **Ollama embeddings** (such as `llama3.1` or `nomic-embed-text`) offer privacy-preserving local alternatives. Both integrate with Chroma through the `embedding_function` parameter, and the choice depends on whether you prioritize cloud performance or data privacy.