How to Use LangChain ChromaDB to Load Data from a Vector Database

To load data from a ChromaDB vector database in LangChain, instantiate the Chroma class from the langchain_chroma package with an embedding model, then use similarity_search() to query existing collections or add_documents() to populate new ones.

The langchain_chroma integration in the langchain-ai/langchain repository provides a first-class wrapper around ChromaDB, enabling seamless interaction with both local persistent storage and remote vector databases. Located in libs/partners/chroma, this implementation bridges LangChain's VectorStore interface with ChromaDB's high-performance embedding storage and retrieval capabilities.

Setting Up the LangChain ChromaDB Integration

The ChromaDB integration resides in the partners package at libs/partners/chroma/langchain_chroma/vectorstores.py. This file implements the Chroma class, which inherits from LangChain's abstract VectorStore base class and translates LangChain-style operations into native ChromaDB API calls.

To begin, install the partner package (typically available as langchain-chroma) and import the Chroma class alongside your preferred embedding provider.

Architecture of the Chroma Vector Store

The Chroma Class

In libs/partners/chroma/langchain_chroma/vectorstores.py, the Chroma class serves as the primary interface for vector operations. It accepts an Embedding implementation (such as OpenAIEmbeddings or HuggingFaceEmbeddings) and manages the conversion of documents into vector representations before storage or query execution.

Client Types and Persistence

The wrapper supports four distinct client configurations for different deployment scenarios:

  • chromadb.PersistentClient – Stores vectors on-disk using SQLite, ideal for production workloads requiring data durability.
  • chromadb.HttpClient – Connects to remote ChromaDB servers via REST API, enabling distributed architectures.
  • chromadb.CloudClient – Interfaces with Chroma Cloud's hosted SaaS offering.
  • chromadb.Client – Provides ephemeral in-memory storage for rapid prototyping and testing.

You can inject a pre-configured client instance through the client parameter or customize behavior via client_settings when initializing the Chroma class.

Collection Management

Each Chroma instance manages a single collection (created lazily upon first access). Collections store embeddings alongside metadata and optional document IDs. While the collection name is user-definable, tenant and database names default to Chroma's built-in constants unless explicitly overridden.

Loading and Querying Data from ChromaDB

Basic Local Persistent Setup

For most applications, you will create a persistent local vector store that survives between program executions. The persist_directory parameter specifies where ChromaDB stores its SQLite files.

from langchain_chroma import Chroma
from langchain_community.embeddings import OpenAIEmbeddings
from langchain.docstore.document import Document

# Initialize the embedding model

embeddings = OpenAIEmbeddings(model="text-embedding-ada-002")

# Create a persistent Chroma vector store

vectorstore = Chroma(
    collection_name="my_docs",
    embedding=embeddings,
    persist_directory="./chroma_db",
)

# Add documents to the vector database

docs = [
    Document(page_content="LangChain enables building LLM applications.", metadata={"source": "blog"}),
    Document(page_content="ChromaDB is a fast vector database.", metadata={"source": "docs"}),
]
vectorstore.add_documents(docs)

# Load data via similarity search

results = vectorstore.similarity_search("What is LangChain?", k=2)
for doc in results:
    print(doc.page_content, doc.metadata)

Connecting to Remote ChromaDB Servers

When operating with a remote ChromaDB HTTP server, instantiate an HttpClient and pass it to the Chroma constructor. This configuration bypasses local persistence and directs all operations to the specified endpoint.

import chromadb
from langchain_chroma import Chroma
from langchain_community.embeddings import HuggingFaceEmbeddings

# Configure remote client

client = chromadb.HttpClient(
    settings=chromadb.config.Settings(
        chroma_api_impl="rest",
        chroma_server_host="my-chroma-server.example.com",
        chroma_server_http_port=8000,
    )
)

# Initialize vector store with remote connection

vectorstore = Chroma(
    collection_name="remote_collection",
    embedding=HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2"),
    client=client,
)

# Query operations work identically to local storage

results = vectorstore.similarity_search("Distributed vector search", k=3)

Advanced Query Methods

Beyond basic similarity search, the Chroma class implements several retrieval strategies available in libs/partners/chroma/langchain_chroma/vectorstores.py:

Maximal Marginal Relevance (MMR) diversifies results to reduce redundancy while maintaining relevance to the query:

query = "How can I store embeddings efficiently?"
diverse_results = vectorstore.max_marginal_relevance_search(query, k=5, fetch_k=20)

for doc in diverse_results:
    print("- ", doc.page_content)

Direct Vector Querying allows you to search using pre-computed embeddings when you already have vector representations:

import numpy as np

# Pre-computed embedding vector (e.g., from a custom model)

precomputed_vector = np.random.rand(1536).tolist()

hits = vectorstore.similarity_search_by_vector(precomputed_vector, k=3)
for hit in hits:
    print(hit.page_content)

Summary

  • Install langchain_chroma from the partners package to access the ChromaDB wrapper implementation in libs/partners/chroma.
  • Choose your persistence model by selecting between PersistentClient (local disk), HttpClient (remote server), CloudClient (hosted), or in-memory Client.
  • Use similarity_search() for standard K-NN retrieval or max_marginal_relevance_search() for diverse result sets.
  • Reference the source at libs/partners/chroma/langchain_chroma/vectorstores.py to understand how the Chroma class translates LangChain operations into ChromaDB API calls.

Frequently Asked Questions

How do I install the LangChain ChromaDB package?

Install the langchain-chroma package using your Python package manager. This partner package is defined in libs/partners/chroma/pyproject.toml and declares chromadb as a dependency alongside the core LangChain interfaces.

Can I use an existing ChromaDB collection with LangChain?

Yes. When instantiating the Chroma class, specify the existing collection name via the collection_name parameter. The wrapper will connect to the existing collection rather than creating a new one, allowing you to query data previously stored via native ChromaDB APIs or other applications.

How do I filter results by metadata when querying?

Pass ChromaDB-native where clauses through the filter parameter in similarity_search() or related methods. The wrapper passes these filters directly to the underlying ChromaDB collection query, enabling metadata-constrained retrieval alongside vector similarity.

What embedding models work with LangChain ChromaDB?

Any LangChain Embeddings implementation is compatible, including OpenAIEmbeddings, HuggingFaceEmbeddings, OllamaEmbeddings, or custom implementations. The Chroma class stores vectors produced by your chosen embedder and automatically handles dimensionality alignment during query time.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →