# How to Use LangChain ChromaDB to Load Data from a Vector Database

> Learn how to use LangChain ChromaDB to load data from your vector database. Get started with Chroma DB for efficient data retrieval and storage with this quick guide.

- Repository: [LangChain/langchain](https://github.com/langchain-ai/langchain)
- Tags: how-to-guide
- Published: 2026-02-14

---

**To load data from a ChromaDB vector database in LangChain, instantiate the `Chroma` class from the `langchain_chroma` package with an embedding model, then use `similarity_search()` to query existing collections or `add_documents()` to populate new ones.**

The `langchain_chroma` integration in the `langchain-ai/langchain` repository provides a first-class wrapper around ChromaDB, enabling seamless interaction with both local persistent storage and remote vector databases. Located in `libs/partners/chroma`, this implementation bridges LangChain's `VectorStore` interface with ChromaDB's high-performance embedding storage and retrieval capabilities.

## Setting Up the LangChain ChromaDB Integration

The ChromaDB integration resides in the **partners package** at [`libs/partners/chroma/langchain_chroma/vectorstores.py`](https://github.com/langchain-ai/langchain/blob/main/libs/partners/chroma/langchain_chroma/vectorstores.py). This file implements the `Chroma` class, which inherits from LangChain's abstract `VectorStore` base class and translates LangChain-style operations into native ChromaDB API calls.

To begin, install the partner package (typically available as `langchain-chroma`) and import the `Chroma` class alongside your preferred embedding provider.

## Architecture of the Chroma Vector Store

### The Chroma Class

In [`libs/partners/chroma/langchain_chroma/vectorstores.py`](https://github.com/langchain-ai/langchain/blob/main/libs/partners/chroma/langchain_chroma/vectorstores.py), the `Chroma` class serves as the primary interface for vector operations. It accepts an **Embedding** implementation (such as `OpenAIEmbeddings` or `HuggingFaceEmbeddings`) and manages the conversion of documents into vector representations before storage or query execution.

### Client Types and Persistence

The wrapper supports four distinct client configurations for different deployment scenarios:

- **`chromadb.PersistentClient`** – Stores vectors on-disk using SQLite, ideal for production workloads requiring data durability.
- **`chromadb.HttpClient`** – Connects to remote ChromaDB servers via REST API, enabling distributed architectures.
- **`chromadb.CloudClient`** – Interfaces with Chroma Cloud's hosted SaaS offering.
- **`chromadb.Client`** – Provides ephemeral in-memory storage for rapid prototyping and testing.

You can inject a pre-configured client instance through the `client` parameter or customize behavior via `client_settings` when initializing the `Chroma` class.

### Collection Management

Each `Chroma` instance manages a single **collection** (created lazily upon first access). Collections store embeddings alongside metadata and optional document IDs. While the collection name is user-definable, tenant and database names default to Chroma's built-in constants unless explicitly overridden.

## Loading and Querying Data from ChromaDB

### Basic Local Persistent Setup

For most applications, you will create a persistent local vector store that survives between program executions. The `persist_directory` parameter specifies where ChromaDB stores its SQLite files.

```python
from langchain_chroma import Chroma
from langchain_community.embeddings import OpenAIEmbeddings
from langchain.docstore.document import Document

# Initialize the embedding model

embeddings = OpenAIEmbeddings(model="text-embedding-ada-002")

# Create a persistent Chroma vector store

vectorstore = Chroma(
    collection_name="my_docs",
    embedding=embeddings,
    persist_directory="./chroma_db",
)

# Add documents to the vector database

docs = [
    Document(page_content="LangChain enables building LLM applications.", metadata={"source": "blog"}),
    Document(page_content="ChromaDB is a fast vector database.", metadata={"source": "docs"}),
]
vectorstore.add_documents(docs)

# Load data via similarity search

results = vectorstore.similarity_search("What is LangChain?", k=2)
for doc in results:
    print(doc.page_content, doc.metadata)

```

### Connecting to Remote ChromaDB Servers

When operating with a remote ChromaDB HTTP server, instantiate an `HttpClient` and pass it to the `Chroma` constructor. This configuration bypasses local persistence and directs all operations to the specified endpoint.

```python
import chromadb
from langchain_chroma import Chroma
from langchain_community.embeddings import HuggingFaceEmbeddings

# Configure remote client

client = chromadb.HttpClient(
    settings=chromadb.config.Settings(
        chroma_api_impl="rest",
        chroma_server_host="my-chroma-server.example.com",
        chroma_server_http_port=8000,
    )
)

# Initialize vector store with remote connection

vectorstore = Chroma(
    collection_name="remote_collection",
    embedding=HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2"),
    client=client,
)

# Query operations work identically to local storage

results = vectorstore.similarity_search("Distributed vector search", k=3)

```

### Advanced Query Methods

Beyond basic similarity search, the `Chroma` class implements several retrieval strategies available in [`libs/partners/chroma/langchain_chroma/vectorstores.py`](https://github.com/langchain-ai/langchain/blob/main/libs/partners/chroma/langchain_chroma/vectorstores.py):

**Maximal Marginal Relevance (MMR)** diversifies results to reduce redundancy while maintaining relevance to the query:

```python
query = "How can I store embeddings efficiently?"
diverse_results = vectorstore.max_marginal_relevance_search(query, k=5, fetch_k=20)

for doc in diverse_results:
    print("- ", doc.page_content)

```

**Direct Vector Querying** allows you to search using pre-computed embeddings when you already have vector representations:

```python
import numpy as np

# Pre-computed embedding vector (e.g., from a custom model)

precomputed_vector = np.random.rand(1536).tolist()

hits = vectorstore.similarity_search_by_vector(precomputed_vector, k=3)
for hit in hits:
    print(hit.page_content)

```

## Summary

- **Install `langchain_chroma`** from the partners package to access the ChromaDB wrapper implementation in `libs/partners/chroma`.
- **Choose your persistence model** by selecting between `PersistentClient` (local disk), `HttpClient` (remote server), `CloudClient` (hosted), or in-memory `Client`.
- **Use `similarity_search()`** for standard K-NN retrieval or `max_marginal_relevance_search()` for diverse result sets.
- **Reference the source** at [`libs/partners/chroma/langchain_chroma/vectorstores.py`](https://github.com/langchain-ai/langchain/blob/main/libs/partners/chroma/langchain_chroma/vectorstores.py) to understand how the `Chroma` class translates LangChain operations into ChromaDB API calls.

## Frequently Asked Questions

### How do I install the LangChain ChromaDB package?

Install the `langchain-chroma` package using your Python package manager. This partner package is defined in [`libs/partners/chroma/pyproject.toml`](https://github.com/langchain-ai/langchain/blob/main/libs/partners/chroma/pyproject.toml) and declares `chromadb` as a dependency alongside the core LangChain interfaces.

### Can I use an existing ChromaDB collection with LangChain?

Yes. When instantiating the `Chroma` class, specify the existing collection name via the `collection_name` parameter. The wrapper will connect to the existing collection rather than creating a new one, allowing you to query data previously stored via native ChromaDB APIs or other applications.

### How do I filter results by metadata when querying?

Pass ChromaDB-native `where` clauses through the `filter` parameter in `similarity_search()` or related methods. The wrapper passes these filters directly to the underlying ChromaDB collection query, enabling metadata-constrained retrieval alongside vector similarity.

### What embedding models work with LangChain ChromaDB?

Any LangChain `Embeddings` implementation is compatible, including `OpenAIEmbeddings`, `HuggingFaceEmbeddings`, `OllamaEmbeddings`, or custom implementations. The `Chroma` class stores vectors produced by your chosen embedder and automatically handles dimensionality alignment during query time.