# Implementing and Managing Long-Term Memory for Agents Using RoleZero in MetaGPT

> Learn how to implement and manage long-term memory for agents with RoleZero in MetaGPT. Explore its hybrid memory architecture for persistent recall and fast access to interactions.

- Repository: [FoundationAgents/MetaGPT](https://github.com/FoundationAgents/MetaGPT)
- Tags: how-to-guide
- Published: 2026-03-04

---

**MetaGPT's RoleZero role supports a hybrid memory architecture that combines an in-process short-term buffer with a persistent Chroma-backed vector store, enabling agents to recall historical context across sessions while maintaining fast access to recent interactions.**

MetaGPT equips autonomous agents with sophisticated memory capabilities through its hierarchical storage system. When implementing and managing long-term memory for agents using RoleZero, developers enable persistent knowledge retention that survives process restarts while keeping active context in high-speed memory. This architecture automatically transfers aging messages from short-term storage to a vector database, ensuring agents maintain relevant historical context without overwhelming LLM prompts.

## Configuring Long-Term Memory in RoleZeroConfig

The memory system is controlled through `RoleZeroConfig`, located in [`metagpt/configs/role_zero_config.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/configs/role_zero_config.py). This configuration class defines the parameters that govern how the hybrid memory system behaves at runtime.

```python

# metagpt/configs/role_zero_config.py

class RoleZeroConfig(YamlModel):
    enable_longterm_memory: bool = Field(default=False, description="Whether to use long‑term memory.")
    longterm_memory_persist_path: str = Field(default=".role_memory_data")
    memory_k: int = Field(default=200, description="Capacity of short‑term memory.")
    similarity_top_k: int = Field(default=5, description="Number of long‑term memories to retrieve.")
    use_llm_ranker: bool = Field(default=False)

```

To activate the feature, set `enable_longterm_memory: true` in your configuration. The `memory_k` parameter defines the short-term buffer capacity (default 200 messages), while `similarity_top_k` controls how many historical memories are retrieved during context building (default 5). The `longterm_memory_persist_path` specifies where ChromaDB stores vector embeddings, defaulting to `.role_memory_data` in the project root.

## The RoleZeroLongTermMemory Engine

The core implementation resides in [`metagpt/memory/role_zero_memory.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/memory/role_zero_memory.py) (lines 1-182), where the `RoleZeroLongTermMemory` class extends the base `Memory` class with RAG (Retrieval-Augmented Generation) capabilities. This engine manages the seamless transfer of data between volatile short-term storage and persistent long-term storage.

### Short-Term Buffer and Overflow Handling

The short-term memory operates as an in-process list providing **O(1)** access to recent messages, which is essential for LLM prompt construction. When the `add()` method receives a new `Message` (lines 71-80), it first stores the item in the short-term buffer via `super().add()`.

Once the buffer exceeds `memory_k` items, the `_should_use_longterm_memory_for_add()` method (lines 97-103) triggers `_transfer_to_longterm_memory()`. This process extracts the oldest message outside the active window using `_get_longterm_memory_item()` and persists it to the Chroma vector store via `_add_to_longterm_memory()` (lines 119-133).

```python

# Simplified logic from metagpt/memory/role_zero_memory.py

def add(self, message: Message):
    super().add(message)  # Add to short-term list

    if self._should_use_longterm_memory_for_add():
        old_message = self._get_longterm_memory_item()
        self._add_to_longterm_memory(old_message)

```

### Retrieval with Similarity Search

When building context for LLM requests, the `get(k)` method (lines 81-95) returns the most recent *k* messages from the short-term buffer. If `_should_use_longterm_memory_for_get()` returns true—meaning the buffer is full and the last message originated from a user requirement (lines 105-117)—the system augments recent context with historical data.

The `_build_longterm_memory_query()` method (lines 173-182) constructs a search query from the most recent user message content. The RAG engine, which lazy-loads a `SimpleEngine` coupling Chroma with an optional LLM Ranker (lines 39-69), then retrieves `similarity_top_k` relevant vectors and injects them into the returned context list.

### Safety and Error Handling

All methods interacting with external RAG components use the `@handle_exception` decorator (lines 131-155). This ensures that vector store failures emit logs rather than crashing the agent, maintaining system stability during memory operations.

## Wiring Memory into the RoleZero Agent

The integration occurs in [`metagpt/roles/di/role_zero.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/roles/di/role_zero.py) through a Pydantic model validator. When a `RoleZero` instance initializes, the `set_longterm_memory` validator (lines 71-91) checks `self.config.role_zero.enable_longterm_memory`. If enabled, it replaces the default `rc.memory` with a `RoleZeroLongTermMemory` instance.

```python

# metagpt/roles/di/role_zero.py (excerpt)

@model_validator(mode="after")
def set_longterm_memory(self) -> "RoleZero":
    if self.config.role_zero.enable_longterm_memory:
        self.rc.memory = RoleZeroLongTermMemory(
            **self.rc.memory.model_dump(),
            persist_path=self.config.role_zero.longterm_memory_persist_path,
            collection_name=self.name.replace(" ", ""),
            memory_k=self.config.role_zero.memory_k,
            similarity_top_k=self.config.role_zero.similarity_top_k,
            use_llm_ranker=self.config.role_zero.use_llm_ranker,
        )
        logger.info(f"Long‑term memory set for role '{self.name}'")
    return self

```

This injection ensures that all subsequent calls to `self.rc.memory.add()` and `self.rc.memory.get()` automatically execute the hybrid memory logic without modifying the role's business logic.

## Runtime Execution Flow

Understanding the runtime behavior helps optimize memory usage in production environments.

1. **Initialization**: The MetaGPT framework loads configuration from [`config/config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config/config2.yaml). If `role_zero.enable_longterm_memory` is true, the validator instantiates the RAG-backed memory engine before the agent processes any messages.

2. **Message Ingestion**: As the role processes user requests, each `Message` passes through `self.rc.memory.add()`. Recent messages remain in fast memory while aged entries automatically migrate to Chroma.

3. **Context Assembly**: Before LLM invocation, the role calls `self.rc.memory.get(k)`. The system returns recent short-term messages augmented with up to `similarity_top_k` relevant long-term memories, keeping token counts efficient while preserving historical context.

4. **Persistence**: Chroma stores vectors in the configured `persist_path`, ensuring data survives process restarts. Agents retain knowledge across sessions, enabling long-running workflows that span days or weeks.

## Implementation Examples

### Enabling Long-Term Memory via Configuration

Create or modify [`config/config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config/config2.yaml) to activate the feature:

```yaml

# config/config2.yaml

role_zero:
  enable_longterm_memory: true
  longterm_memory_persist_path: .role_memory_data
  memory_k: 200
  similarity_top_k: 5
  use_llm_ranker: false

```

### Instantiating a RoleZero Agent with Persistent Memory

```python
from metagpt.configs import Config
from metagpt.roles.di.role_zero import RoleZero

cfg = Config.from_yaml("config/config2.yaml")
role = RoleZero(config=cfg)  # Validator automatically configures memory

```

The `set_longterm_memory` validator runs automatically during instantiation, though you can verify activation by checking the role logs for "Long-term memory set for role" messages.

### Adding Messages and Triggering Transfer

```python
from metagpt.schema import UserMessage

# Add to memory - automatically handles short-term vs long-term routing

msg = UserMessage(content="Analyze the Q3 financial reports.")
role.rc.memory.add(msg)

# When buffer exceeds memory_k (200), oldest messages move to Chroma

```

### Retrieving Augmented Context

```python

# Retrieve recent context plus relevant historical memories

context = role.rc.memory.get(k=10)  

# Returns up to 10 recent messages + up to 5 similar long-term items

```

The returned list combines recent short-term interactions with semantically relevant historical data, ready for LLM prompt construction.

### Inspecting Stored Vectors (Debugging)

For troubleshooting or analysis, access the underlying Chroma collection directly:

```python

# Access the RAG engine's retriever

engine = role.rc.memory.rag_engine
stored_ids = engine.retriever.collection.get(include=["ids"]).ids
print(f"Stored memory count: {len(stored_ids)}")

```

## Summary

- **Configuration**: Enable long-term memory in `RoleZeroConfig` via `enable_longterm_memory: true` in [`config/config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config/config2.yaml)
- **Architecture**: The `RoleZeroLongTermMemory` class in [`metagpt/memory/role_zero_memory.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/memory/role_zero_memory.py) manages a hybrid system combining fast short-term buffers with Chroma-backed vector storage
- **Integration**: The `set_longterm_memory` validator in [`metagpt/roles/di/role_zero.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/roles/di/role_zero.py) automatically wires the memory engine into the agent's runtime context
- **Operation**: Messages flow through `add()` for storage and `get()` for retrieval, with automatic overflow handling that preserves older data in the vector store while keeping recent context in memory
- **Persistence**: Data survives process restarts through ChromaDB's persistent storage at the configured `persist_path`

## Frequently Asked Questions

### How does RoleZero decide when to move messages from short-term to long-term memory?

The `RoleZeroLongTermMemory` class uses the `_should_use_longterm_memory_for_add()` method (lines 97-103 in [`role_zero_memory.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/role_zero_memory.py)) to monitor buffer capacity. When the short-term list exceeds `memory_k` items (default 200), the system automatically transfers the oldest message to the Chroma vector store via `_transfer_to_longterm_memory()`. This overflow handling ensures the fast memory buffer remains bounded while preserving historical context for future retrieval.

### What vector database does MetaGPT use for RoleZero's long-term memory?

MetaGPT uses **ChromaDB** as the underlying vector store for `RoleZeroLongTermMemory`. The implementation initializes a `SimpleEngine` (lines 39-69) that couples Chroma with optional LLM-based ranking. Vector embeddings persist to disk at the path specified by `longterm_memory_persist_path` (default `.role_memory_data`), enabling data survival across process restarts according to the MetaGPT source code.

### Can I adjust how many historical memories are retrieved during context building?

Yes. The `similarity_top_k` parameter in `RoleZeroConfig` controls retrieval volume, defaulting to 5 items. When the `get()` method detects a full buffer and user requirement context (as determined by `_should_use_longterm_memory_for_get`), it queries the RAG engine for this number of semantically similar historical messages. Increasing this value provides more historical context but consumes additional tokens in the LLM prompt.

### What happens if the Chroma vector store connection fails during runtime?

The memory implementation includes the `@handle_exception` decorator on methods interacting with external RAG components (lines 131-155). These decorators ensure that vector store failures emit log messages rather than raising exceptions that would crash the agent. The system degrades gracefully by continuing with short-term memory only until the external service recovers.