Implementing and Managing Long-Term Memory for Agents Using RoleZero in MetaGPT
MetaGPT's RoleZero role supports a hybrid memory architecture that combines an in-process short-term buffer with a persistent Chroma-backed vector store, enabling agents to recall historical context across sessions while maintaining fast access to recent interactions.
MetaGPT equips autonomous agents with sophisticated memory capabilities through its hierarchical storage system. When implementing and managing long-term memory for agents using RoleZero, developers enable persistent knowledge retention that survives process restarts while keeping active context in high-speed memory. This architecture automatically transfers aging messages from short-term storage to a vector database, ensuring agents maintain relevant historical context without overwhelming LLM prompts.
Configuring Long-Term Memory in RoleZeroConfig
The memory system is controlled through RoleZeroConfig, located in metagpt/configs/role_zero_config.py. This configuration class defines the parameters that govern how the hybrid memory system behaves at runtime.
# metagpt/configs/role_zero_config.py
class RoleZeroConfig(YamlModel):
enable_longterm_memory: bool = Field(default=False, description="Whether to use long‑term memory.")
longterm_memory_persist_path: str = Field(default=".role_memory_data")
memory_k: int = Field(default=200, description="Capacity of short‑term memory.")
similarity_top_k: int = Field(default=5, description="Number of long‑term memories to retrieve.")
use_llm_ranker: bool = Field(default=False)
To activate the feature, set enable_longterm_memory: true in your configuration. The memory_k parameter defines the short-term buffer capacity (default 200 messages), while similarity_top_k controls how many historical memories are retrieved during context building (default 5). The longterm_memory_persist_path specifies where ChromaDB stores vector embeddings, defaulting to .role_memory_data in the project root.
The RoleZeroLongTermMemory Engine
The core implementation resides in metagpt/memory/role_zero_memory.py (lines 1-182), where the RoleZeroLongTermMemory class extends the base Memory class with RAG (Retrieval-Augmented Generation) capabilities. This engine manages the seamless transfer of data between volatile short-term storage and persistent long-term storage.
Short-Term Buffer and Overflow Handling
The short-term memory operates as an in-process list providing O(1) access to recent messages, which is essential for LLM prompt construction. When the add() method receives a new Message (lines 71-80), it first stores the item in the short-term buffer via super().add().
Once the buffer exceeds memory_k items, the _should_use_longterm_memory_for_add() method (lines 97-103) triggers _transfer_to_longterm_memory(). This process extracts the oldest message outside the active window using _get_longterm_memory_item() and persists it to the Chroma vector store via _add_to_longterm_memory() (lines 119-133).
# Simplified logic from metagpt/memory/role_zero_memory.py
def add(self, message: Message):
super().add(message) # Add to short-term list
if self._should_use_longterm_memory_for_add():
old_message = self._get_longterm_memory_item()
self._add_to_longterm_memory(old_message)
Retrieval with Similarity Search
When building context for LLM requests, the get(k) method (lines 81-95) returns the most recent k messages from the short-term buffer. If _should_use_longterm_memory_for_get() returns true—meaning the buffer is full and the last message originated from a user requirement (lines 105-117)—the system augments recent context with historical data.
The _build_longterm_memory_query() method (lines 173-182) constructs a search query from the most recent user message content. The RAG engine, which lazy-loads a SimpleEngine coupling Chroma with an optional LLM Ranker (lines 39-69), then retrieves similarity_top_k relevant vectors and injects them into the returned context list.
Safety and Error Handling
All methods interacting with external RAG components use the @handle_exception decorator (lines 131-155). This ensures that vector store failures emit logs rather than crashing the agent, maintaining system stability during memory operations.
Wiring Memory into the RoleZero Agent
The integration occurs in metagpt/roles/di/role_zero.py through a Pydantic model validator. When a RoleZero instance initializes, the set_longterm_memory validator (lines 71-91) checks self.config.role_zero.enable_longterm_memory. If enabled, it replaces the default rc.memory with a RoleZeroLongTermMemory instance.
# metagpt/roles/di/role_zero.py (excerpt)
@model_validator(mode="after")
def set_longterm_memory(self) -> "RoleZero":
if self.config.role_zero.enable_longterm_memory:
self.rc.memory = RoleZeroLongTermMemory(
**self.rc.memory.model_dump(),
persist_path=self.config.role_zero.longterm_memory_persist_path,
collection_name=self.name.replace(" ", ""),
memory_k=self.config.role_zero.memory_k,
similarity_top_k=self.config.role_zero.similarity_top_k,
use_llm_ranker=self.config.role_zero.use_llm_ranker,
)
logger.info(f"Long‑term memory set for role '{self.name}'")
return self
This injection ensures that all subsequent calls to self.rc.memory.add() and self.rc.memory.get() automatically execute the hybrid memory logic without modifying the role's business logic.
Runtime Execution Flow
Understanding the runtime behavior helps optimize memory usage in production environments.
-
Initialization: The MetaGPT framework loads configuration from
config/config2.yaml. Ifrole_zero.enable_longterm_memoryis true, the validator instantiates the RAG-backed memory engine before the agent processes any messages. -
Message Ingestion: As the role processes user requests, each
Messagepasses throughself.rc.memory.add(). Recent messages remain in fast memory while aged entries automatically migrate to Chroma. -
Context Assembly: Before LLM invocation, the role calls
self.rc.memory.get(k). The system returns recent short-term messages augmented with up tosimilarity_top_krelevant long-term memories, keeping token counts efficient while preserving historical context. -
Persistence: Chroma stores vectors in the configured
persist_path, ensuring data survives process restarts. Agents retain knowledge across sessions, enabling long-running workflows that span days or weeks.
Implementation Examples
Enabling Long-Term Memory via Configuration
Create or modify config/config2.yaml to activate the feature:
# config/config2.yaml
role_zero:
enable_longterm_memory: true
longterm_memory_persist_path: .role_memory_data
memory_k: 200
similarity_top_k: 5
use_llm_ranker: false
Instantiating a RoleZero Agent with Persistent Memory
from metagpt.configs import Config
from metagpt.roles.di.role_zero import RoleZero
cfg = Config.from_yaml("config/config2.yaml")
role = RoleZero(config=cfg) # Validator automatically configures memory
The set_longterm_memory validator runs automatically during instantiation, though you can verify activation by checking the role logs for "Long-term memory set for role" messages.
Adding Messages and Triggering Transfer
from metagpt.schema import UserMessage
# Add to memory - automatically handles short-term vs long-term routing
msg = UserMessage(content="Analyze the Q3 financial reports.")
role.rc.memory.add(msg)
# When buffer exceeds memory_k (200), oldest messages move to Chroma
Retrieving Augmented Context
# Retrieve recent context plus relevant historical memories
context = role.rc.memory.get(k=10)
# Returns up to 10 recent messages + up to 5 similar long-term items
The returned list combines recent short-term interactions with semantically relevant historical data, ready for LLM prompt construction.
Inspecting Stored Vectors (Debugging)
For troubleshooting or analysis, access the underlying Chroma collection directly:
# Access the RAG engine's retriever
engine = role.rc.memory.rag_engine
stored_ids = engine.retriever.collection.get(include=["ids"]).ids
print(f"Stored memory count: {len(stored_ids)}")
Summary
- Configuration: Enable long-term memory in
RoleZeroConfigviaenable_longterm_memory: trueinconfig/config2.yaml - Architecture: The
RoleZeroLongTermMemoryclass inmetagpt/memory/role_zero_memory.pymanages a hybrid system combining fast short-term buffers with Chroma-backed vector storage - Integration: The
set_longterm_memoryvalidator inmetagpt/roles/di/role_zero.pyautomatically wires the memory engine into the agent's runtime context - Operation: Messages flow through
add()for storage andget()for retrieval, with automatic overflow handling that preserves older data in the vector store while keeping recent context in memory - Persistence: Data survives process restarts through ChromaDB's persistent storage at the configured
persist_path
Frequently Asked Questions
How does RoleZero decide when to move messages from short-term to long-term memory?
The RoleZeroLongTermMemory class uses the _should_use_longterm_memory_for_add() method (lines 97-103 in role_zero_memory.py) to monitor buffer capacity. When the short-term list exceeds memory_k items (default 200), the system automatically transfers the oldest message to the Chroma vector store via _transfer_to_longterm_memory(). This overflow handling ensures the fast memory buffer remains bounded while preserving historical context for future retrieval.
What vector database does MetaGPT use for RoleZero's long-term memory?
MetaGPT uses ChromaDB as the underlying vector store for RoleZeroLongTermMemory. The implementation initializes a SimpleEngine (lines 39-69) that couples Chroma with optional LLM-based ranking. Vector embeddings persist to disk at the path specified by longterm_memory_persist_path (default .role_memory_data), enabling data survival across process restarts according to the MetaGPT source code.
Can I adjust how many historical memories are retrieved during context building?
Yes. The similarity_top_k parameter in RoleZeroConfig controls retrieval volume, defaulting to 5 items. When the get() method detects a full buffer and user requirement context (as determined by _should_use_longterm_memory_for_get), it queries the RAG engine for this number of semantically similar historical messages. Increasing this value provides more historical context but consumes additional tokens in the LLM prompt.
What happens if the Chroma vector store connection fails during runtime?
The memory implementation includes the @handle_exception decorator on methods interacting with external RAG components (lines 131-155). These decorators ensure that vector store failures emit log messages rather than raising exceptions that would crash the agent. The system degrades gracefully by continuing with short-term memory only until the external service recovers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →