Building Agent Memory Systems: A Complete MemGPT-Style Hybrid Memory Implementation
The rohitg00/ai-engineering-from-scratch repository provides a production-grade implementation of a hybrid memory architecture that combines vector similarity, key-value lookup, and graph-based reasoning behind a unified interface with configurable fusion scoring.
Agent memory systems enable large language models to persist context across sessions and recall facts with semantic, exact-match, and relational precision. The reference implementation in rohitg00/ai-engineering-from-scratch demonstrates a MemGPT-style hybrid memory pattern that orchestrates three complementary storage backends behind a single add() and search() API. This architecture prevents data leakage across users while supporting temporal reasoning and importance-based recall.
Architectural Overview of the Three-Store System
The implementation in phases/14-agent-engineering/09-hybrid-memory-mem0/code/main.py defines a modular architecture where each storage backend handles a specific retrieval pattern. The Mem0 class acts as a facade that coordinates writes and fuses retrieval results using a weighted scoring function.
VectorStore for Semantic Similarity
The VectorStore class (lines 26-47) provides semantic search capabilities using a token-overlap metric as a lightweight stand-in for dense embeddings. In production deployments, this component swaps seamlessly for vector databases like Qdrant or Pinecone while maintaining the same interface.
KVStore for Exact-Match Facts
The KVStore class (lines 56-68) maintains an O(1) lookup table keyed by (user_id, fact_type, entity) tuples. This store handles factual precision for attributes like locations, preferences, or account statuses that require exact retrieval rather than approximate similarity.
GraphStore for Relationship Reasoning
The GraphStore class (lines 71-95) maintains typed edges between entities and supports temporal invalidation. When add_edge() detects contradictory information (lines 84-88), it flags the existing edge as valid=False rather than deleting it, enabling "as-of-time" queries and contradiction detection without destroying historical context.
Fusion Scoring and Retrieval Logic
The hybrid retrieval mechanism in Mem0.search() (lines 45-53) aggregates candidates from all three stores and ranks them using a configurable weighted sum:
score = w_relevance × rel + w_importance × imp + w_recency × rec
This fusion scoring approach ensures that semantically relevant results, critically important facts, and recent updates all contribute to the final ranking. Weights are configurable via Mem0Config, allowing agents to prioritize recency for conversational contexts or importance for long-term knowledge retention.
Scope Isolation and Security
The implementation enforces a scope taxonomy that segregates memory by user, session, and agent (lines 13-15 and 34-44). When calling Mem0.add() or Mem0.search(), the system automatically injects scope constraints, ensuring that queries return only records matching the specified user_id and session_id. This prevents cross-user data leakage at the storage layer rather than relying solely on application-level filtering.
Implementing the Hybrid Memory System
The following examples demonstrate the complete workflow using classes defined in phases/14-agent-engineering/09-hybrid-memory-mem0/code/main.py.
Initializing the Memory Facade
from code.main import Mem0
mem = Mem0() # Uses default fusion weights
Writing to All Three Stores
The add() method simultaneously populates the vector, key-value, and graph stores:
# Write semantic, factual, and relational data in one call
mem.add(
"ava lives in Berlin",
user_id="ava",
session_id="s001",
importance=0.6,
kv_triples=(("city", "Berlin"),),
graph_triples=(("ava", "lives_in", "Berlin"),)
)
# Temporal update invalidates previous graph edge
mem.add(
"ava moved to Lisbon last month",
user_id="ava",
session_id="s002",
importance=0.8,
kv_triples=(("city", "Lisbon"),),
graph_triples=(("ava", "lives_in", "Lisbon"),)
)
Fused Retrieval Across Backends
The search() method queries all stores and returns fused results:
# Query touches vector, KV, and graph stores automatically
for score, rec in mem.search(
"where does ava live", user_id="ava", top_k=3
):
print(f"{score:.3f} | {rec.rid} | {rec.text}")
Debugging Individual Stores
For analysis or specialized queries, access each store directly:
# Vector-only semantic search
for s, r in mem.vector.search("writing style preferences", top_k=3):
print(s, r.rid, r.text)
# KV-only factual lookup
for r in mem.kv.by_user("ava"):
print(r.rid, r.text)
# Graph traversal with validity checking
for e in mem.graph.neighbors("ava", valid_only=False):
status = "VALID" if e.valid else "INVALID"
print(f"{status}: {e.subject} --{e.relation}--> {e.obj}")
Verifying Scope Isolation
# Confirm Bob's data is isolated from Ava's queries
hits = mem.search("refund invoice", user_id="ava")
assert len(hits) == 0 # No cross-user leakage
Production Deployment Considerations
While the repository uses in-memory data structures and token-overlap similarity for educational clarity, the architecture supports enterprise scaling through dependency injection. Replace the VectorStore with embedding models from Sentence-Transformers or OpenAI, swap the KVStore with Redis or PostgreSQL, and replace GraphStore with Neo4j or Amazon Neptune. The Mem0 facade maintains the same interface, allowing prototypes to evolve into production agent memory systems without rewriting application logic.
Summary
- The three-store architecture (vector, KV, graph) handles semantic, exact-match, and relational queries through a unified interface defined in
code/main.py. - Fusion scoring combines relevance, importance, and recency using configurable weights defined in
Mem0Config. - Scope isolation at the storage layer prevents cross-user data leakage via automatic
user_id,session_id, andagentfiltering inMem0.add()andMem0.search(). - Temporal invalidation in
GraphStore.add_edge()(lines 84-88) preserves historical facts while flagging contradictions, enabling "as-of-time" reasoning. - The implementation in
rohitg00/ai-engineering-from-scratchprovides drop-in classes that scale from educational prototypes to production deployments by swapping storage backends.
Frequently Asked Questions
What is the difference between MemGPT and this hybrid memory implementation?
MemGPT introduced the concept of virtual-context memory management for language models, while this implementation follows the Mem0 design pattern (Chhikara et al., 2025) by specifically combining three distinct storage backends—vector, key-value, and graph—behind a unified retrieval interface with weighted fusion scoring. The Mem0 approach explicitly separates semantic recall from factual lookup and relational reasoning.
How does the fusion scoring algorithm work?
The algorithm defined in Mem0.search() (lines 45-53) calculates a composite score for each memory record using the formula score = w_relevance·rel + w_importance·imp + w_recency·rec, where weights are configurable parameters. This ensures retrieved memories balance semantic similarity with factual priority and temporal freshness, allowing agents to surface the most contextually appropriate information.
Can this system handle contradictory information over time?
Yes. The GraphStore class implements temporal invalidation in add_edge() (lines 84-88) by setting valid=False on superseded edges rather than deleting them. This preserves the historical record while ensuring current queries receive only valid relationships, enabling agents to reason about when facts changed and detect contradictions in user statements.
How do I prevent user data from leaking between different agent sessions?
The system enforces scope isolation through mandatory user_id and session_id parameters in Mem0.add() and Mem0.search() (lines 13-15, 34-44). These scopes act as logical partitions at the storage layer, ensuring queries return only records matching the specified user and session contexts, preventing accidental exposure of one user's data to another.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →