Building Agent Memory Systems: A Complete MemGPT-Style Hybrid Memory Implementation

The rohitg00/ai-engineering-from-scratch repository provides a production-grade implementation of a hybrid memory architecture that combines vector similarity, key-value lookup, and graph-based reasoning behind a unified interface with configurable fusion scoring.

Agent memory systems enable large language models to persist context across sessions and recall facts with semantic, exact-match, and relational precision. The reference implementation in rohitg00/ai-engineering-from-scratch demonstrates a MemGPT-style hybrid memory pattern that orchestrates three complementary storage backends behind a single add() and search() API. This architecture prevents data leakage across users while supporting temporal reasoning and importance-based recall.

Architectural Overview of the Three-Store System

The implementation in phases/14-agent-engineering/09-hybrid-memory-mem0/code/main.py defines a modular architecture where each storage backend handles a specific retrieval pattern. The Mem0 class acts as a facade that coordinates writes and fuses retrieval results using a weighted scoring function.

VectorStore for Semantic Similarity

The VectorStore class (lines 26-47) provides semantic search capabilities using a token-overlap metric as a lightweight stand-in for dense embeddings. In production deployments, this component swaps seamlessly for vector databases like Qdrant or Pinecone while maintaining the same interface.

KVStore for Exact-Match Facts

The KVStore class (lines 56-68) maintains an O(1) lookup table keyed by (user_id, fact_type, entity) tuples. This store handles factual precision for attributes like locations, preferences, or account statuses that require exact retrieval rather than approximate similarity.

GraphStore for Relationship Reasoning

The GraphStore class (lines 71-95) maintains typed edges between entities and supports temporal invalidation. When add_edge() detects contradictory information (lines 84-88), it flags the existing edge as valid=False rather than deleting it, enabling "as-of-time" queries and contradiction detection without destroying historical context.

Fusion Scoring and Retrieval Logic

The hybrid retrieval mechanism in Mem0.search() (lines 45-53) aggregates candidates from all three stores and ranks them using a configurable weighted sum:


score = w_relevance × rel + w_importance × imp + w_recency × rec

This fusion scoring approach ensures that semantically relevant results, critically important facts, and recent updates all contribute to the final ranking. Weights are configurable via Mem0Config, allowing agents to prioritize recency for conversational contexts or importance for long-term knowledge retention.

Scope Isolation and Security

The implementation enforces a scope taxonomy that segregates memory by user, session, and agent (lines 13-15 and 34-44). When calling Mem0.add() or Mem0.search(), the system automatically injects scope constraints, ensuring that queries return only records matching the specified user_id and session_id. This prevents cross-user data leakage at the storage layer rather than relying solely on application-level filtering.

Implementing the Hybrid Memory System

The following examples demonstrate the complete workflow using classes defined in phases/14-agent-engineering/09-hybrid-memory-mem0/code/main.py.

Initializing the Memory Facade

from code.main import Mem0

mem = Mem0()  # Uses default fusion weights

Writing to All Three Stores

The add() method simultaneously populates the vector, key-value, and graph stores:


# Write semantic, factual, and relational data in one call

mem.add(
    "ava lives in Berlin",
    user_id="ava",
    session_id="s001",
    importance=0.6,
    kv_triples=(("city", "Berlin"),),
    graph_triples=(("ava", "lives_in", "Berlin"),)
)

# Temporal update invalidates previous graph edge

mem.add(
    "ava moved to Lisbon last month",
    user_id="ava",
    session_id="s002",
    importance=0.8,
    kv_triples=(("city", "Lisbon"),),
    graph_triples=(("ava", "lives_in", "Lisbon"),)
)

Fused Retrieval Across Backends

The search() method queries all stores and returns fused results:


# Query touches vector, KV, and graph stores automatically

for score, rec in mem.search(
    "where does ava live", user_id="ava", top_k=3
):
    print(f"{score:.3f} | {rec.rid} | {rec.text}")

Debugging Individual Stores

For analysis or specialized queries, access each store directly:


# Vector-only semantic search

for s, r in mem.vector.search("writing style preferences", top_k=3):
    print(s, r.rid, r.text)

# KV-only factual lookup

for r in mem.kv.by_user("ava"):
    print(r.rid, r.text)

# Graph traversal with validity checking

for e in mem.graph.neighbors("ava", valid_only=False):
    status = "VALID" if e.valid else "INVALID"
    print(f"{status}: {e.subject} --{e.relation}--> {e.obj}")

Verifying Scope Isolation


# Confirm Bob's data is isolated from Ava's queries

hits = mem.search("refund invoice", user_id="ava")
assert len(hits) == 0  # No cross-user leakage

Production Deployment Considerations

While the repository uses in-memory data structures and token-overlap similarity for educational clarity, the architecture supports enterprise scaling through dependency injection. Replace the VectorStore with embedding models from Sentence-Transformers or OpenAI, swap the KVStore with Redis or PostgreSQL, and replace GraphStore with Neo4j or Amazon Neptune. The Mem0 facade maintains the same interface, allowing prototypes to evolve into production agent memory systems without rewriting application logic.

Summary

  • The three-store architecture (vector, KV, graph) handles semantic, exact-match, and relational queries through a unified interface defined in code/main.py.
  • Fusion scoring combines relevance, importance, and recency using configurable weights defined in Mem0Config.
  • Scope isolation at the storage layer prevents cross-user data leakage via automatic user_id, session_id, and agent filtering in Mem0.add() and Mem0.search().
  • Temporal invalidation in GraphStore.add_edge() (lines 84-88) preserves historical facts while flagging contradictions, enabling "as-of-time" reasoning.
  • The implementation in rohitg00/ai-engineering-from-scratch provides drop-in classes that scale from educational prototypes to production deployments by swapping storage backends.

Frequently Asked Questions

What is the difference between MemGPT and this hybrid memory implementation?

MemGPT introduced the concept of virtual-context memory management for language models, while this implementation follows the Mem0 design pattern (Chhikara et al., 2025) by specifically combining three distinct storage backends—vector, key-value, and graph—behind a unified retrieval interface with weighted fusion scoring. The Mem0 approach explicitly separates semantic recall from factual lookup and relational reasoning.

How does the fusion scoring algorithm work?

The algorithm defined in Mem0.search() (lines 45-53) calculates a composite score for each memory record using the formula score = w_relevance·rel + w_importance·imp + w_recency·rec, where weights are configurable parameters. This ensures retrieved memories balance semantic similarity with factual priority and temporal freshness, allowing agents to surface the most contextually appropriate information.

Can this system handle contradictory information over time?

Yes. The GraphStore class implements temporal invalidation in add_edge() (lines 84-88) by setting valid=False on superseded edges rather than deleting them. This preserves the historical record while ensuring current queries receive only valid relationships, enabling agents to reason about when facts changed and detect contradictions in user statements.

How do I prevent user data from leaking between different agent sessions?

The system enforces scope isolation through mandatory user_id and session_id parameters in Mem0.add() and Mem0.search() (lines 13-15, 34-44). These scopes act as logical partitions at the storage layer, ensuring queries return only records matching the specified user and session contexts, preventing accidental exposure of one user's data to another.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →