# Building Agent Memory Systems: A Complete MemGPT-Style Hybrid Memory Implementation

> Build advanced agent memory systems with a MemGPT-style hybrid approach. Explore a production-grade implementation combining vector search, key-value lookup, and graph reasoning.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-26

---

**The rohitg00/ai-engineering-from-scratch repository provides a production-grade implementation of a hybrid memory architecture that combines vector similarity, key-value lookup, and graph-based reasoning behind a unified interface with configurable fusion scoring.**

Agent memory systems enable large language models to persist context across sessions and recall facts with semantic, exact-match, and relational precision. The reference implementation in `rohitg00/ai-engineering-from-scratch` demonstrates a **MemGPT-style hybrid memory** pattern that orchestrates three complementary storage backends behind a single `add()` and `search()` API. This architecture prevents data leakage across users while supporting temporal reasoning and importance-based recall.

## Architectural Overview of the Three-Store System

The implementation in [`phases/14-agent-engineering/09-hybrid-memory-mem0/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/09-hybrid-memory-mem0/code/main.py) defines a modular architecture where each storage backend handles a specific retrieval pattern. The `Mem0` class acts as a facade that coordinates writes and fuses retrieval results using a weighted scoring function.

### VectorStore for Semantic Similarity

The `VectorStore` class (lines 26-47) provides semantic search capabilities using a token-overlap metric as a lightweight stand-in for dense embeddings. In production deployments, this component swaps seamlessly for vector databases like Qdrant or Pinecone while maintaining the same interface.

### KVStore for Exact-Match Facts

The `KVStore` class (lines 56-68) maintains an O(1) lookup table keyed by `(user_id, fact_type, entity)` tuples. This store handles factual precision for attributes like locations, preferences, or account statuses that require exact retrieval rather than approximate similarity.

### GraphStore for Relationship Reasoning

The `GraphStore` class (lines 71-95) maintains typed edges between entities and supports temporal invalidation. When `add_edge()` detects contradictory information (lines 84-88), it flags the existing edge as `valid=False` rather than deleting it, enabling "as-of-time" queries and contradiction detection without destroying historical context.

## Fusion Scoring and Retrieval Logic

The hybrid retrieval mechanism in `Mem0.search()` (lines 45-53) aggregates candidates from all three stores and ranks them using a configurable weighted sum:

```

score = w_relevance × rel + w_importance × imp + w_recency × rec

```

This **fusion scoring** approach ensures that semantically relevant results, critically important facts, and recent updates all contribute to the final ranking. Weights are configurable via `Mem0Config`, allowing agents to prioritize recency for conversational contexts or importance for long-term knowledge retention.

## Scope Isolation and Security

The implementation enforces a **scope taxonomy** that segregates memory by `user`, `session`, and `agent` (lines 13-15 and 34-44). When calling `Mem0.add()` or `Mem0.search()`, the system automatically injects scope constraints, ensuring that queries return only records matching the specified `user_id` and `session_id`. This prevents cross-user data leakage at the storage layer rather than relying solely on application-level filtering.

## Implementing the Hybrid Memory System

The following examples demonstrate the complete workflow using classes defined in [`phases/14-agent-engineering/09-hybrid-memory-mem0/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/09-hybrid-memory-mem0/code/main.py).

### Initializing the Memory Facade

```python
from code.main import Mem0

mem = Mem0()  # Uses default fusion weights

```

### Writing to All Three Stores

The `add()` method simultaneously populates the vector, key-value, and graph stores:

```python

# Write semantic, factual, and relational data in one call

mem.add(
    "ava lives in Berlin",
    user_id="ava",
    session_id="s001",
    importance=0.6,
    kv_triples=(("city", "Berlin"),),
    graph_triples=(("ava", "lives_in", "Berlin"),)
)

# Temporal update invalidates previous graph edge

mem.add(
    "ava moved to Lisbon last month",
    user_id="ava",
    session_id="s002",
    importance=0.8,
    kv_triples=(("city", "Lisbon"),),
    graph_triples=(("ava", "lives_in", "Lisbon"),)
)

```

### Fused Retrieval Across Backends

The `search()` method queries all stores and returns fused results:

```python

# Query touches vector, KV, and graph stores automatically

for score, rec in mem.search(
    "where does ava live", user_id="ava", top_k=3
):
    print(f"{score:.3f} | {rec.rid} | {rec.text}")

```

### Debugging Individual Stores

For analysis or specialized queries, access each store directly:

```python

# Vector-only semantic search

for s, r in mem.vector.search("writing style preferences", top_k=3):
    print(s, r.rid, r.text)

# KV-only factual lookup

for r in mem.kv.by_user("ava"):
    print(r.rid, r.text)

# Graph traversal with validity checking

for e in mem.graph.neighbors("ava", valid_only=False):
    status = "VALID" if e.valid else "INVALID"
    print(f"{status}: {e.subject} --{e.relation}--> {e.obj}")

```

### Verifying Scope Isolation

```python

# Confirm Bob's data is isolated from Ava's queries

hits = mem.search("refund invoice", user_id="ava")
assert len(hits) == 0  # No cross-user leakage

```

## Production Deployment Considerations

While the repository uses in-memory data structures and token-overlap similarity for educational clarity, the architecture supports enterprise scaling through dependency injection. Replace the `VectorStore` with embedding models from Sentence-Transformers or OpenAI, swap the `KVStore` with Redis or PostgreSQL, and replace `GraphStore` with Neo4j or Amazon Neptune. The `Mem0` facade maintains the same interface, allowing prototypes to evolve into production **agent memory systems** without rewriting application logic.

## Summary

- The **three-store architecture** (vector, KV, graph) handles semantic, exact-match, and relational queries through a unified interface defined in [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py).
- **Fusion scoring** combines relevance, importance, and recency using configurable weights defined in `Mem0Config`.
- **Scope isolation** at the storage layer prevents cross-user data leakage via automatic `user_id`, `session_id`, and `agent` filtering in `Mem0.add()` and `Mem0.search()`.
- **Temporal invalidation** in `GraphStore.add_edge()` (lines 84-88) preserves historical facts while flagging contradictions, enabling "as-of-time" reasoning.
- The implementation in `rohitg00/ai-engineering-from-scratch` provides drop-in classes that scale from educational prototypes to production deployments by swapping storage backends.

## Frequently Asked Questions

### What is the difference between MemGPT and this hybrid memory implementation?

MemGPT introduced the concept of virtual-context memory management for language models, while this implementation follows the Mem0 design pattern (Chhikara et al., 2025) by specifically combining three distinct storage backends—vector, key-value, and graph—behind a unified retrieval interface with weighted fusion scoring. The Mem0 approach explicitly separates semantic recall from factual lookup and relational reasoning.

### How does the fusion scoring algorithm work?

The algorithm defined in `Mem0.search()` (lines 45-53) calculates a composite score for each memory record using the formula `score = w_relevance·rel + w_importance·imp + w_recency·rec`, where weights are configurable parameters. This ensures retrieved memories balance semantic similarity with factual priority and temporal freshness, allowing agents to surface the most contextually appropriate information.

### Can this system handle contradictory information over time?

Yes. The `GraphStore` class implements temporal invalidation in `add_edge()` (lines 84-88) by setting `valid=False` on superseded edges rather than deleting them. This preserves the historical record while ensuring current queries receive only valid relationships, enabling agents to reason about when facts changed and detect contradictions in user statements.

### How do I prevent user data from leaking between different agent sessions?

The system enforces scope isolation through mandatory `user_id` and `session_id` parameters in `Mem0.add()` and `Mem0.search()` (lines 13-15, 34-44). These scopes act as logical partitions at the storage layer, ensuring queries return only records matching the specified user and session contexts, preventing accidental exposure of one user's data to another.