Building Agent Memory Systems from Scratch: The MemGPT Pattern Explained
Building agent memory systems from scratch requires implementing a two-tier architecture that treats the LLM prompt as fixed-size working memory (RAM) and external searchable stores as long-term memory (disk), with explicit paging mechanisms to move data between these tiers.
The rohitg00/ai-engineering-from-scratch repository provides a comprehensive 20-phase curriculum for building production-grade autonomous agents. Phase 14 ("Agent Engineering") introduces the core concepts for memory-augmented agents through Lesson 07 ("Virtual Context and Memory Paging"), which demonstrates the seminal MemGPT pattern using pure Python without external dependencies. This educational implementation exposes the underlying mechanics of virtual memory systems that power modern frameworks like Letta and Mem0.
The Virtual Memory Problem in LLM Agents
Expanding context windows alone cannot solve agent memory management. Infinite context leads to prompt dilution, increased latency, and lack of persistence across sessions. The MemGPT pattern, as implemented in phases/14-agent-engineering/07-memory-virtual-context-memgpt/code/main.py, solves this by adapting operating system virtual memory principles to LLM agents.
The architecture distinguishes between volatile working memory (the prompt buffer) and durable long-term memory (external databases). When the working memory fills, older content undergoes eviction rather than deletion, moving to archival storage where it remains searchable and retrievable.
Two-Tier Memory Architecture
MainContext: The Prompt as Working Memory
The MainContext class implements the fixed-size prompt buffer, analogous to RAM in computer architecture. Defined in the lesson's reference implementation, this dataclass maintains a max_messages limit (defaulting to 3 in the educational example) and automatically evicts older messages when the buffer exceeds capacity.
Key attributes include:
core: A dictionary storing persistent persona instructions and behavioral rulesmessages: A rolling list of recent conversation turnsevicted: A buffer holding popped messages before archival consolidation
When append() triggers eviction, messages move from messages to evicted, simulating the transition from active memory to swap space.
ArchivalStore: Persistent Long-Term Memory
The ArchivalStore class provides unlimited storage capacity using a BM25-inspired token overlap scoring mechanism. This in-memory implementation demonstrates how production vector databases function without requiring external dependencies.
The store manages ArchivalRecord objects containing:
rid: Unique record identifiertext: Semantic contenttags: Categorical metadata for filteringsession_idandturn_id: Temporal provenance tracking
Implementing the Memory System in Python
The following runnable implementation demonstrates the core MemGPT pattern. This code from phases/14-agent-engineering/07-memory-virtual-context-memgpt/code/main.py requires only Python's standard library:
from dataclasses import dataclass, field
from typing import Any
# -------------------------------------------------
# Core data structures (trimmed for brevity)
# -------------------------------------------------
@dataclass
class Message:
role: str
text: str
@dataclass
class MainContext:
core: dict[str, str] = field(default_factory=dict)
messages: list[Message] = field(default_factory=list)
max_messages: int = 3 # fixed-size prompt buffer
evicted: list[Message] = field(default_factory=list)
def append(self, role: str, text: str) -> None:
self.messages.append(Message(role, text))
while len(self.messages) > self.max_messages:
self.evicted.append(self.messages.pop(0))
def render(self) -> str:
parts = ["[core]"]
for k, v in sorted(self.core.items()):
parts.append(f" {k}: {v}")
parts.append("[messages]")
for m in self.messages:
parts.append(f" {m.role}: {m.text}")
return "\n".join(parts)
# -------------------------------------------------
# Simple archival store (token-overlap BM25 stub)
# -------------------------------------------------
@dataclass
class ArchivalRecord:
rid: str
text: str
tags: tuple[str, ...] = ()
session_id: str = "s0"
turn_id: int = 0
class ArchivalStore:
def __init__(self) -> None:
self._records: list[ArchivalRecord] = []
self._counter = 0
def insert(self, text: str, *, tags: tuple[str, ...] = ()) -> str:
self._counter += 1
rid = f"a{self._counter:03d}"
self._records.append(ArchivalRecord(rid, text, tags))
return rid
def search(self, query: str, top_k: int = 3) -> list[ArchivalRecord]:
q = set(query.lower().split())
scored = []
for r in self._records:
overlap = len(q & set(r.text.lower().split()))
if overlap:
score = overlap / (len(q) + len(r.text.split()) - overlap)
scored.append((score, r))
scored.sort(key=lambda x: -x[0])
return [r for _, r in scored[:top_k]]
Memory Tools and the Interrupt Pattern
Memory operations expose system call-like interfaces through the MemoryTools class. These tools implement the interrupt pattern: the agent halts execution to request data, the runtime fetches from external storage, and results splice back into the next turn's context.
Core Memory Operations
core_memory_append(section, text): Appends content to persistent sections of the prompt that survive message eviction. This implements procedural memory (behavioral rules) and semantic memory (user facts) that remain constantly available.
Archival Memory Operations
archival_memory_insert(text, tags): Persists facts to long-term storage with optional metadata for categorical retrieval. Returns a record ID for future reference.
archival_memory_search(query, top_k): Retrieves relevant records using BM25-style token overlap scoring. This paging operation simulates disk I/O, fetching external context and injecting it into the working memory for the next reasoning step.
From Educational Prototype to Production
The rohitg00/ai-engineering-from-scratch implementation generalizes to production systems through specific architectural extensions:
- Letta: Adds a third recall tier between working and archival memory
- Mem0: Fuses vector, key-value, and graph stores for multi-modal retrieval
- Zep: Implements temporal knowledge graphs for conversation history
- OpenAI Assistants: Embeds similar paging mechanisms via the
file_searchtool API
The educational code deliberately uses Python's standard library to expose mechanics hidden by frameworks like LangChain or LlamaIndex. The AGENTS.md file in the repository root provides curriculum context, while site/data.js catalogs lesson dependencies.
Summary
- Treat prompts as fixed buffers: Configure
max_messageslimits and implement eviction policies to prevent context overflow. - Provide bidirectional paging tools: Implement explicit read/write functions (
archival_memory_search,archival_memory_insert) rather than implicit retrieval. - Use interrupt-driven execution: Structure agents to pause for memory operations, mimicking system calls that fetch external data before continuing reasoning.
- Maintain memory hygiene: Consolidate archival stores periodically to prevent knowledge rot and poisoning.
- Start framework-free: The
phases/14-agent-engineering/07-memory-virtual-context-memgpt/implementation demonstrates core patterns without external dependencies, making mechanics inspectable.
Frequently Asked Questions
What is the MemGPT pattern in agent memory systems?
The MemGPT pattern treats the LLM prompt as operating system working memory (RAM) and external databases as long-term memory (disk). It implements explicit paging tools that allow agents to move data between these tiers through function calls, solving the fixed context window limitation without losing information. According to the rohitg00/ai-engineering-from-scratch source code, this architecture prevents prompt dilution while maintaining unlimited storage capacity through the ArchivalStore class.
How does message eviction work in the MainContext implementation?
When MainContext.append() exceeds the max_messages threshold, the oldest messages undergo FIFO eviction from the active messages list to a separate evicted buffer. This mimics how operating systems swap memory pages to disk. In production systems, evicted content typically undergoes summarization or compression before archival storage, though the educational implementation preserves full text in phases/14-agent-engineering/07-memory-virtual-context-memgpt/code/main.py to demonstrate the mechanics clearly.
Why use BM25 scoring instead of vector embeddings in the reference implementation?
The ArchivalStore.search() method uses token-overlap BM25 scoring to remain dependency-free and deterministic. Vector embeddings require external libraries (sentence-transformers, OpenAI API) that obscure the underlying retrieval mechanics. The rohitg00/ai-engineering-from-scratch repository maintains pure Python implementations to teach concepts without framework opacity, though production deployments should replace this with vector stores like Pinecone or Weaviate for semantic similarity search.
How does this architecture scale to production multi-agent systems?
The two-tier pattern generalizes through three extensions: adding a recall tier (Letta) for frequently accessed memories, implementing multi-store fusion (Mem0) combining vector, KV, and graph databases, and deploying sleep-time compute agents that consolidate and clean archival memory during idle periods. The interrupt pattern remains consistent—agents invoke memory.search tools that the runtime executes before continuing the reasoning loop, whether in single agents or orchestrated swarms.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →