How the MemPalace 4-Layer Memory Stack (L0–L3) Works

MemPalace organizes long-term memory into a four-layer hierarchy—L0 (Identity), L1 (Essential Story), L2 (On-Demand Retrieval), and L3 (Deep Search)—that keeps system prompts small while loading richer context only when needed.

The MemPalace/mempalace project implements a lightweight, hierarchical 4-layer memory stack that lets large language models access only the most relevant user memories for any given conversation. Defined in mempalace/layers.py, the L0–L3 architecture keeps token usage predictable by loading static identity and essential summaries at start-up, while fetching deeper context on demand.

What Is the MemPalace 4-Layer Memory Stack?

The 4-layer memory stack is a tiered retrieval system wrapped by the MemoryStack class in mempalace/layers.py. It serves a condensed system prompt to the LLM by default and expands context only when the conversation references specific projects, topics, or explicit search queries.

L0 – Identity: The Static Baseline

L0 reserves roughly 100 tokens and is always loaded at start-up. It reads the user’s static identity file from ~/.mempalace/identity.txt verbatim and renders that text as the first chunk of the system prompt. In mempalace/layers.py, this behavior is implemented by class Layer0 at line 34.

L1 – Essential Story: The Wake-Up Summary

L1 consumes approximately 500–800 tokens and is always loaded immediately after L0, forming the wake-up text. It auto-generates a summary of the most salient memory drawers by scanning the ChromaDB collection, scoring each drawer by importance and recentness, and grouping results by room. This layer is implemented in mempalace/layers.py as class Layer1 at line 76.

L2 – On-Demand Retrieval: Filtered Context by Wing or Room

L2 contributes roughly 200–500 tokens per request and is fetched only when the user mentions a specific wing (project) or room (topic). It executes a filtered ChromaDB query through helpers in mempalace/searcher.py and returns matching drawers without pulling the entire palace. The implementation lives in mempalace/layers.py as class Layer2 at line 87.

L3 – Deep Search: Full Semantic Lookup

L3 has no fixed token cap and runs only when an explicit semantic search is requested. It performs a full-text semantic search across the entire ChromaDB collection—optionally filtered by wing or room—and formats results with similarity scores and source file hints. In mempalace/layers.py, this is handled by class Layer3 at line 47.

How the MemoryStack API Unifies the Layers

The MemoryStack class provides a simple façade over the four layers:

  • wake_up() → renders L0 + L1 (≈ 600–900 tokens).
  • recall(wing, room) → invokes L2 to fetch on-demand drawers.
  • search(query, wing, room) → runs L3 semantic search.

This design keeps the baseline system prompt around 1 k tokens, ensuring low latency and predictable costs. The following example shows how to initialize the stack and call each API method:

from mempalace.layers import MemoryStack

# Initialise the stack (uses the default palace path and identity file)

stack = MemoryStack()

# 1️⃣ Wake‑up: L0 + L1 (typical start of a conversation)

print("--- Wake‑up text ---")
print(stack.wake_up())          # ~600‑900 tokens

# 2️⃣ On‑demand retrieval for a specific project (wing)

print("\n--- On‑demand L2 for wing='my_app' ---")
print(stack.recall(wing="my_app"))   # returns ~200‑500 tokens of relevant drawers

# 3️⃣ Deep semantic search (L3) across the whole palace

print("\n--- Deep L3 search ---")
print(stack.search("pricing change"))   # returns the top 5 most similar drawers

The CLI entry point inside mempalace/layers.py (if __name__ == "__main__":) exposes the same commands—wake-up, recall, search, and status—for quick manual testing.

Key Files in the Memory Stack Implementation

Several modules work together to support the L0–L3 architecture:

Summary

  • MemPalace uses a hierarchical 4-layer memory stack in mempalace/layers.py to balance context richness with token efficiency.
  • L0 and L1 form the always-loaded wake-up text (≈ 600–900 tokens) from static identity and auto-generated summaries.
  • L2 pulls filtered project or topic context on demand (≈ 200–500 tokens) via recall().
  • L3 executes full semantic search across the entire palace when search() is called, with no fixed token limit.
  • The MemoryStack class exposes wake_up(), recall(), and search() to keep the public API simple.

Frequently Asked Questions

What is the combined token budget for L0 and L1 in MemPalace?

The combined wake-up text from L0 and L1 typically stays between 600 and 900 tokens. L0 contributes roughly 100 tokens from the static identity file, while L1 adds 500–800 tokens of summarized essential story.

L2 triggers only when a specific wing or room is referenced and runs a filtered ChromaDB query scoped to that metadata, returning about 200–500 tokens. L3 requires an explicit semantic search request, scans the entire collection regardless of topic, and returns results ranked by similarity with no fixed token cap.

Which source file defines the Layer classes and the MemoryStack façade?

All four layer classes and the MemoryStack façade are defined in mempalace/layers.py. This file also contains the CLI entry point for manual stack testing.

Can users edit the identity file used by L0?

Yes. L0 reads the static identity file directly from ~/.mempalace/identity.txt, so users can edit this file to update their baseline profile information. Changes are reflected the next time MemoryStack.wake_up() is invoked.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →