# How the MemPalace 4-Layer Memory Stack (L0–L3) Works

> Explore the MemPalace 4 layer memory stack L0 L3. Learn how L1 L2 L3 layers provide rich context efficiently, keeping prompts small and improving retrieval. Essential for AI development.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: internals
- Published: 2026-06-06

---

**MemPalace organizes long-term memory into a four-layer hierarchy—L0 (Identity), L1 (Essential Story), L2 (On-Demand Retrieval), and L3 (Deep Search)—that keeps system prompts small while loading richer context only when needed.**

The `MemPalace/mempalace` project implements a lightweight, hierarchical **4-layer memory stack** that lets large language models access only the most relevant user memories for any given conversation. Defined in [`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py), the L0–L3 architecture keeps token usage predictable by loading static identity and essential summaries at start-up, while fetching deeper context on demand.

## What Is the MemPalace 4-Layer Memory Stack?

The **4-layer memory stack** is a tiered retrieval system wrapped by the `MemoryStack` class in [`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py). It serves a condensed system prompt to the LLM by default and expands context only when the conversation references specific projects, topics, or explicit search queries.

## L0 – Identity: The Static Baseline

**L0** reserves roughly **100 tokens** and is always loaded at start-up. It reads the user’s static identity file from `~/.mempalace/identity.txt` verbatim and renders that text as the first chunk of the system prompt. In [`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py), this behavior is implemented by `class Layer0` at line 34.

## L1 – Essential Story: The Wake-Up Summary

**L1** consumes approximately **500–800 tokens** and is always loaded immediately after L0, forming the **wake-up text**. It auto-generates a summary of the most salient memory drawers by scanning the ChromaDB collection, scoring each drawer by importance and recentness, and grouping results by room. This layer is implemented in [`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py) as `class Layer1` at line 76.

## L2 – On-Demand Retrieval: Filtered Context by Wing or Room

**L2** contributes roughly **200–500 tokens per request** and is fetched only when the user mentions a specific **wing** (project) or **room** (topic). It executes a filtered ChromaDB query through helpers in [`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py) and returns matching drawers without pulling the entire palace. The implementation lives in [`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py) as `class Layer2` at line 87.

## L3 – Deep Search: Full Semantic Lookup

**L3** has no fixed token cap and runs only when an explicit semantic search is requested. It performs a full-text semantic search across the entire ChromaDB collection—optionally filtered by wing or room—and formats results with similarity scores and source file hints. In [`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py), this is handled by `class Layer3` at line 47.

## How the MemoryStack API Unifies the Layers

The `MemoryStack` class provides a simple façade over the four layers:

- **`wake_up()`** → renders L0 + L1 (≈ 600–900 tokens).
- **`recall(wing, room)`** → invokes L2 to fetch on-demand drawers.
- **`search(query, wing, room)`** → runs L3 semantic search.

This design keeps the baseline system prompt around **1 k tokens**, ensuring low latency and predictable costs. The following example shows how to initialize the stack and call each API method:

```python
from mempalace.layers import MemoryStack

# Initialise the stack (uses the default palace path and identity file)

stack = MemoryStack()

# 1️⃣ Wake‑up: L0 + L1 (typical start of a conversation)

print("--- Wake‑up text ---")
print(stack.wake_up())          # ~600‑900 tokens

# 2️⃣ On‑demand retrieval for a specific project (wing)

print("\n--- On‑demand L2 for wing='my_app' ---")
print(stack.recall(wing="my_app"))   # returns ~200‑500 tokens of relevant drawers

# 3️⃣ Deep semantic search (L3) across the whole palace

print("\n--- Deep L3 search ---")
print(stack.search("pricing change"))   # returns the top 5 most similar drawers

```

The CLI entry point inside [`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py) (`if __name__ == "__main__":`) exposes the same commands—`wake-up`, `recall`, `search`, and `status`—for quick manual testing.

## Key Files in the Memory Stack Implementation

Several modules work together to support the L0–L3 architecture:

- **[`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py)** — Defines the four `Layer` classes and the `MemoryStack` façade.
- **[`mempalace/palace.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/palace.py)** — Provides `_get_collection` to access the ChromaDB collection.
- **[`mempalace/searcher.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/searcher.py)** — Builds filtered queries and extracts result fields for L2 and L3.
- **[`mempalace/config.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/config.py)** — Handles paths to the palace directory and identity file.
- **[`mempalace/backends/chroma.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/backends/chroma.py)** — Stores the drawers (text chunks) and metadata in ChromaDB.

## Summary

- MemPalace uses a **hierarchical 4-layer memory stack** in [`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py) to balance context richness with token efficiency.
- **L0** and **L1** form the always-loaded wake-up text (≈ 600–900 tokens) from static identity and auto-generated summaries.
- **L2** pulls filtered project or topic context on demand (≈ 200–500 tokens) via `recall()`.
- **L3** executes full semantic search across the entire palace when `search()` is called, with no fixed token limit.
- The `MemoryStack` class exposes `wake_up()`, `recall()`, and `search()` to keep the public API simple.

## Frequently Asked Questions

### What is the combined token budget for L0 and L1 in MemPalace?

The combined wake-up text from L0 and L1 typically stays between **600 and 900 tokens**. L0 contributes roughly 100 tokens from the static identity file, while L1 adds 500–800 tokens of summarized essential story.

### How does L2 on-demand retrieval differ from L3 deep search?

**L2** triggers only when a specific wing or room is referenced and runs a filtered ChromaDB query scoped to that metadata, returning about 200–500 tokens. **L3** requires an explicit semantic search request, scans the entire collection regardless of topic, and returns results ranked by similarity with no fixed token cap.

### Which source file defines the Layer classes and the MemoryStack façade?

All four layer classes and the `MemoryStack` façade are defined in **[`mempalace/layers.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/layers.py)**. This file also contains the CLI entry point for manual stack testing.

### Can users edit the identity file used by L0?

Yes. L0 reads the static identity file directly from **`~/.mempalace/identity.txt`**, so users can edit this file to update their baseline profile information. Changes are reflected the next time `MemoryStack.wake_up()` is invoked.