Claude-Mem 3-Layer Progressive Disclosure Workflow: Implementation Guide

Claude-Mem implements a 3-layer progressive disclosure workflow that reduces token consumption by up to 90% through lightweight search indexing, contextual timeline expansion, and batched full-payload retrieval.

The thedotmack/claude-mem repository introduces a sophisticated memory retrieval pattern designed specifically for large language model contexts. This Claude-Mem 3-layer progressive disclosure workflow ensures that expensive observation payloads are only fetched after precise filtering, keeping interactions fast and cost-efficient.

The Three-Layer Architecture

The workflow operates as a funnel, progressively revealing more data only when necessary. Each layer corresponds to a specific MCP tool registered in src/servers/mcp-server.ts.

Layer 1: Search (Lightweight Index)

The search tool performs broad queries against the memory store and returns a minimal index. Each result contains only essential metadata—observation IDs, titles, and timestamps—consuming approximately 50–100 tokens per result rather than the full 500–1000 token payload.

In src/servers/mcp-server.ts (lines 71–76), the implementation forwards the request to the worker's /api/search endpoint:

// src/servers/mcp-server.ts
server.setRequestHandler(CallToolRequestSchema, async (request) => {
  if (request.params.name === "search") {
    const response = await fetch(`${WORKER_URL}/api/search`, {
      method: "POST",
      body: JSON.stringify(request.params.arguments),
    });
    // Process and return lightweight index
  }
});

Layer 2: Timeline (Contextual Expansion)

Once the model identifies a relevant observation ID from the search results, the timeline tool expands the context chronologically. This layer retrieves "before" and "after" entries relative to an anchor point, providing temporal context without requiring multiple full-payload fetches.

The implementation (lines 80–85 in src/servers/mcp-server.ts) calls the worker's /api/timeline endpoint:

if (request.params.name === "timeline") {
  const { anchor, depth_before = 3, depth_after = 3 } = request.params.arguments;
  const response = await fetch(`${WORKER_URL}/api/timeline`, {
    method: "POST",
    body: JSON.stringify({ anchor, depth_before, depth_after }),
  });
}

Layer 3: Fetch Observations (Full Payload)

Only after the model has filtered the timeline results to specific observation IDs does the workflow retrieve full payloads. The get_observations tool batches these requests, fetching complete observation data exclusively for the selected identifiers.

This final layer (lines 92–99 in src/servers/mcp-server.ts) posts to /api/observations/batch:

if (request.params.name === "get_observations") {
  const { ids } = request.params.arguments;
  const response = await fetch(`${WORKER_URL}/api/observations/batch`, {
    method: "POST",
    body: JSON.stringify({ ids }),
  });
  // Return full observation payloads
}

Enforcement Through the MCP Protocol

The workflow is not merely documented but enforced through the MCP server's __IMPORTANT tool (lines 59–66 in src/servers/mcp-server.ts). When invoked, this tool returns the strict workflow rules:

3-LAYER WORKFLOW (ALWAYS FOLLOW):

  1. search(query) → index with IDs (~50-100 tokens/result)
  2. timeline(anchor=ID) → context around interesting results
  3. get_observations([IDs]) → fetch full details ONLY for filtered IDs Never fetch full details without filtering first – up to 10× token savings.

This programmatic enforcement ensures that Claude-Code instances interacting with the repository automatically adhere to the progressive disclosure pattern.

Token Economy and Performance Benefits

The Claude-Mem 3-layer progressive disclosure workflow delivers measurable efficiency gains:

  • Token Reduction: By filtering through search (50–100 tokens) and timeline layers before fetching full observations (500–1000 tokens each), the system achieves up to 90% token savings on large memory stores.
  • Latency Optimization: Index queries return in milliseconds, allowing rapid iteration on search results before committing to expensive full-payload transfers.
  • Context Safety: The explicit selection mechanism in layer 2 prevents accidental inclusion of irrelevant observations in the final context window.

Code Example: Executing the Workflow

The following TypeScript implementation demonstrates the complete 3-layer pattern using the MCP tools:

import { callTool } from '@claude-mem/client'

// Layer 1: Search - Retrieve lightweight index
const searchResult = await callTool('search', {
  query: 'authentication middleware configuration',
  limit: 20,
  project: 'claude-mem'
})

// Select promising candidate from index
const targetId = searchResult.table[0].id

// Layer 2: Timeline - Expand chronological context
const timeline = await callTool('timeline', {
  anchor: targetId,
  depth_before: 3,
  depth_after: 3
})

// Filter timeline to relevant entries
const relevantIds = timeline.entries
  .filter(e => e.title.includes('middleware'))
  .map(e => e.id)

// Layer 3: Fetch Observations - Retrieve full payloads
const observations = await callTool('get_observations', {
  ids: relevantIds
})

// Process complete observation data
console.log(observations)

Summary

  • Claude-Mem's 3-layer progressive disclosure workflow minimizes token consumption through staged data retrieval.
  • Layer 1 (search) provides lightweight indices (50–100 tokens) via src/servers/mcp-server.ts lines 71–76.
  • Layer 2 (timeline) expands chronological context around selected anchors via lines 80–85.
  • Layer 3 (get_observations) fetches full payloads only for filtered IDs via lines 92–99.
  • The __IMPORTANT tool (lines 59–66) programmatically enforces this workflow, ensuring up to 10× token savings.

Frequently Asked Questions

How does the 3-layer workflow reduce token costs?

The workflow reduces token costs by deferring full payload retrieval until the final step. Layer 1 returns only metadata (50–100 tokens per result) rather than complete observations (500–1000 tokens). By filtering through Layer 2 before invoking Layer 3, the system typically fetches full data for fewer than 10% of initial matches, achieving up to 90% token reduction.

What happens if I skip the timeline layer and fetch observations directly?

Skipping the timeline layer violates the protocol enforced by the __IMPORTANT tool in src/servers/mcp-server.ts (lines 59–66). While technically possible, this bypasses the contextual filtering that ensures relevance. Fetching observations directly after search risks including irrelevant data in the context window, eliminating the token savings and potentially introducing noise that degrades model performance.

Where is the progressive disclosure logic actually implemented?

The MCP server in src/servers/mcp-server.ts defines the three tools and their orchestration (lines 59–99), while the actual data retrieval logic resides in src/services/worker-service.ts. The worker implements the /api/search, /api/timeline, and /api/observations/batch endpoints that the MCP tools invoke. This separation allows the MCP layer to enforce the workflow while the worker handles storage-specific optimizations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →