How Agent Memory Is Condensed in Munder Difflin to Prevent Unbounded Growth

Munder Difflin prevents unbounded agent memory growth through a MemoryReflector service that periodically compresses oversized memory.md files into a three-region structure: pinned durable facts, an LLM-summarized condensed history, and verbatim recent sections.

Munder Difflin stores each agent's long-term notes in plain-text memory.md files. Without intervention, these files would exhaust disk space and degrade semantic-search performance. The solution is a deterministic memory condensation pipeline that runs in the Electron main process, enforcing a hard size budget while preserving critical information.

The Three-Region Memory Structure

After condensation, every memory.md follows a fixed shape:

Region Purpose
Pinned facts Durable information that must never be removed—kept verbatim
Condensed history A short LLM-generated summary of older content
Recent sections The newest K sections (verbatim) the user is actively working with

This structure ensures bounded growth while maintaining both long-term knowledge and immediate context.

Trigger Detection: When to Condense

The MemoryReflector scans agent files on a configurable interval. Condensation triggers when either threshold is met, as implemented in shouldCondense at src/main/reflect.ts【74-84】:

// Constants from the source
const BUDGET_BYTES = 131_072; // 128 KB hard ceiling

// Trigger conditions (configurable)
- byteTriggerPct: 80  // 80% of budget = ~102 KB
- sectionTrigger: 150 // number of '## ' headings

- minBytes: 4_096     // ignore files under 4 KB

A file condenses when:

  1. Its size exceeds byteTriggerPct percent of BUDGET_BYTES, or
  2. Its section count exceeds sectionTrigger and size is above minBytes

This dual-threshold prevents premature condensation of small, active files while catching verbose histories early.

The 10-Step Condensation Pipeline

1. Parsing and Partitioning

The parseMemory function at src/main/reflect.ts【28-38】 splits markdown into the three regions and extracts all ## sections. The Recent region divides into:

  • Keep: newest recentKeep sections (default: 5) stay verbatim
  • Evict: older sections marked for summarization

2. Atomic Backup

Before any modification, the original copies to hive/backups/<timestamp>/<agent>/memory.md at src/main/reflect.ts【200-206】. This enables recovery from LLM failures or verification rejects.

3. LLM Summarization

The evict tail plus any existing condensed summary feeds to a headless Claude-Haiku model (claude-haiku-4-5) via runHiddenClaude. The system prompt (CONDENSE_SYSTEM) at src/main/reflect.ts【44-56】 demands strict JSON output:

{
  "condensed": "Brief narrative summary...",
  "hoisted": ["Critical fact to pin", "Another durable insight"]
}

4. Pinned-Line Hoisting

New durable facts merge into the pinned block via mergePinned at src/main/reflect.ts【47-55】, with duplicate detection. This ensures extracted wisdom survives future condensations.

5. Re-assembly and Verification

The rebuild function at src/main/reflect.ts【58-68】 stitches header, updated pinned block, new condensed summary, and kept recent sections. Then verify at src/main/reflect.ts【71-90】 enforces strict invariants:

  • Valid three-region structure
  • Non-empty condensed summary
  • New file strictly smaller than old (≤95% of original size)
  • All original pinned lines retained

6. Atomic Write

Passing verification triggers atomicWrite at src/main/reflect.ts【36-43】: write to temp sibling, fsync, then rename over original. This eliminates partial-write corruption.

7. Event Logging

Results emit as condense or condense-abort events for UI visibility, logged at src/main/reflect.ts【44-50】.

Manual Invocation Example

Trigger condensation on-demand for debugging or UI controls:

import { MemoryReflector } from './reflect';

const reflector = new MemoryReflector(
  () => '/Users/me/.munder-difflin',   // getHome
  () => 'claude',                      // getCommand
  () => process.env.MEMPALACE_PALACE_PATH 
    ? { MEMPALACE_PALACE_PATH: process.env.MEMPALACE_PALACE_PATH } 
    : {},
  () => ({
    enabled: true,
    intervalMs: 300_000,  // 5 minutes
    byteTriggerPct: 80,
    sectionTrigger: 150,
    recentKeep: 5,
    minBytes: 4_096,
  }),
  (evt) => console.log('Memory event:', evt)
);

// Condense single agent immediately
reflector.reflectNow('alice').then((results) => {
  console.log('Condensation result:', results[0]);
});

Configuration Reference

Parameter Default Purpose
intervalMs 300000 Scan frequency (5 min)
byteTriggerPct 80 Size threshold (% of 128 KB)
sectionTrigger 150 Section count threshold
recentKeep 5 Verbatim sections to preserve
minBytes 4096 Floor for section-based trigger

Key Implementation Files

File Responsibility
src/main/reflect.ts MemoryReflector core—detection, backup, LLM calls, verification, atomic rewrite
src/main/memory.ts Semantic memory (MemPalace) coordination; provides memory.start() hook
src/main/hiddenClaude.ts Headless Claude model wrapper for summarization
src/main/palaceReap.ts Quarantine cleanup of temporary palace copies
src/main/triggerHistory.ts Condensed change storage for UI tooling

Summary

  • Bounded by design: Hard 128 KB budget with configurable triggers
  • Three-region architecture: Pinned facts survive forever; condensed history shrinks indefinitely; recent context stays fresh
  • Verification-guarded: Every rewrite must pass structural, size, and content checks
  • Atomic and logged: Backup, fsync, and event emission ensure reliability and observability
  • LLM-powered but deterministic: Claude-Haiku summarizes, but strict JSON contracts and verification prevent drift

Frequently Asked Questions

What happens if the LLM produces invalid JSON or loses pinned facts?

The verify function at src/main/reflect.ts【71-90】 rejects any rewrite that fails structural validation, drops pinned lines, or exceeds size limits. The original file remains untouched, and a condense-abort event logs the specific failure reason. The backup at hive/backups/ preserves pre-condensation state for manual recovery.

Can users disable memory condensation entirely?

Yes—set enabled: false in ReflectSettings. However, this risks unbounded memory.md growth and degraded semantic search performance. The 128 KB BUDGET_BYTES constant is currently hardcoded, though byteTriggerPct offers proportional control.

How does Munder Difflin choose what to pin versus summarize?

Durable fact extraction is delegated to the Claude-Haiku model via the CONDENSE_SYSTEM prompt, which instructs identification of "critical, timeless information." These "hoisted" facts merge into the pinned block with duplicate detection (mergePinned). Users can also manually pin lines using 📌 markers in memory.md.

Why use Claude-Haiku specifically for condensation?

The codebase selects claude-haiku-4-5 as a cost-efficient, fast summarization backend. The runHiddenClaude wrapper at src/main/hiddenClaude.ts launches this model headlessly, avoiding UI latency while maintaining sufficient quality for narrative condensation. The hidden invocation prevents user-facing streaming overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →