How Agent Memory Is Condensed in Munder Difflin to Prevent Unbounded Growth
Munder Difflin prevents unbounded agent memory growth through a MemoryReflector service that periodically compresses oversized memory.md files into a three-region structure: pinned durable facts, an LLM-summarized condensed history, and verbatim recent sections.
Munder Difflin stores each agent's long-term notes in plain-text memory.md files. Without intervention, these files would exhaust disk space and degrade semantic-search performance. The solution is a deterministic memory condensation pipeline that runs in the Electron main process, enforcing a hard size budget while preserving critical information.
The Three-Region Memory Structure
After condensation, every memory.md follows a fixed shape:
| Region | Purpose |
|---|---|
| Pinned facts | Durable information that must never be removed—kept verbatim |
| Condensed history | A short LLM-generated summary of older content |
| Recent sections | The newest K sections (verbatim) the user is actively working with |
This structure ensures bounded growth while maintaining both long-term knowledge and immediate context.
Trigger Detection: When to Condense
The MemoryReflector scans agent files on a configurable interval. Condensation triggers when either threshold is met, as implemented in shouldCondense at src/main/reflect.ts【74-84】:
// Constants from the source
const BUDGET_BYTES = 131_072; // 128 KB hard ceiling
// Trigger conditions (configurable)
- byteTriggerPct: 80 // 80% of budget = ~102 KB
- sectionTrigger: 150 // number of '## ' headings
- minBytes: 4_096 // ignore files under 4 KB
A file condenses when:
- Its size exceeds
byteTriggerPctpercent ofBUDGET_BYTES, or - Its section count exceeds
sectionTriggerand size is aboveminBytes
This dual-threshold prevents premature condensation of small, active files while catching verbose histories early.
The 10-Step Condensation Pipeline
1. Parsing and Partitioning
The parseMemory function at src/main/reflect.ts【28-38】 splits markdown into the three regions and extracts all ## sections. The Recent region divides into:
- Keep: newest
recentKeepsections (default: 5) stay verbatim - Evict: older sections marked for summarization
2. Atomic Backup
Before any modification, the original copies to hive/backups/<timestamp>/<agent>/memory.md at src/main/reflect.ts【200-206】. This enables recovery from LLM failures or verification rejects.
3. LLM Summarization
The evict tail plus any existing condensed summary feeds to a headless Claude-Haiku model (claude-haiku-4-5) via runHiddenClaude. The system prompt (CONDENSE_SYSTEM) at src/main/reflect.ts【44-56】 demands strict JSON output:
{
"condensed": "Brief narrative summary...",
"hoisted": ["Critical fact to pin", "Another durable insight"]
}
4. Pinned-Line Hoisting
New durable facts merge into the pinned block via mergePinned at src/main/reflect.ts【47-55】, with duplicate detection. This ensures extracted wisdom survives future condensations.
5. Re-assembly and Verification
The rebuild function at src/main/reflect.ts【58-68】 stitches header, updated pinned block, new condensed summary, and kept recent sections. Then verify at src/main/reflect.ts【71-90】 enforces strict invariants:
- Valid three-region structure
- Non-empty condensed summary
- New file strictly smaller than old (≤95% of original size)
- All original pinned lines retained
6. Atomic Write
Passing verification triggers atomicWrite at src/main/reflect.ts【36-43】: write to temp sibling, fsync, then rename over original. This eliminates partial-write corruption.
7. Event Logging
Results emit as condense or condense-abort events for UI visibility, logged at src/main/reflect.ts【44-50】.
Manual Invocation Example
Trigger condensation on-demand for debugging or UI controls:
import { MemoryReflector } from './reflect';
const reflector = new MemoryReflector(
() => '/Users/me/.munder-difflin', // getHome
() => 'claude', // getCommand
() => process.env.MEMPALACE_PALACE_PATH
? { MEMPALACE_PALACE_PATH: process.env.MEMPALACE_PALACE_PATH }
: {},
() => ({
enabled: true,
intervalMs: 300_000, // 5 minutes
byteTriggerPct: 80,
sectionTrigger: 150,
recentKeep: 5,
minBytes: 4_096,
}),
(evt) => console.log('Memory event:', evt)
);
// Condense single agent immediately
reflector.reflectNow('alice').then((results) => {
console.log('Condensation result:', results[0]);
});
Configuration Reference
| Parameter | Default | Purpose |
|---|---|---|
intervalMs |
300000 | Scan frequency (5 min) |
byteTriggerPct |
80 | Size threshold (% of 128 KB) |
sectionTrigger |
150 | Section count threshold |
recentKeep |
5 | Verbatim sections to preserve |
minBytes |
4096 | Floor for section-based trigger |
Key Implementation Files
| File | Responsibility |
|---|---|
src/main/reflect.ts |
MemoryReflector core—detection, backup, LLM calls, verification, atomic rewrite |
src/main/memory.ts |
Semantic memory (MemPalace) coordination; provides memory.start() hook |
src/main/hiddenClaude.ts |
Headless Claude model wrapper for summarization |
src/main/palaceReap.ts |
Quarantine cleanup of temporary palace copies |
src/main/triggerHistory.ts |
Condensed change storage for UI tooling |
Summary
- Bounded by design: Hard 128 KB budget with configurable triggers
- Three-region architecture: Pinned facts survive forever; condensed history shrinks indefinitely; recent context stays fresh
- Verification-guarded: Every rewrite must pass structural, size, and content checks
- Atomic and logged: Backup, fsync, and event emission ensure reliability and observability
- LLM-powered but deterministic: Claude-Haiku summarizes, but strict JSON contracts and verification prevent drift
Frequently Asked Questions
What happens if the LLM produces invalid JSON or loses pinned facts?
The verify function at src/main/reflect.ts【71-90】 rejects any rewrite that fails structural validation, drops pinned lines, or exceeds size limits. The original file remains untouched, and a condense-abort event logs the specific failure reason. The backup at hive/backups/ preserves pre-condensation state for manual recovery.
Can users disable memory condensation entirely?
Yes—set enabled: false in ReflectSettings. However, this risks unbounded memory.md growth and degraded semantic search performance. The 128 KB BUDGET_BYTES constant is currently hardcoded, though byteTriggerPct offers proportional control.
How does Munder Difflin choose what to pin versus summarize?
Durable fact extraction is delegated to the Claude-Haiku model via the CONDENSE_SYSTEM prompt, which instructs identification of "critical, timeless information." These "hoisted" facts merge into the pinned block with duplicate detection (mergePinned). Users can also manually pin lines using 📌 markers in memory.md.
Why use Claude-Haiku specifically for condensation?
The codebase selects claude-haiku-4-5 as a cost-efficient, fast summarization backend. The runHiddenClaude wrapper at src/main/hiddenClaude.ts launches this model headlessly, avoiding UI latency while maintaining sufficient quality for narrative condensation. The hidden invocation prevents user-facing streaming overhead.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →