How the L0-L3 Memory Distillation Pipeline Works in TencentDB Agent Memory
The L0-L3 Memory Distillation Pipeline in TencentDB Agent Memory transforms raw conversation logs into hierarchical, reusable knowledge assets through an asynchronous four-layer process that progresses from immutable chat records to long-term persona profiles.
The TencentDB Agent Memory system solves a critical problem in AI agent architectures: how to convert ephemeral dialogue into structured, retrievable knowledge that persists across sessions. The L0-L3 Memory Distillation Pipeline implements this through progressive refinement layers, each building upon the previous to extract increasing semantic value.
The Four Layers of Memory Distillation
The pipeline organizes memory into four distinct tiers, from raw logs to synthesized personas:
L0 Conversation: The Immutable Source
L0 Conversation stores the complete, timestamped record of every message exchanged. The Conversation Recorder (MemoryCore/src/core/conversation/l0-recorder.ts) writes each interaction to JSONL files as it occurs.
This layer serves as the audit trail and replay source. Nothing in L0 is ever modified—subsequent layers derive from it, ensuring traceability.
L1 Atom: Extracted Facts and Constraints
L1 Atom contains structured data points extracted from L0 logs. The Atomizer worker (MemoryCore/src/core/atomizer/atomizer_worker.ts) processes raw conversations asynchronously, invoking the configured LLM (Claude, DeepSeek, etc.) with a prompt designed to isolate concrete items: user preferences, deadlines, technical constraints, events.
The LLM returns JSON that the atomizer persists as discrete Atom assets. This enables precise fact lookup without scanning entire conversation histories.
L2 Scenario: Contextual Workstream Snapshots
L2 Scenario groups related atoms into cohesive knowledge blocks. The Scenario Builder (MemoryCore/src/core/scenario/scenario_worker.ts) aggregates atoms sharing a scenario_id or topical cluster, attaching linked Wiki or CodeGraph pages referenced in those atoms.
Scenarios provide agents with ready-made contextual snapshots—for example, "Product Launch v2" containing all relevant preferences, constraints, and documentation in one retrievable unit.
L3 Core / Persona: Long-Term Cognitive Profiles
L3 Core / Persona captures stable patterns across all scenarios for a user or team. The Persona Synthesizer (MemoryCore/src/core/persona/persona_worker.ts) runs periodically or on-demand, summarizing recurring preferences, style guidelines, and domain expertise into a Persona asset.
This layer allows newly created agents to bootstrap with rich, pre-populated mindsets rather than learning from scratch.
Asynchronous Pipeline Execution
The L0-L3 distillation operates through four asynchronous stages managed by background workers in MemoryCore/src/core/worker/*:
- Ingestion – Conversation completion triggers the L0 recorder to persist the raw log
- Distillation Trigger – A watcher detects new L0 data and enqueues a "distill" task
- Layer Generation – Workers sequentially produce L1 atoms, L2 scenarios, and L3 personas, persisting each to the hub database
- Binding – The Binding Manager (
MemoryCore/src/core/binding/binding_manager.ts) enforces ACL rules determining which layers each agent may access
Layered Retrieval Strategy
When an agent queries the memory hub, the Retriever (MemoryCore/src/core/retrieval/retriever.ts) implements a priority fallback:
- Primary: Query L2 scenarios and L3 personas for immediate contextual bootstrap
- Fallback: If specific facts are missing, execute BM25 + vector search across L1 atoms and L0 conversations
- Safety: Result sets are capped to keep context windows within LLM limits
This design prioritizes pre-distilled, semantically dense layers while preserving access to raw source material when needed.
Practical Implementation with the Python SDK
The following examples demonstrate triggering distillation and retrieving layered assets:
Recording L0 and Triggering Distillation
from tencentdb_agent_memory.v3 import client
# Initialize client pointing to Memory Proxy
mem = client.MemoryClient(base_url="http://localhost:8125/api/v3")
# Define conversation exchange
conversation = [
{"role": "user", "content": "We need a React app that talks to a MySQL DB."},
{"role": "assistant", "content": "Sure, I can scaffold that for you."},
]
# Import raw log — hub automatically enqueues background distillation
mem.import_conversation(agent_id="builder-01", messages=conversation)
Retrieving Distilled Layers
# Fetch L1 atoms after background processing completes
atoms = mem.list_atoms(agent_id="builder-01")
print("Extracted facts:", atoms)
# Retrieve L2 scenario grouping those atoms
scenarios = mem.list_scenarios(agent_id="builder-01")
print("Workstream context:", scenarios)
# Access L3 persona for long-term team knowledge
persona = mem.get_persona(team_id="my-team")
print("Team cognitive profile:", persona)
Each method maps to REST endpoints (/v3/atoms, /v3/scenarios, /v3/personas) that return the latest available distilled layers.
Key Source Files
| Component | Path | Responsibility |
|---|---|---|
| L0 Recorder | MemoryCore/src/core/conversation/l0-recorder.ts |
Persists raw chat logs, notifies workers |
| L1 Atomizer | MemoryCore/src/core/atomizer/atomizer_worker.ts |
LLM-based fact extraction |
| L2 Scenario Builder | MemoryCore/src/core/scenario/scenario_worker.ts |
Atom grouping and context assembly |
| L3 Persona Synthesizer | MemoryCore/src/core/persona/persona_worker.ts |
Long-term pattern summarization |
| Retrieval Engine | MemoryCore/src/core/retrieval/retriever.ts |
Layered fallback search implementation |
| Binding & ACL | MemoryCore/src/core/binding/binding_manager.ts |
Agent memory access control |
Summary
- L0-L3 Memory Distillation progressively refines raw conversation into structured, reusable knowledge
- Four layers: L0 (immutable logs) → L1 (extracted facts) → L2 (scenario snapshots) → L3 (persona profiles)
- Asynchronous workers handle distillation without blocking agent operations
- Layered retrieval prioritizes L2/L3 for speed, falling back to L1/L0 with BM25+vector search when needed
- ACL-enforced binding ensures agents access only authorized memory assets
Frequently Asked Questions
How long does the L0-L3 distillation process take?
Distillation latency depends on LLM provider response times and queue depth. L0 writes are immediate; L1-L3 generation typically completes within seconds to minutes based on conversation complexity. The Python SDK's list_atoms() and related methods transparently return available layers while background workers continue processing.
Can agents access L0 directly, or must they use distilled layers?
Agents can access L0 through the retrieval fallback path, but the system design discourages this. According to MemoryCore/src/core/retrieval/retriever.ts, the retrieval engine attempts L2/L3 first and only descends to L1/L0 when higher layers lack needed information. The Binding Manager enforces granular ACL controls per layer.
What determines when L3 personas update?
The Persona Synthesizer supports both periodic scheduled runs and on-demand invocation. Trigger conditions are configurable; common patterns include nightly batch jobs for stable teams or immediate synthesis after significant scenario accumulation. The synthesizer analyzes cross-scenario patterns rather than reacting to single events.
How does the pipeline handle conflicting information across conversations?
The Atomizer timestamps all extracted facts with provenance to L0 records. When conflicts emerge, the Scenario Builder and Persona Synthesizer apply recency-weighted resolution strategies, with newer atoms generally superseding older ones unless explicitly marked as persistent constraints. The immutable L0 layer preserves full auditability for dispute resolution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →