How the L0-L3 Memory Distillation Pipeline Works in TencentDB Agent Memory

The L0-L3 Memory Distillation Pipeline in TencentDB Agent Memory transforms raw conversation logs into hierarchical, reusable knowledge assets through an asynchronous four-layer process that progresses from immutable chat records to long-term persona profiles.

The TencentDB Agent Memory system solves a critical problem in AI agent architectures: how to convert ephemeral dialogue into structured, retrievable knowledge that persists across sessions. The L0-L3 Memory Distillation Pipeline implements this through progressive refinement layers, each building upon the previous to extract increasing semantic value.

The Four Layers of Memory Distillation

The pipeline organizes memory into four distinct tiers, from raw logs to synthesized personas:

L0 Conversation: The Immutable Source

L0 Conversation stores the complete, timestamped record of every message exchanged. The Conversation Recorder (MemoryCore/src/core/conversation/l0-recorder.ts) writes each interaction to JSONL files as it occurs.

This layer serves as the audit trail and replay source. Nothing in L0 is ever modified—subsequent layers derive from it, ensuring traceability.

L1 Atom: Extracted Facts and Constraints

L1 Atom contains structured data points extracted from L0 logs. The Atomizer worker (MemoryCore/src/core/atomizer/atomizer_worker.ts) processes raw conversations asynchronously, invoking the configured LLM (Claude, DeepSeek, etc.) with a prompt designed to isolate concrete items: user preferences, deadlines, technical constraints, events.

The LLM returns JSON that the atomizer persists as discrete Atom assets. This enables precise fact lookup without scanning entire conversation histories.

L2 Scenario: Contextual Workstream Snapshots

L2 Scenario groups related atoms into cohesive knowledge blocks. The Scenario Builder (MemoryCore/src/core/scenario/scenario_worker.ts) aggregates atoms sharing a scenario_id or topical cluster, attaching linked Wiki or CodeGraph pages referenced in those atoms.

Scenarios provide agents with ready-made contextual snapshots—for example, "Product Launch v2" containing all relevant preferences, constraints, and documentation in one retrievable unit.

L3 Core / Persona: Long-Term Cognitive Profiles

L3 Core / Persona captures stable patterns across all scenarios for a user or team. The Persona Synthesizer (MemoryCore/src/core/persona/persona_worker.ts) runs periodically or on-demand, summarizing recurring preferences, style guidelines, and domain expertise into a Persona asset.

This layer allows newly created agents to bootstrap with rich, pre-populated mindsets rather than learning from scratch.

Asynchronous Pipeline Execution

The L0-L3 distillation operates through four asynchronous stages managed by background workers in MemoryCore/src/core/worker/*:

  1. Ingestion – Conversation completion triggers the L0 recorder to persist the raw log
  2. Distillation Trigger – A watcher detects new L0 data and enqueues a "distill" task
  3. Layer Generation – Workers sequentially produce L1 atoms, L2 scenarios, and L3 personas, persisting each to the hub database
  4. Binding – The Binding Manager (MemoryCore/src/core/binding/binding_manager.ts) enforces ACL rules determining which layers each agent may access

Layered Retrieval Strategy

When an agent queries the memory hub, the Retriever (MemoryCore/src/core/retrieval/retriever.ts) implements a priority fallback:

  • Primary: Query L2 scenarios and L3 personas for immediate contextual bootstrap
  • Fallback: If specific facts are missing, execute BM25 + vector search across L1 atoms and L0 conversations
  • Safety: Result sets are capped to keep context windows within LLM limits

This design prioritizes pre-distilled, semantically dense layers while preserving access to raw source material when needed.

Practical Implementation with the Python SDK

The following examples demonstrate triggering distillation and retrieving layered assets:

Recording L0 and Triggering Distillation

from tencentdb_agent_memory.v3 import client

# Initialize client pointing to Memory Proxy

mem = client.MemoryClient(base_url="http://localhost:8125/api/v3")

# Define conversation exchange

conversation = [
    {"role": "user", "content": "We need a React app that talks to a MySQL DB."},
    {"role": "assistant", "content": "Sure, I can scaffold that for you."},
]

# Import raw log — hub automatically enqueues background distillation

mem.import_conversation(agent_id="builder-01", messages=conversation)

Retrieving Distilled Layers


# Fetch L1 atoms after background processing completes

atoms = mem.list_atoms(agent_id="builder-01")
print("Extracted facts:", atoms)

# Retrieve L2 scenario grouping those atoms

scenarios = mem.list_scenarios(agent_id="builder-01")
print("Workstream context:", scenarios)

# Access L3 persona for long-term team knowledge

persona = mem.get_persona(team_id="my-team")
print("Team cognitive profile:", persona)

Each method maps to REST endpoints (/v3/atoms, /v3/scenarios, /v3/personas) that return the latest available distilled layers.

Key Source Files

Component Path Responsibility
L0 Recorder MemoryCore/src/core/conversation/l0-recorder.ts Persists raw chat logs, notifies workers
L1 Atomizer MemoryCore/src/core/atomizer/atomizer_worker.ts LLM-based fact extraction
L2 Scenario Builder MemoryCore/src/core/scenario/scenario_worker.ts Atom grouping and context assembly
L3 Persona Synthesizer MemoryCore/src/core/persona/persona_worker.ts Long-term pattern summarization
Retrieval Engine MemoryCore/src/core/retrieval/retriever.ts Layered fallback search implementation
Binding & ACL MemoryCore/src/core/binding/binding_manager.ts Agent memory access control

Summary

  • L0-L3 Memory Distillation progressively refines raw conversation into structured, reusable knowledge
  • Four layers: L0 (immutable logs) → L1 (extracted facts) → L2 (scenario snapshots) → L3 (persona profiles)
  • Asynchronous workers handle distillation without blocking agent operations
  • Layered retrieval prioritizes L2/L3 for speed, falling back to L1/L0 with BM25+vector search when needed
  • ACL-enforced binding ensures agents access only authorized memory assets

Frequently Asked Questions

How long does the L0-L3 distillation process take?

Distillation latency depends on LLM provider response times and queue depth. L0 writes are immediate; L1-L3 generation typically completes within seconds to minutes based on conversation complexity. The Python SDK's list_atoms() and related methods transparently return available layers while background workers continue processing.

Can agents access L0 directly, or must they use distilled layers?

Agents can access L0 through the retrieval fallback path, but the system design discourages this. According to MemoryCore/src/core/retrieval/retriever.ts, the retrieval engine attempts L2/L3 first and only descends to L1/L0 when higher layers lack needed information. The Binding Manager enforces granular ACL controls per layer.

What determines when L3 personas update?

The Persona Synthesizer supports both periodic scheduled runs and on-demand invocation. Trigger conditions are configurable; common patterns include nightly batch jobs for stable teams or immediate synthesis after significant scenario accumulation. The synthesizer analyzes cross-scenario patterns rather than reacting to single events.

How does the pipeline handle conflicting information across conversations?

The Atomizer timestamps all extracted facts with provenance to L0 records. When conflicts emerge, the Scenario Builder and Persona Synthesizer apply recency-weighted resolution strategies, with newer atoms generally superseding older ones unless explicitly marked as persistent constraints. The immutable L0 layer preserves full auditability for dispute resolution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →