How the Layered Memory System (L0-L4) Works in GenericAgent: A Deep Dive into the Five-Level Architecture
The Layered Memory System in GenericAgent organizes persistent knowledge into five distinct levels (L0-L4), where L4 automatically compresses raw session logs into searchable archives every 12 hours via a background scheduler, enabling long-term recall while minimizing token usage and storage costs.
The GenericAgent repository implements a sophisticated Layered Memory System (L0-L4) to manage persistent knowledge across a hierarchical five-level stack. This architecture separates immutable behavioral constraints from dynamic session histories, optimizing for both immediate retrieval and long-term storage efficiency. While L0 through L3 handle static configurations, environment facts, and task-specific procedures, the L4 layer serves as the automated archival engine that transforms raw dialogue logs into compressed, searchable histories.
The Five-Level Memory Hierarchy (L0-L4)
GenericAgent partitions all persistent knowledge into five layers, each designed to minimize token consumption and hallucination risks:
- L0 – Meta Rules: Hard-coded constants embedded directly in the source code (few KB). These contain core behavioral constraints and safety rules that never change during runtime.
- L1 – Insight Index: Stored in
global_mem_insight.txt(≤30 lines, <1 KB). This ultra-compact "pointer map" routes queries to deeper layers without loading full contexts. - L2 – Global Facts: Stored in
global_mem.txt(few KB to MB). Holds stable environment facts such as file paths, credentials, and configuration constants. - L3 – Task-Level Records: Stored in the
memory/folder as Markdown and Python files (unbounded). Contains reusable SOPs, scripts, or specialized data for particular tasks. - L4 – Session Archive: Stored in
memory/L4_raw_sessions/(GB-scale when archived). Contains compressed logs of completed sessions, turned into a searchable history for long-horizon recall.
The L4 layer is the focus of most archival operations—it automatically processes raw model-response logs into compact, indexed archives.
How L4 Session Archiving Is Triggered
A background scheduler in reflect/scheduler.py manages the L4 compression cycle. Every 12 hours (43,200 seconds), it invokes the archiver:
# reflect/scheduler.py (excerpt)
if _time.time() - _l4_t > 43200: # 12 h
_l4_t = _time.time()
import sys; sys.path.insert(0, os.path.join(_dir, '../memory/L4_raw_sessions'))
from compress_session import batch_process
raw_dir = os.path.join(_dir, '../temp/model_responses')
r = batch_process(raw_dir, dry_run=False) # real run
print(f'[L4 cron] {r}')
The scheduler dynamically inserts the L4_raw_sessions package into the import path, loads compress_session.batch_process, and points it at temp/model_responses where the model writes raw logs.
The Compression Pipeline: Inside compress_session.py
The memory/L4_raw_sessions/compress_session.py file implements the core L4 logic through four primary operations.
Parsing and Compressing Raw Sessions
The compress_session function handles individual log files:
def compress_session(src, dst_dir=None):
"""
• Reads a `model_responses_*.txt` file.
• Detects its format (JSON‑style or “raw” with markers).
• Extracts timestamps to build a filename: MMDD_HHMM‑MMDD_HHMM.txt
• If raw, calls `_compress_raw` to strip the duplicated system‑prompt /
assistant‑echo sections.
• Writes the compressed text to `dst_dir` (default = L4 directory).
"""
The function returns the destination path and a statistics dictionary containing original size, new size, and compression ratio.
Stripping Redundant Content
Raw model outputs often contain duplicated content. The _compress_raw function removes the assistant echo:
def _compress_raw(text):
"""
Raw format (B) looks like:
=== Prompt === …
=== USER === …
=== ASSISTANT === … ← echo of the prompt, not needed
=== Response === …
The function keeps only Prompt, USER, and Response blocks,
removing the assistant‑echo block.
"""
By discarding the === ASSISTANT === echo block while preserving the essential Prompt, USER, and Response sections, the system dramatically reduces file size without losing conversational context.
Extracting History for Searchable Recall
The extract_history function prepares data for fast lookup:
def extract_history(src, session_name=None):
"""
Reads a compressed L4 file and pulls all `<history>` XML blocks.
Returns a deduplicated list of "[USER]" / "[Agent]" lines.
"""
The extracted history is appended to all_histories.txt, enabling the agent to search across all past sessions without loading individual files.
Batch Processing and Monthly Archiving
The batch_process function orchestrates the complete pipeline through six phases:
- Scan: Identifies all
model_responses_*.txtfiles, skipping those modified within the last 2 hours - Compress & Extract: Invokes
compress_sessionandextract_historyfor each file - Append History: Merges extracted histories into
all_histories.txt - Archive: Packages sessions into monthly zip files (
YYYY-MM.zip) - Cleanup: Deletes original raw files and duplicates to free space
When dry_run=False, the scheduler triggers this pipeline to persist changes permanently.
Integration With Higher Memory Layers
The L4 archive maintains bidirectional relationships with upper layers:
- L3 → L4: When tasks complete, the agent writes SOPs and scripts to
memory/(L3). The raw model outputs that generated these artifacts remain intemp/model_responses/until the L4 archiver compresses them, making the origin of every L3 artifact traceable through the session archive. - L1 Indexing: According to
memory_management_sop.md, the L1 Insight Index must contain pointers to new L3 or L4 entries. After a successful L4 batch run, maintenance routines updateglobal_mem_insight.txtwith concise key-to-session mappings. - L2 Facts: Stable environment facts discovered during sessions (API endpoints, credentials) are written to
global_mem.txt. The L4 archive stores the conversational context that led to these facts, providing provenance for future reasoning.
Practical Examples: Manual L4 Operations
While the scheduler handles automatic archiving, developers can interact with L4 directly.
Running the Archiver On Demand
Trigger the compression pipeline manually for specific directories:
# Example: compress and archive all raw model responses in ./temp/model_responses
from memory.L4_raw_sessions.compress_session import batch_process
raw_folder = 'temp/model_responses' # path relative to repository root
report = batch_process(raw_folder, dry_run=False) # set dry_run=False to apply
print('L4 archiving finished:')
print(f" Processed: {report['processed']}")
print(f" New sessions added: {report['new_sessions']}")
print(f" Raw files deleted: {report.get('deleted_raw', 0)}")
Setting dry_run=False persists the compressed archives and deletes the original raw files, identical to the scheduler's cron job.
Retrieving Historical Sessions
Query the consolidated history file to recall previous conversations:
import os
L4_DIR = os.path.join('memory', 'L4_raw_sessions')
history_path = os.path.join(L4_DIR, 'all_histories.txt')
def search_history(keyword: str):
"""Return all session blocks that contain the keyword."""
with open(history_path, 'r', encoding='utf-8') as f:
content = f.read()
matches = [block for block in content.split('============================================================')
if keyword.lower() in block.lower()]
return matches
# Example usage:
for block in search_history('wechat'):
print('--- SESSION START ---')
print(block.strip())
print('--- SESSION END ---\n')
The search_history function scans all_histories.txt, returning any session blocks where the keyword appears. This provides the primary mechanism for the L4 layer to supply long-term recall to the agent's reasoning engine.
Key Implementation Files
| File | Purpose |
|---|---|
memory/memory_management_sop.md |
Architectural specification defining all memory layers (L0-L4) and policies governing their interactions. |
reflect/scheduler.py |
Background scheduler that triggers the L4 archiver every 12 hours via the batch_process function. |
memory/L4_raw_sessions/compress_session.py |
Core implementation containing compress_session, _compress_raw, extract_history, and batch_process functions. |
memory/L4_raw_sessions/all_histories.txt |
Consolidated, deduplicated history file enabling fast keyword search across all archived sessions. |
README.md |
High-level overview of the five-layer memory system for new contributors. |
Summary
- The Layered Memory System (L0-L4) in GenericAgent separates knowledge into five distinct tiers, from hard-coded meta rules (L0) to compressed session archives (L4).
- L4 automation runs every 12 hours via
reflect/scheduler.py, invokingcompress_session.pyto process raw logs without manual intervention. - Compression removes redundancy by stripping assistant echo blocks while preserving essential Prompt, USER, and Response sections, dramatically reducing storage footprint.
- Searchable history is extracted into
all_histories.txt, enabling fast keyword-based recall across all past sessions without loading individual archive files. - Integration with upper layers ensures that L3 artifacts remain traceable to their L4 origins, while L1 and L2 layers receive updated pointers and facts derived from archived sessions.
Frequently Asked Questions
What is the difference between L3 and L4 in GenericAgent's memory system?
L3 (Task-Level Records) stores reusable SOPs, scripts, and specialized data files in the memory/ folder as Markdown and Python files, representing the agent's procedural knowledge. L4 (Session Archive) stores the raw conversational logs and model responses that generated those L3 artifacts, compressed and archived in memory/L4_raw_sessions/. While L3 contains the "what" (final procedures), L4 contains the "why" and "how" (conversational context and reasoning traces).
How often does the L4 archiver run automatically?
The L4 archiver runs automatically every 12 hours (43,200 seconds) via the background scheduler defined in reflect/scheduler.py. This cron-like mechanism checks the timestamp difference _time.time() - _l4_t > 43200 and invokes compress_session.batch_process when the threshold is met, ensuring that raw model logs are regularly compressed without manual intervention.
Can I manually trigger L4 compression for specific sessions?
Yes, you can manually invoke the L4 pipeline by importing batch_process from memory.L4_raw_sessions.compress_session and calling it with your target directory. Set dry_run=False to apply the compression and delete the original raw files, or dry_run=True to preview the changes. This allows on-demand archiving outside the 12-hour scheduler cycle, useful for debugging or immediate storage cleanup.
How does GenericAgent prevent token overflow with the Layered Memory System?
The Layered Memory System (L0-L4) prevents token overflow through strict size constraints and selective retrieval. L0 (Meta Rules) and L1 (Insight Index) are kept under a few KB and 30 lines respectively, loading instantly into context. L4 archives are excluded from active context unless specifically queried via search_history() on the consolidated all_histories.txt file. By compressing raw logs (stripping redundant assistant echoes) and archiving older sessions into monthly zip files, the system ensures only the most relevant, compact data (L1-L3) consumes active token budget while maintaining GB-scale history in cold storage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →