How GenericAgent Optimizes Token Usage with Its Layered Memory System Under 30K Context

GenericAgent stays well under the 30K token limit by routing information through a four-layer memory hierarchy that compresses chat history, loads task knowledge on demand, and keeps only a 1KB routing index in the active prompt.

Managing long-running AI agent conversations typically leads to bloated prompts that exceed model context limits and degrade performance. The lsdefine/GenericAgent repository solves this through a sophisticated layered memory architecture that aggressively optimizes token usage while maintaining instant access to historical context and task-specific knowledge.

The Four-Layer Memory Architecture

GenericAgent isolates information across four distinct layers, preventing rarely-used data from consuming valuable prompt space. This design ensures the active context remains lean while preserving deep knowledge in readily accessible storage.

L4 – Raw Session History (Disk-Only Archive)

Located in memory/L4_raw_sessions/compress_session.py, this layer stores full raw logs of every interaction as compressed session files on disk. L4 never loads into RAM or the prompt—it serves purely as an archival reference that keeps historic data completely outside the token budget.

L3 – Task-Level Knowledge Base (On-Demand Loading)

Defined in memory/memory_management_sop.md (lines 50-60), this layer contains task-specific SOPs, scripts, or data files too large for active memory but essential for specific workflows. Rather than prepending this content to every request, the system accesses it via file_read tool calls only when trigger keywords match, keeping the prompt short during unrelated tasks.

L2 – Global Fact Store (Semi-Static Context)

Stored in memory/global_mem.txt (generated at runtime), this layer holds environment-specific facts such as paths, configuration values, and credentials-free constants. The system inserts these into prompts as a single compact block wrapped in <facts>…</facts> tags, referenced by pointers from the L1 layer to minimize redundancy.

L1 – Insight Index (Active Routing Table)

The memory/global_mem_insight.txt file maintains an ultra-compact mapping of high-frequency trigger keywords to their corresponding L2/L3 locations. Limited to ≤30 lines (approximately 1KB), this index is prepended to every request via the get_global_memory() function. Acting as a routing table, it eliminates the need for the model to scan the entire knowledge base when searching for relevant context.

Token Compression Mechanisms in llmcore.py

The core token optimization engine resides in llmcore.py, which implements aggressive compression algorithms to enforce hard context window limits.

Tag-Based History Compression

The compress_history_tags function (lines 23-55) scans the chat history every fifth turn and truncates verbose content within <thinking>, <tool_use>, <tool_result>, and <history> tags to a configurable max_len (default 800 characters). It replaces the middle section of long strings with […] markers:


# llmcore.py – compress older messages

def compress_history_tags(messages, keep_recent=10, max_len=800, force=False):
    ...
    _trunc_str = lambda s: s[:max_len//2] + '\n...[Truncated]...\n' + s[-max_len//2:]

This process preserves recent context while compressing older interaction metadata that remains technically accessible but rarely needed in full.

Context Window Trimming

The trim_messages_history function (lines 74-87) first invokes compress_history_tags, then removes the oldest non-user messages when the cumulative token count exceeds three times the configured window (context_win * 3). It also sanitizes the leading user message to strip stray <tool_result> blocks, guaranteeing a lean prompt payload:


# llmcore.py – drop oldest messages when over budget

def trim_messages_history(history, context_win):
    compress_history_tags(history)
    cost = sum(len(json.dumps(m, ensure_ascii=False)) for m in history)
    ...

Memory-Aware Prompting with get_global_memory()

In ga.py (line 565), the get_global_memory() function reads the L1 insight file and injects it into the system prompt using the [Memory] marker:


# ga.py – insert L1 index into the prompt

def get_global_memory():
    with open(os.path.join(script_dir, 'memory/global_mem_insight.txt'), 'r', ...) as f:
        insight = f.read()
    return f"\n[Memory] (../memory)\n{insight}"

Because L1 is strictly limited to approximately 1KB, this injection adds minimal overhead while providing the model with a complete map to locate L2 facts and L3 task files as needed.

Lazy Loading for L3 Task Files

When tools like file_read or web_scan are invoked, the agent first consults the L1 Insight Index for matching keywords. Only if the keyword exists does the system read the corresponding L2/L3 file, implementing lazy loading that avoids unnecessary token consumption. Additionally, the request payload in llmcore.py (line 470) sets the minimum necessary max_tokens value derived from user configuration in mykey_template.py, preventing LLM providers from over-allocating context space.

Practical Implementation Examples

Compressing Chat History Before LLM Calls

from llmcore import compress_history_tags, trim_messages_history

# `chat` is a list of message dicts from previous turns

chat = [...]  

# First compress tags (keeps the last 10 messages intact)

compress_history_tags(chat, keep_recent=10, max_len=800)

# Then trim the history to fit a 8,000-token window

trim_messages_history(chat, context_win=8000)

# `chat` is now token-lean and ready for the LLM call

Adding L1 Index Entries for New L2 Facts

from ga import get_global_memory, log_memory_access

# Suppose we added a new environment variable to global_mem.txt via file_patch

new_key = "my_special_path"
new_value = "/opt/special"

# Update L1 manually (or via a tool that patches the insight file)

insight_line = f"{new_key} → global_mem.txt#{new_key}"
with open("memory/global_mem_insight.txt", "a", encoding="utf-8") as f:
    f.write(insight_line + "\n")

# Log the change for debugging / audit

log_memory_access("memory/global_mem_insight.txt")

Triggering L3 SOP Loading on Demand

from ga import GenericAgentHandler

handler = GenericAgentHandler(parent=None)

# The user asks: "How do I scrape a WeChat login QR code?"

# L1 contains the keyword "wechat_qr"

prompt = get_global_memory()  # inject L1

# The agent sees the keyword, then reads the SOP from L3

sop_path = "../memory/wechat_qr_sop.md"
with open(sop_path, "r", encoding="utf-8") as f:
    sop_content = f.read()

# The SOP text is now inserted as a tool result (tiny token cost)

Summary

  • Four-layer isolation routes raw history (L4), task scripts (L3), semi-static facts (L2), and a routing index (L1) to appropriate storage tiers, keeping active prompts minimal.
  • History compression via compress_history_tags and trim_messages_history enforces hard token budgets by truncating verbose XML tags and removing aged messages.
  • L1 Insight Index replaces full knowledge base scanning with a ~1KB routing table prepended to every request.
  • Lazy loading mechanisms ensure task-specific content only consumes tokens when explicitly triggered by keywords.

Frequently Asked Questions

What is the maximum token budget GenericAgent targets?

While the system architecture supports context windows up to 30K tokens, GenericAgent actively maintains active prompts under approximately 3,000 tokens through aggressive compression and the layered memory hierarchy. The 30K limit serves as a hard ceiling rarely approached during normal operation.

How does GenericAgent handle very long conversation histories?

The compress_history_tags function in llmcore.py scans chat history every fifth turn and truncates verbose XML tags—including <thinking>, <tool_use>, and <tool_result>—to 800 characters by default. It replaces the middle content with […] markers while preserving the beginning and end of each tag.

When does the agent load content from L3 task files?

L3 content loads only when the L1 Insight Index contains a matching trigger keyword for the specific task. This lazy loading approach prevents unused task documentation from consuming tokens during unrelated conversations, as implemented in the file_read tool logic.

Can the token compression settings be customized?

Yes, the compress_history_tags function accepts configurable parameters including keep_recent (default 10 messages) and max_len (default 800 characters), allowing developers to adjust compression aggressiveness based on specific model context window requirements or verbosity needs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →