Patterns for Agent Self-Evolution and Continuous Learning: A Deep Dive into the ai-agent-book Repository

The bojieli/ai-agent-book repository implements two core patterns for autonomous improvement: versioned external-memory arms with configurable update policies (static, append-only, evolving) and a closed-loop proposer-reviewer mechanism that enables agents to iteratively edit their own code based on critique feedback.

This repository explores how autonomous agents can evolve their behavior over time without human intervention. Through practical implementations in chapter9/self-evolution-eval/ and chapter9/hermes-self-evolution/, it demonstrates production-ready patterns for agent self-evolution and continuous learning that separate knowledge storage from reasoning, enabling systematic improvement from experience.

External-Memory Arms Pattern

The first pattern treats agent memory as a versioned rule store that operates independently of the LLM's weights. This architecture allows agents to adapt to new information while preserving audit trails of how their behavior changed.

Architecture and Memory Model

At the heart of this pattern lies a dictionary-based memory system defined in chapter9/self-evolution-eval/agent.py. The ReferenceAgent and OpenAILongitudinalAgent classes maintain a memory: Dict[str, MemoryEntry] structure where each entry stores both a value and a version number.


# Memory structure from agent.py

memory = {
    "rule_id": MemoryEntry(value="answer_23kg", version=2)
}

The repository implements three distinct profiles that control how this memory evolves:

  • Static: Memory is read-only; the agent never updates its rule store
  • Append-only: Records every observation but preserves the first version permanently
  • Evolving: Replaces older versions when a higher version number appears, maintaining audit trails while allowing updates

Action and Observation Flow

The separation of acting and observing creates a clean interface for learning. When the agent acts, it queries memory; when it observes, it processes learning signals.

During the action phase in ReferenceAgent.act(), the agent looks up rules by rule_id. If no entry exists and the profile isn't static, it falls back to BASELINE_ACTIONS indexed by task family:


# From chapter9/self-evolution-eval/agent.py

entry = self.memory.get(task["rule_id"])
used_memory = entry is not None and self.profile != "static"
action = entry.value if used_memory else BASELINE_ACTIONS[task["family"]]

The observation phase in ReferenceAgent.observe() handles the actual learning. When a task includes a learning_signal containing a new value and version, the agent decides whether to update its memory based on the profile:

signal = task.get("learning_signal")
if signal and self.profile != "static":
    can_write = (self.profile == "evolving" and int(signal["version"]) > current.version) \
                or current is None
    if can_write:
        self.memory[task["rule_id"]] = MemoryEntry(signal["value"], int(signal["version"]))

Longitudinal Evaluation Framework

The LongitudinalEvaluator in chapter9/self-evolution-eval/harness.py tests these memory arms through a rigorous four-phase task stream: learning, transfer, rule-change, and retention. This measures not just immediate accuracy but how well agents adapt to shifting requirements while maintaining previously learned capabilities.

Each experiment run generates SHA-256 receipts logging every request, response, timestamp, and token usage, enabling reproducible audits of the self-evolution process.

Proposer-Reviewer Self-Evolution Loop (Hermes)

The second pattern moves beyond external memory into meta-cognitive self-improvement. Named Hermes after the messenger god, this implementation allows an agent to propose changes to its own source code, receive critique from a reviewer, and iterate based on feedback.

The Closed-Loop Improvement Cycle

Implemented in chapter9/hermes-self-evolution/run_experiment_8_6.py, this pattern mimics the scientific method: hypothesis (code change) → experiment (execution) → evaluation (review) → feedback (critique). The cycle continues until the reviewer accepts the patch or reaches a maximum iteration count.

The orchestration script follows this sequence:

  1. Propose: The LLM analyzes the entire codebase (including its own implementation) and generates a patch (hermes-self-evolution.patch)
  2. Review: A separate LLM instance evaluates the patch against acceptance criteria defined in acceptance_review.md
  3. Learn: If rejected, the review critique becomes a learning signal fed back into the next iteration
  4. Converge: The loop terminates when the reviewer accepts the patch or hits the max_rounds limit

Implementation Details

The Hermes agent receives both the book content and its own source code as context. Unlike the external-memory pattern which updates discrete rules, this pattern modifies the agent's actual implementation files, enabling architectural self-improvement rather than just behavioral adjustment.

The task.md file formalizes the "book-driven self-evolution" objective, while run_experiment_8_6.py handles the orchestration between proposer and reviewer models:


# Simplified usage of the Hermes pattern

run_self_evolution(
    model="gpt-5.6",            # LLM that edits its own code

    reviewer="gpt-5.6",         # Same model used as the reviewer

    max_rounds=10,              # Stop after 10 attempts if not accepted

)

Practical Implementation Examples

Beyond the conceptual framework, the repository provides runnable implementations demonstrating both patterns.

Using the evolving memory profile:

from chapter9.self_evolution_eval.agent import ReferenceAgent

agent = ReferenceAgent(profile="evolving")
task = {
    "id": "t1",
    "family": "baggage",
    "rule_id": "R1",
    "input": "How many kg can I check in?",
}

# First action uses baseline (memory is empty)

result = agent.act(task)

# Inject learning signal

task["learning_signal"] = {"value": "answer_23kg", "version": 2}
agent.observe(task)

# Subsequent actions use the learned rule

updated_result = agent.act(task)

Running full longitudinal experiments:


# Install dependencies

uv sync --locked --extra ch8

# Run four-phase evaluation

cd chapter9/self-evolution-eval
python run_experiment_8_7.py \
    --provider ark --model doubao-seed-1-6-250615 \
    --seeds 8601,8602,8603 --workers 6

Summary

  • Versioned external memory isolates learning from execution, preventing training data leakage into the model's reasoning process
  • Three update policies (static, append-only, evolving) provide experimental controls for measuring different learning rates and memory retention strategies
  • Longitudinal evaluation validates true self-evolution through multi-phase testing that includes rule changes and retention checks, not just one-shot accuracy
  • Proposer-reviewer loops enable code-level self-modification through iterative critique, implementing a scientific discovery cycle for architectural improvement
  • Reproducible auditing via SHA-256 receipts and harness logging ensures continuous learning systems remain observable and debuggable

Frequently Asked Questions

What is the difference between the external-memory and proposer-reviewer patterns?

The external-memory pattern stores behavioral rules outside the LLM weights, allowing the agent to update its responses without retraining the underlying model. It acts like a dynamic lookup table that evolves through observation. The proposer-reviewer pattern enables the agent to modify its own source code, effectively changing its architecture and reasoning process rather than just its knowledge base.

How does the "evolving" profile prevent catastrophic forgetting?

In chapter9/self-evolution-eval/agent.py, the evolving profile uses version numbers to manage rule updates. When a new learning_signal arrives, the observe() method only overwrites the existing memory entry if the incoming version is strictly greater than the current version. This preserves an audit trail while ensuring the agent always uses the most recent valid rule, allowing it to adapt to new requirements without losing the ability to track when and how its behavior changed.

Can the Hermes self-evolution loop modify any part of the codebase?

According to the implementation in run_experiment_8_6.py, the Hermes agent receives its own source code as context and can propose patches to any file it has access to. However, the terminal reviewer (defined in acceptance_review.md) validates these changes against safety and correctness criteria before they are applied. If rejected, the critique feeds back into the next proposal iteration, creating a safeguarded but autonomous improvement loop.

What metrics does the longitudinal evaluator track to measure self-evolution?

The LongitudinalEvaluator in harness.py tracks performance across four phases: initial learning, transfer to related tasks, adaptation to rule changes, and retention of previously learned behaviors. It logs accuracy per phase, token consumption, latency, and generates cryptographic receipts (SHA-256) for each interaction. This comprehensive measurement ensures that continuous learning doesn't come at the cost of stability or previously mastered capabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →