# Patterns for Agent Self-Evolution and Continuous Learning: A Deep Dive into the ai-agent-book Repository

> Explore agent self-evolution and continuous learning patterns. Discover versioned memory and a proposer-reviewer loop for iterative code improvement in the ai-agent-book repository.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-17

---

**The bojieli/ai-agent-book repository implements two core patterns for autonomous improvement: versioned external-memory arms with configurable update policies (static, append-only, evolving) and a closed-loop proposer-reviewer mechanism that enables agents to iteratively edit their own code based on critique feedback.**

This repository explores how autonomous agents can evolve their behavior over time without human intervention. Through practical implementations in `chapter9/self-evolution-eval/` and `chapter9/hermes-self-evolution/`, it demonstrates production-ready patterns for **agent self-evolution and continuous learning** that separate knowledge storage from reasoning, enabling systematic improvement from experience.

## External-Memory Arms Pattern

The first pattern treats agent memory as a versioned rule store that operates independently of the LLM's weights. This architecture allows agents to adapt to new information while preserving audit trails of how their behavior changed.

### Architecture and Memory Model

At the heart of this pattern lies a dictionary-based memory system defined in [`chapter9/self-evolution-eval/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/self-evolution-eval/agent.py). The `ReferenceAgent` and `OpenAILongitudinalAgent` classes maintain a `memory: Dict[str, MemoryEntry]` structure where each entry stores both a value and a version number.

```python

# Memory structure from agent.py

memory = {
    "rule_id": MemoryEntry(value="answer_23kg", version=2)
}

```

The repository implements three distinct profiles that control how this memory evolves:

- **Static**: Memory is read-only; the agent never updates its rule store
- **Append-only**: Records every observation but preserves the first version permanently
- **Evolving**: Replaces older versions when a higher version number appears, maintaining audit trails while allowing updates

### Action and Observation Flow

The separation of **acting** and **observing** creates a clean interface for learning. When the agent acts, it queries memory; when it observes, it processes learning signals.

During the action phase in `ReferenceAgent.act()`, the agent looks up rules by `rule_id`. If no entry exists and the profile isn't static, it falls back to `BASELINE_ACTIONS` indexed by task family:

```python

# From chapter9/self-evolution-eval/agent.py

entry = self.memory.get(task["rule_id"])
used_memory = entry is not None and self.profile != "static"
action = entry.value if used_memory else BASELINE_ACTIONS[task["family"]]

```

The observation phase in `ReferenceAgent.observe()` handles the actual learning. When a task includes a `learning_signal` containing a new value and version, the agent decides whether to update its memory based on the profile:

```python
signal = task.get("learning_signal")
if signal and self.profile != "static":
    can_write = (self.profile == "evolving" and int(signal["version"]) > current.version) \
                or current is None
    if can_write:
        self.memory[task["rule_id"]] = MemoryEntry(signal["value"], int(signal["version"]))

```

### Longitudinal Evaluation Framework

The `LongitudinalEvaluator` in [`chapter9/self-evolution-eval/harness.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/self-evolution-eval/harness.py) tests these memory arms through a rigorous **four-phase task stream**: learning, transfer, rule-change, and retention. This measures not just immediate accuracy but how well agents adapt to shifting requirements while maintaining previously learned capabilities.

Each experiment run generates SHA-256 receipts logging every request, response, timestamp, and token usage, enabling reproducible audits of the self-evolution process.

## Proposer-Reviewer Self-Evolution Loop (Hermes)

The second pattern moves beyond external memory into meta-cognitive self-improvement. Named Hermes after the messenger god, this implementation allows an agent to propose changes to its own source code, receive critique from a reviewer, and iterate based on feedback.

### The Closed-Loop Improvement Cycle

Implemented in [`chapter9/hermes-self-evolution/run_experiment_8_6.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/hermes-self-evolution/run_experiment_8_6.py), this pattern mimics the scientific method: hypothesis (code change) → experiment (execution) → evaluation (review) → feedback (critique). The cycle continues until the reviewer accepts the patch or reaches a maximum iteration count.

The orchestration script follows this sequence:

1. **Propose**: The LLM analyzes the entire codebase (including its own implementation) and generates a patch (`hermes-self-evolution.patch`)
2. **Review**: A separate LLM instance evaluates the patch against acceptance criteria defined in [`acceptance_review.md`](https://github.com/bojieli/ai-agent-book/blob/main/acceptance_review.md)
3. **Learn**: If rejected, the review critique becomes a learning signal fed back into the next iteration
4. **Converge**: The loop terminates when the reviewer accepts the patch or hits the `max_rounds` limit

### Implementation Details

The Hermes agent receives both the book content and its own source code as context. Unlike the external-memory pattern which updates discrete rules, this pattern modifies the agent's actual implementation files, enabling architectural self-improvement rather than just behavioral adjustment.

The [`task.md`](https://github.com/bojieli/ai-agent-book/blob/main/task.md) file formalizes the "book-driven self-evolution" objective, while [`run_experiment_8_6.py`](https://github.com/bojieli/ai-agent-book/blob/main/run_experiment_8_6.py) handles the orchestration between proposer and reviewer models:

```python

# Simplified usage of the Hermes pattern

run_self_evolution(
    model="gpt-5.6",            # LLM that edits its own code

    reviewer="gpt-5.6",         # Same model used as the reviewer

    max_rounds=10,              # Stop after 10 attempts if not accepted

)

```

## Practical Implementation Examples

Beyond the conceptual framework, the repository provides runnable implementations demonstrating both patterns.

Using the evolving memory profile:

```python
from chapter9.self_evolution_eval.agent import ReferenceAgent

agent = ReferenceAgent(profile="evolving")
task = {
    "id": "t1",
    "family": "baggage",
    "rule_id": "R1",
    "input": "How many kg can I check in?",
}

# First action uses baseline (memory is empty)

result = agent.act(task)

# Inject learning signal

task["learning_signal"] = {"value": "answer_23kg", "version": 2}
agent.observe(task)

# Subsequent actions use the learned rule

updated_result = agent.act(task)

```

Running full longitudinal experiments:

```bash

# Install dependencies

uv sync --locked --extra ch8

# Run four-phase evaluation

cd chapter9/self-evolution-eval
python run_experiment_8_7.py \
    --provider ark --model doubao-seed-1-6-250615 \
    --seeds 8601,8602,8603 --workers 6

```

## Summary

- **Versioned external memory** isolates learning from execution, preventing training data leakage into the model's reasoning process
- **Three update policies** (static, append-only, evolving) provide experimental controls for measuring different learning rates and memory retention strategies
- **Longitudinal evaluation** validates true self-evolution through multi-phase testing that includes rule changes and retention checks, not just one-shot accuracy
- **Proposer-reviewer loops** enable code-level self-modification through iterative critique, implementing a scientific discovery cycle for architectural improvement
- **Reproducible auditing** via SHA-256 receipts and harness logging ensures continuous learning systems remain observable and debuggable

## Frequently Asked Questions

### What is the difference between the external-memory and proposer-reviewer patterns?

The **external-memory pattern** stores behavioral rules outside the LLM weights, allowing the agent to update its responses without retraining the underlying model. It acts like a dynamic lookup table that evolves through observation. The **proposer-reviewer pattern** enables the agent to modify its own source code, effectively changing its architecture and reasoning process rather than just its knowledge base.

### How does the "evolving" profile prevent catastrophic forgetting?

In [`chapter9/self-evolution-eval/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/self-evolution-eval/agent.py), the evolving profile uses version numbers to manage rule updates. When a new `learning_signal` arrives, the `observe()` method only overwrites the existing memory entry if the incoming version is strictly greater than the current version. This preserves an audit trail while ensuring the agent always uses the most recent valid rule, allowing it to adapt to new requirements without losing the ability to track when and how its behavior changed.

### Can the Hermes self-evolution loop modify any part of the codebase?

According to the implementation in [`run_experiment_8_6.py`](https://github.com/bojieli/ai-agent-book/blob/main/run_experiment_8_6.py), the Hermes agent receives its own source code as context and can propose patches to any file it has access to. However, the terminal reviewer (defined in [`acceptance_review.md`](https://github.com/bojieli/ai-agent-book/blob/main/acceptance_review.md)) validates these changes against safety and correctness criteria before they are applied. If rejected, the critique feeds back into the next proposal iteration, creating a safeguarded but autonomous improvement loop.

### What metrics does the longitudinal evaluator track to measure self-evolution?

The `LongitudinalEvaluator` in [`harness.py`](https://github.com/bojieli/ai-agent-book/blob/main/harness.py) tracks performance across four phases: initial learning, transfer to related tasks, adaptation to rule changes, and retention of previously learned behaviors. It logs accuracy per phase, token consumption, latency, and generates cryptographic receipts (SHA-256) for each interaction. This comprehensive measurement ensures that **continuous learning** doesn't come at the cost of stability or previously mastered capabilities.