# Agent Continuous Evolution from Run Trajectories: 7 Techniques for Self-Improving AI Systems

> Discover 7 techniques for AI agent continuous evolution. Learn how agents analyze run trajectories to self-improve without human intervention. Boost your AI system's performance.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-06

---

**AI agents can evolve automatically by analyzing their own execution traces, diagnosing failures, and generating validated patches—creating a closed-loop system that improves without human intervention.**

This guide examines the **agent continuous evolution** pipeline implemented in `bojieli/ai-agent-book`, a reference implementation showing how production agents learn from **run trajectories** (execution traces) to iteratively upgrade their control logic. The system combines deterministic rules, LLM-driven code generation, and multi-layer safety gates to ensure reliable self-modification.

## Trajectory Diagnosis: Finding Root Causes in Failure Patterns

The evolution cycle begins with **trajectory diagnosis**, which scans failure trajectories to identify repeated, non-retryable error patterns.

In `evolution.py → diagnose()` (lines 62-105), the system:

- Parses JSON trajectory logs from previous agent runs
- Clusters similar failure signatures
- Extracts the root-cause location (e.g., the retry-policy module)
- Outputs a structured diagnosis with `change_required: bool`

Only data-backed patterns trigger evolution—preventing spurious changes from noise.

## Candidate Generation: Two Complementary Approaches

The repository supports both **deterministic** and **LLM-driven** patch generation, selectable based on failure complexity.

### Deterministic Candidate Generation

For well-understood failure modes, `evolution.py → generate_candidate()` (lines 54-71) applies textual replacements:

- Adding new error codes to handler lists
- Adjusting retry count thresholds
- Bumping version strings

This path is fast, reproducible, and requires no external API calls.

### LLM-Driven Candidate Generation

For novel or complex failures, `llm_generator.py → generate_with_openai()` (lines 66-84) constructs a detailed prompt for OpenAI, OpenRouter, or Ark models. The LLM returns:

- A rewritten module source
- An **impact prediction** describing expected behavioral changes

The response is parsed into a candidate object via `candidate_from_source()`.

## Sandboxed Validation: Multi-Layer Safety Gates

Before any candidate reaches production, `evolution.py → validate_candidate()` (lines 102-124) executes a **sandboxed validation** sequence:

1. **Static compilation** — ensures syntactically valid Python
2. **Security AST scan** — detects dangerous imports or eval usage
3. **Docker sandbox execution** — runs candidate against historical trajectories
4. **API compatibility check** — verifies interface contracts

The result is a `checks` dictionary with boolean gate results—every gate must pass for progression.

## Behavior Metric Collection: Measuring Real Impact

Validation alone isn't sufficient. `evolution.py → behavior_metrics()` (lines 46-66) executes the candidate to collect concrete performance data:

- **Mean non-retryable calls** — efficiency improvement indicator
- **Temporary-error recovery rate** — resilience metric
- **Old-task regressions** — backward compatibility score

These metrics enable data-driven release decisions, not just structural correctness.

## Release Manifest & Deployment Gates

The `evolution.py → release_manifest()` (lines 68-106) packages all evidence into an auditable artifact containing:

| Field | Purpose |
|-------|---------|
| `diff` | Unified diff against stable code |
| `impact_prediction` | LLM-generated or rule-based change description |
| `validation_results` | Boolean gate outcomes |
| `behavior_metrics` | Quantified performance deltas |
| `provenance` | SHA-256 hashes and timestamps |
| `decision` | `"release_to_canary"` or `"reject_candidate"` |

This manifest drives **canary deployment** and **automated rollback** workflows.

## The Continuous Loop: Closing the Feedback Cycle

The entire pipeline reruns automatically after each experiment batch. As shown in [`demo.py`](https://github.com/bojieli/ai-agent-book/blob/main/demo.py) (lines 15-32):

```python

# Load fresh trajectories from recent agent runs

trajectories = load_trajectories("failure_trajectories.json")

# Diagnose → generate → validate → decide

diagnosis = diagnose(trajectories)
candidate = generate_candidate(stable_source, diagnosis)
checks = validate_candidate(candidate["source"], trajectories)
metrics = behavior_metrics(candidate["source"], trajectories)
manifest = release_manifest(stable_source, candidate, diagnosis, checks)

if manifest["decision"] == "release_to_canary":
    deploy_to_canary(candidate)
else:
    log_rejection(manifest["rejection_reason"])

```

New trajectories feed new diagnoses, creating **agent continuous evolution** without human-written patches.

## End-to-End Implementation Example

```python
from pathlib import Path
import json
from evolution import diagnose, generate_candidate, validate_candidate
from evolution import behavior_metrics, release_manifest
from llm_generator import generate_with_openai

# 1. Load execution trajectories

traj_path = Path("chapter8/self-modifying-agent/failure_trajectories.json")
trajectories = json.loads(traj_path.read_text(encoding="utf-8"))

# 2. Diagnose root cause

diagnosis = diagnose(trajectories)
if not diagnosis["change_required"]:
    print("No evolution triggered")
    exit()

# 3. Generate patch (deterministic or LLM-based)

stable_path = Path("chapter8/self-modifying-agent/stable/retry_policy.py")
stable_source = stable_path.read_text(encoding="utf-8")

# Option A: Rule-based generation

candidate = generate_candidate(stable_source, diagnosis)

# Option B: LLM-based generation (uncomment to use)

# candidate = generate_with_openai(

#     stable_source, diagnosis, model="gpt-4o-mini"

# )

# 4. Validate and measure

checks = validate_candidate(candidate["source"], trajectories)
metrics = behavior_metrics(candidate["source"], trajectories)

# 5. Build release decision

manifest = release_manifest(stable_source, candidate, diagnosis, checks)
print(f"Decision: {manifest['decision']}")

```

## Summary

- **Trajectory diagnosis** in `evolution.py → diagnose()` identifies root-cause locations from failure patterns.
- **Dual generation paths** support both deterministic patches (`generate_candidate()`) and LLM-driven rewrites ([`llm_generator.py`](https://github.com/bojieli/ai-agent-book/blob/main/llm_generator.py)).
- **Sandboxed validation** combines static analysis, security scanning, and Docker isolation before any code deployment.
- **Behavior metrics** quantify real performance impact beyond structural correctness.
- **Release manifests** create auditable, hash-verified records enabling canary rollouts and automatic rollback.
- **Continuous loop execution** in [`demo.py`](https://github.com/bojieli/ai-agent-book/blob/main/demo.py) enables fully autonomous agent improvement from run trajectories.

## Frequently Asked Questions

### What is a "run trajectory" in agent evolution?

A run trajectory is a structured execution trace—typically JSON lines—recording an agent's step-by-step behavior, tool calls, errors, and outcomes during task execution. The `ai-agent-book` repository uses these trajectories as the input signal for detecting failure patterns and triggering evolutionary improvements.

### How does the system prevent malicious self-modifications?

Multiple safety layers enforce boundaries: **AST security scans** in `validate_candidate()` detect dangerous code patterns; **Docker sandboxing** isolates candidate execution; **API compatibility checks** ensure interface contracts hold; and **behavior metrics** verify functional equivalence. No candidate reaches production without passing all gates.

### Can this pipeline work without cloud LLM APIs?

Yes. The **deterministic candidate generation** path in `generate_candidate()` operates entirely locally using textual replacements and rule-based transformations. LLM generation is optional for novel failures requiring semantic understanding beyond predefined patterns.

### What's the difference between validation and behavior metrics?

**Validation** (`validate_candidate()`) checks structural and security properties: does the code compile? Is it safe to execute? **Behavior metrics** (`behavior_metrics()`) measure functional outcomes: does the candidate actually reduce errors? Does it maintain performance on previously-solved tasks? Both are required for release approval.