Agent Continuous Evolution from Run Trajectories: 7 Techniques for Self-Improving AI Systems

AI agents can evolve automatically by analyzing their own execution traces, diagnosing failures, and generating validated patches—creating a closed-loop system that improves without human intervention.

This guide examines the agent continuous evolution pipeline implemented in bojieli/ai-agent-book, a reference implementation showing how production agents learn from run trajectories (execution traces) to iteratively upgrade their control logic. The system combines deterministic rules, LLM-driven code generation, and multi-layer safety gates to ensure reliable self-modification.

Trajectory Diagnosis: Finding Root Causes in Failure Patterns

The evolution cycle begins with trajectory diagnosis, which scans failure trajectories to identify repeated, non-retryable error patterns.

In evolution.py → diagnose() (lines 62-105), the system:

  • Parses JSON trajectory logs from previous agent runs
  • Clusters similar failure signatures
  • Extracts the root-cause location (e.g., the retry-policy module)
  • Outputs a structured diagnosis with change_required: bool

Only data-backed patterns trigger evolution—preventing spurious changes from noise.

Candidate Generation: Two Complementary Approaches

The repository supports both deterministic and LLM-driven patch generation, selectable based on failure complexity.

Deterministic Candidate Generation

For well-understood failure modes, evolution.py → generate_candidate() (lines 54-71) applies textual replacements:

  • Adding new error codes to handler lists
  • Adjusting retry count thresholds
  • Bumping version strings

This path is fast, reproducible, and requires no external API calls.

LLM-Driven Candidate Generation

For novel or complex failures, llm_generator.py → generate_with_openai() (lines 66-84) constructs a detailed prompt for OpenAI, OpenRouter, or Ark models. The LLM returns:

  • A rewritten module source
  • An impact prediction describing expected behavioral changes

The response is parsed into a candidate object via candidate_from_source().

Sandboxed Validation: Multi-Layer Safety Gates

Before any candidate reaches production, evolution.py → validate_candidate() (lines 102-124) executes a sandboxed validation sequence:

  1. Static compilation — ensures syntactically valid Python
  2. Security AST scan — detects dangerous imports or eval usage
  3. Docker sandbox execution — runs candidate against historical trajectories
  4. API compatibility check — verifies interface contracts

The result is a checks dictionary with boolean gate results—every gate must pass for progression.

Behavior Metric Collection: Measuring Real Impact

Validation alone isn't sufficient. evolution.py → behavior_metrics() (lines 46-66) executes the candidate to collect concrete performance data:

  • Mean non-retryable calls — efficiency improvement indicator
  • Temporary-error recovery rate — resilience metric
  • Old-task regressions — backward compatibility score

These metrics enable data-driven release decisions, not just structural correctness.

Release Manifest & Deployment Gates

The evolution.py → release_manifest() (lines 68-106) packages all evidence into an auditable artifact containing:

Field Purpose
diff Unified diff against stable code
impact_prediction LLM-generated or rule-based change description
validation_results Boolean gate outcomes
behavior_metrics Quantified performance deltas
provenance SHA-256 hashes and timestamps
decision "release_to_canary" or "reject_candidate"

This manifest drives canary deployment and automated rollback workflows.

The Continuous Loop: Closing the Feedback Cycle

The entire pipeline reruns automatically after each experiment batch. As shown in demo.py (lines 15-32):


# Load fresh trajectories from recent agent runs

trajectories = load_trajectories("failure_trajectories.json")

# Diagnose → generate → validate → decide

diagnosis = diagnose(trajectories)
candidate = generate_candidate(stable_source, diagnosis)
checks = validate_candidate(candidate["source"], trajectories)
metrics = behavior_metrics(candidate["source"], trajectories)
manifest = release_manifest(stable_source, candidate, diagnosis, checks)

if manifest["decision"] == "release_to_canary":
    deploy_to_canary(candidate)
else:
    log_rejection(manifest["rejection_reason"])

New trajectories feed new diagnoses, creating agent continuous evolution without human-written patches.

End-to-End Implementation Example

from pathlib import Path
import json
from evolution import diagnose, generate_candidate, validate_candidate
from evolution import behavior_metrics, release_manifest
from llm_generator import generate_with_openai

# 1. Load execution trajectories

traj_path = Path("chapter8/self-modifying-agent/failure_trajectories.json")
trajectories = json.loads(traj_path.read_text(encoding="utf-8"))

# 2. Diagnose root cause

diagnosis = diagnose(trajectories)
if not diagnosis["change_required"]:
    print("No evolution triggered")
    exit()

# 3. Generate patch (deterministic or LLM-based)

stable_path = Path("chapter8/self-modifying-agent/stable/retry_policy.py")
stable_source = stable_path.read_text(encoding="utf-8")

# Option A: Rule-based generation

candidate = generate_candidate(stable_source, diagnosis)

# Option B: LLM-based generation (uncomment to use)

# candidate = generate_with_openai(

#     stable_source, diagnosis, model="gpt-4o-mini"

# )

# 4. Validate and measure

checks = validate_candidate(candidate["source"], trajectories)
metrics = behavior_metrics(candidate["source"], trajectories)

# 5. Build release decision

manifest = release_manifest(stable_source, candidate, diagnosis, checks)
print(f"Decision: {manifest['decision']}")

Summary

  • Trajectory diagnosis in evolution.py → diagnose() identifies root-cause locations from failure patterns.
  • Dual generation paths support both deterministic patches (generate_candidate()) and LLM-driven rewrites (llm_generator.py).
  • Sandboxed validation combines static analysis, security scanning, and Docker isolation before any code deployment.
  • Behavior metrics quantify real performance impact beyond structural correctness.
  • Release manifests create auditable, hash-verified records enabling canary rollouts and automatic rollback.
  • Continuous loop execution in demo.py enables fully autonomous agent improvement from run trajectories.

Frequently Asked Questions

What is a "run trajectory" in agent evolution?

A run trajectory is a structured execution trace—typically JSON lines—recording an agent's step-by-step behavior, tool calls, errors, and outcomes during task execution. The ai-agent-book repository uses these trajectories as the input signal for detecting failure patterns and triggering evolutionary improvements.

How does the system prevent malicious self-modifications?

Multiple safety layers enforce boundaries: AST security scans in validate_candidate() detect dangerous code patterns; Docker sandboxing isolates candidate execution; API compatibility checks ensure interface contracts hold; and behavior metrics verify functional equivalence. No candidate reaches production without passing all gates.

Can this pipeline work without cloud LLM APIs?

Yes. The deterministic candidate generation path in generate_candidate() operates entirely locally using textual replacements and rule-based transformations. LLM generation is optional for novel failures requiring semantic understanding beyond predefined patterns.

What's the difference between validation and behavior metrics?

Validation (validate_candidate()) checks structural and security properties: does the code compile? Is it safe to execute? Behavior metrics (behavior_metrics()) measure functional outcomes: does the candidate actually reduce errors? Does it maintain performance on previously-solved tasks? Both are required for release approval.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →