Implementing Agent Continuous Evolution from Runtime Trajectories: A Production-Ready Self-Modifying System

The bojieli/ai-agent-book repository implements a fully automated closed-loop system where AI agents evolve their own control logic by analyzing runtime trajectories, validating patches in sandboxed environments, and releasing changes through multi-gate canary deployments.

This article examines how production AI agents can achieve continuous evolution from runtime trajectories without human intervention. The implementation in chapter9/self-modifying-agent/ demonstrates a complete engineering pipeline—from failure detection to safe code deployment—built for real-world reliability and auditability.

Four Stages of the Self-Modification Pipeline

The core workflow operates as a tightly-coupled four-stage pipeline. Each stage produces structured artifacts that feed into the next, creating an auditable chain of evidence.

Stage 1: Trajectory Collection and Diagnosis

Runtime failures are captured as structured trajectory records that encode the full execution context.

import json
import pathlib

ROOT = pathlib.Path(__file__).parent
trajectories = json.loads((ROOT / "failure_trajectories.json").read_text())

from evolution import diagnose
diag = diagnose(trajectories)
print(diag["target"])  # → stable/retry_policy.py

The diagnose function (in chapter9/self-modifying-agent/evolution.py, lines 62-87) performs intelligent clustering on repeated non-retryable errors. It identifies:

  • The target file requiring modification (stable/retry_policy.py)
  • Tool-error pairs that trigger the failure pattern
  • A change contract describing expected fixes and potential regressions

This diagnosis becomes the authoritative specification for all downstream generation and validation.

Stage 2: Deterministic Candidate Generation

Rather than using open-ended generation, the system produces minimal, deterministic patches through structured string replacement.

stable_src = (ROOT / "stable" / "retry_policy.py").read_text()

from evolution import generate_candidate
candidate = generate_candidate(stable_src, diag)

The generate_candidate function (evolution.py, lines 54-62 and 158-161) applies exact transformations:

Pattern Replacement Purpose
OLD_CODES NEW_CODES Update error code mappings
OLD_RETRY NEW_RETRY Adjust retry policy constants
OLD_BREAKER NEW_BREAKER Modify circuit breaker thresholds

Each candidate includes provenance metadata via candidate_from_source, tracking the source diagnosis, version bump, and transformation lineage.

Stage 3: Sandboxed Multi-Gate Validation

No candidate touches production without surviving a hardened Docker sandbox with eleven sequential validation gates.

from evolution import validate_candidate
checks = validate_candidate(candidate["source"], trajectories)
print(checks)  # All gates must be True

The validation sequence (evolution.py, lines 128-143) enforces increasing rigor:

  1. Static compile and AST security scan (_safe_ast, lines 109-122) block syntax errors and forbidden constructs
  2. Sandboxed execution (run_in_sandbox) runs the candidate against original trajectories
  3. Behavioral verification through nine specialized gates:
    • public_api_compatible — interface stability
    • failure_replay — original failures resolved
    • nonretryable_circuit — proper error classification
    • temporary_recovery — transient error handling
    • old_task_regression — backward compatibility
    • canary_ready and rollback_ready — operational safety

Quantitative signals feed the canary decision through behavior_metrics (lines 46-53), extracting mean non-retryable calls, temporary-error recovery rates, and regression counts.

Stage 4: Release Manifest and Canary Deployment

Successful candidates produce a cryptographically verifiable release manifest:

from evolution import release_manifest

manifest = release_manifest(stable_src, candidate, diag, checks)
print(manifest["decision"])  # "release_to_canary" or "reject_candidate"

The manifest (evolution.py, lines 90-106) permanently records:

  • Failure cluster metadata and source trajectories
  • Inferential root cause and target component
  • Complete diff with impact prediction
  • SHA-256 hashes of stable, candidate, and rollback versions
  • Canary eligibility flag and automated rollback conditions

Orthogonal Quality Control: The Trajectory Verifier

A separate deterministic verifier provides human-review gating for uncertain trajectories.

from chapter9.trajectory-verifier.verifier import TrajectoryVerifier

verifier = TrajectoryVerifier()
score, needs_review = verifier.assess(trajectory)
if needs_review:
    route_to_human_queue(trajectory)

The TrajectoryVerifier (chapter9/trajectory-verifier/verifier.py, lines 73-100) supplies quality signals used by downstream experiments including Hermes self-evolution. High-risk failures or low-confidence diagnoses trigger mandatory human oversight before entering the learning pipeline.

Complete Integration Example

The following code demonstrates the full continuous evolution cycle:

import json
import pathlib
from evolution import diagnose, generate_candidate, validate_candidate
from evolution import behavior_metrics, release_manifest

ROOT = pathlib.Path(__file__).parent

# Load runtime evidence

trajectories = json.loads(
    (ROOT / "failure_trajectories.json").read_text()
)

# Diagnose and generate

diag = diagnose(trajectories)
stable_src = (ROOT / "stable" / "retry_policy.py").read_text()
candidate = generate_candidate(stable_src, diag)

# Validate exhaustively

checks = validate_candidate(candidate["source"], trajectories)
assert all(checks.values()), "Validation failed"

# Quantify impact

metrics = behavior_metrics(candidate["source"], trajectories)

# Prepare production release

manifest = release_manifest(stable_src, candidate, diag, checks)

if manifest["decision"] == "release_to_canary":
    deploy_to_canary(manifest)
else:
    reject_and_queue_for_analysis(candidate)

Auditability and Trust Architecture

Every stage maintains cryptographic provenance:

  • SHA-256 hashes of all code versions (stable, candidate, rollback)
  • Version short-IDs linking to original evidence JSON
  • Immutable change contracts specifying expected behavior

This prevents unauthorized modifications of the trusted root code and enables forensic reconstruction of any evolution decision.

Key Implementation Files

Component Path Role
Core evolution logic chapter9/self-modifying-agent/evolution.py Diagnosis, candidate generation, sandbox validation, manifest creation
Trajectory verification chapter9/trajectory-verifier/verifier.py Deterministic quality scoring and human-review routing
Integration tests chapter9/self-modifying-agent/test_evolution.py Full-pipeline validation
Sample trajectories chapter9/self-modifying-agent/failure_trajectories.json Learning signal examples
Target control code chapter9/self-modifying-agent/stable/retry_policy.py Production policy subject to modification
Sandbox execution chapter9/self-modifying-agent/llm_generator.py Docker-based isolated execution
Experiment runner chapter9/self-modifying-agent/run_experiment_9_6.py CLI for end-to-end cycles

Summary

  • Runtime trajectories become structured learning signals through JSON-based failure recording
  • Deterministic patch generation replaces open-ended code synthesis with verified string transformations
  • Eleven-gate sandbox validation ensures correctness, security, and backward compatibility before any production exposure
  • Cryptographic manifests create permanent audit trails with rollback capabilities
  • Orthogonal verification enables human oversight for uncertain evolutions

Frequently Asked Questions

How does the system prevent self-modifying agents from introducing security vulnerabilities?

The validate_candidate function employs defense in depth: static AST scanning (_safe_ast) blocks forbidden constructs before execution, the Docker sandbox isolates all candidate runtime with strict resource limits, and the security_scan gate enforces additional policy checks. According to the bojieli/ai-agent-book source code, no candidate reaches canary without passing all eleven validation gates including explicit security verification.

What makes this approach "continuous" rather than just automated patching?

The system closes a full feedback loop: runtime execution produces trajectories, trajectories trigger diagnosis, diagnosis generates patches, patches validate against historical data, and successful manifests deploy to canary environments with automated rollback. This runtime → collect → diagnose → generate → validate → canary → rollback → repeat cycle operates without human intervention, enabling true continuous evolution rather than one-off fixes.

When does the trajectory verifier route failures to human review?

The TrajectoryVerifier in verifier.py (lines 73-100) assesses trajectory reliability through multi-layer deterministic scoring. High-risk failure patterns, low-confidence diagnoses, or anomalous tool-error combinations return needs_review=True, triggering human queue routing. This safeguards the learning pipeline from noise and adversarial trajectories that could corrupt the agent's evolution.

Can the system handle multiple concurrent failure patterns?

The diagnose function clusters trajectories by repeated non-retryable errors, producing separate change contracts for each distinct root cause. Each cluster generates an independent candidate track with its own validation pipeline and manifest, enabling parallel evolution of multiple control policy components while maintaining isolation between unrelated changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →