Implementing Agent Continuous Evolution from Runtime Trajectories: A Production-Ready Self-Modifying System
The bojieli/ai-agent-book repository implements a fully automated closed-loop system where AI agents evolve their own control logic by analyzing runtime trajectories, validating patches in sandboxed environments, and releasing changes through multi-gate canary deployments.
This article examines how production AI agents can achieve continuous evolution from runtime trajectories without human intervention. The implementation in chapter9/self-modifying-agent/ demonstrates a complete engineering pipeline—from failure detection to safe code deployment—built for real-world reliability and auditability.
Four Stages of the Self-Modification Pipeline
The core workflow operates as a tightly-coupled four-stage pipeline. Each stage produces structured artifacts that feed into the next, creating an auditable chain of evidence.
Stage 1: Trajectory Collection and Diagnosis
Runtime failures are captured as structured trajectory records that encode the full execution context.
import json
import pathlib
ROOT = pathlib.Path(__file__).parent
trajectories = json.loads((ROOT / "failure_trajectories.json").read_text())
from evolution import diagnose
diag = diagnose(trajectories)
print(diag["target"]) # → stable/retry_policy.py
The diagnose function (in chapter9/self-modifying-agent/evolution.py, lines 62-87) performs intelligent clustering on repeated non-retryable errors. It identifies:
- The target file requiring modification (
stable/retry_policy.py) - Tool-error pairs that trigger the failure pattern
- A change contract describing expected fixes and potential regressions
This diagnosis becomes the authoritative specification for all downstream generation and validation.
Stage 2: Deterministic Candidate Generation
Rather than using open-ended generation, the system produces minimal, deterministic patches through structured string replacement.
stable_src = (ROOT / "stable" / "retry_policy.py").read_text()
from evolution import generate_candidate
candidate = generate_candidate(stable_src, diag)
The generate_candidate function (evolution.py, lines 54-62 and 158-161) applies exact transformations:
| Pattern | Replacement | Purpose |
|---|---|---|
OLD_CODES |
NEW_CODES |
Update error code mappings |
OLD_RETRY |
NEW_RETRY |
Adjust retry policy constants |
OLD_BREAKER |
NEW_BREAKER |
Modify circuit breaker thresholds |
Each candidate includes provenance metadata via candidate_from_source, tracking the source diagnosis, version bump, and transformation lineage.
Stage 3: Sandboxed Multi-Gate Validation
No candidate touches production without surviving a hardened Docker sandbox with eleven sequential validation gates.
from evolution import validate_candidate
checks = validate_candidate(candidate["source"], trajectories)
print(checks) # All gates must be True
The validation sequence (evolution.py, lines 128-143) enforces increasing rigor:
- Static compile and AST security scan (
_safe_ast, lines 109-122) block syntax errors and forbidden constructs - Sandboxed execution (
run_in_sandbox) runs the candidate against original trajectories - Behavioral verification through nine specialized gates:
public_api_compatible— interface stabilityfailure_replay— original failures resolvednonretryable_circuit— proper error classificationtemporary_recovery— transient error handlingold_task_regression— backward compatibilitycanary_readyandrollback_ready— operational safety
Quantitative signals feed the canary decision through behavior_metrics (lines 46-53), extracting mean non-retryable calls, temporary-error recovery rates, and regression counts.
Stage 4: Release Manifest and Canary Deployment
Successful candidates produce a cryptographically verifiable release manifest:
from evolution import release_manifest
manifest = release_manifest(stable_src, candidate, diag, checks)
print(manifest["decision"]) # "release_to_canary" or "reject_candidate"
The manifest (evolution.py, lines 90-106) permanently records:
- Failure cluster metadata and source trajectories
- Inferential root cause and target component
- Complete diff with impact prediction
- SHA-256 hashes of stable, candidate, and rollback versions
- Canary eligibility flag and automated rollback conditions
Orthogonal Quality Control: The Trajectory Verifier
A separate deterministic verifier provides human-review gating for uncertain trajectories.
from chapter9.trajectory-verifier.verifier import TrajectoryVerifier
verifier = TrajectoryVerifier()
score, needs_review = verifier.assess(trajectory)
if needs_review:
route_to_human_queue(trajectory)
The TrajectoryVerifier (chapter9/trajectory-verifier/verifier.py, lines 73-100) supplies quality signals used by downstream experiments including Hermes self-evolution. High-risk failures or low-confidence diagnoses trigger mandatory human oversight before entering the learning pipeline.
Complete Integration Example
The following code demonstrates the full continuous evolution cycle:
import json
import pathlib
from evolution import diagnose, generate_candidate, validate_candidate
from evolution import behavior_metrics, release_manifest
ROOT = pathlib.Path(__file__).parent
# Load runtime evidence
trajectories = json.loads(
(ROOT / "failure_trajectories.json").read_text()
)
# Diagnose and generate
diag = diagnose(trajectories)
stable_src = (ROOT / "stable" / "retry_policy.py").read_text()
candidate = generate_candidate(stable_src, diag)
# Validate exhaustively
checks = validate_candidate(candidate["source"], trajectories)
assert all(checks.values()), "Validation failed"
# Quantify impact
metrics = behavior_metrics(candidate["source"], trajectories)
# Prepare production release
manifest = release_manifest(stable_src, candidate, diag, checks)
if manifest["decision"] == "release_to_canary":
deploy_to_canary(manifest)
else:
reject_and_queue_for_analysis(candidate)
Auditability and Trust Architecture
Every stage maintains cryptographic provenance:
- SHA-256 hashes of all code versions (stable, candidate, rollback)
- Version short-IDs linking to original evidence JSON
- Immutable change contracts specifying expected behavior
This prevents unauthorized modifications of the trusted root code and enables forensic reconstruction of any evolution decision.
Key Implementation Files
| Component | Path | Role |
|---|---|---|
| Core evolution logic | chapter9/self-modifying-agent/evolution.py |
Diagnosis, candidate generation, sandbox validation, manifest creation |
| Trajectory verification | chapter9/trajectory-verifier/verifier.py |
Deterministic quality scoring and human-review routing |
| Integration tests | chapter9/self-modifying-agent/test_evolution.py |
Full-pipeline validation |
| Sample trajectories | chapter9/self-modifying-agent/failure_trajectories.json |
Learning signal examples |
| Target control code | chapter9/self-modifying-agent/stable/retry_policy.py |
Production policy subject to modification |
| Sandbox execution | chapter9/self-modifying-agent/llm_generator.py |
Docker-based isolated execution |
| Experiment runner | chapter9/self-modifying-agent/run_experiment_9_6.py |
CLI for end-to-end cycles |
Summary
- Runtime trajectories become structured learning signals through JSON-based failure recording
- Deterministic patch generation replaces open-ended code synthesis with verified string transformations
- Eleven-gate sandbox validation ensures correctness, security, and backward compatibility before any production exposure
- Cryptographic manifests create permanent audit trails with rollback capabilities
- Orthogonal verification enables human oversight for uncertain evolutions
Frequently Asked Questions
How does the system prevent self-modifying agents from introducing security vulnerabilities?
The validate_candidate function employs defense in depth: static AST scanning (_safe_ast) blocks forbidden constructs before execution, the Docker sandbox isolates all candidate runtime with strict resource limits, and the security_scan gate enforces additional policy checks. According to the bojieli/ai-agent-book source code, no candidate reaches canary without passing all eleven validation gates including explicit security verification.
What makes this approach "continuous" rather than just automated patching?
The system closes a full feedback loop: runtime execution produces trajectories, trajectories trigger diagnosis, diagnosis generates patches, patches validate against historical data, and successful manifests deploy to canary environments with automated rollback. This runtime → collect → diagnose → generate → validate → canary → rollback → repeat cycle operates without human intervention, enabling true continuous evolution rather than one-off fixes.
When does the trajectory verifier route failures to human review?
The TrajectoryVerifier in verifier.py (lines 73-100) assesses trajectory reliability through multi-layer deterministic scoring. High-risk failure patterns, low-confidence diagnoses, or anomalous tool-error combinations return needs_review=True, triggering human queue routing. This safeguards the learning pipeline from noise and adversarial trajectories that could corrupt the agent's evolution.
Can the system handle multiple concurrent failure patterns?
The diagnose function clusters trajectories by repeated non-retryable errors, producing separate change contracts for each distinct root cause. Each cluster generates an independent candidate track with its own validation pipeline and manifest, enabling parallel evolution of multiple control policy components while maintaining isolation between unrelated changes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →