How context.md Tracks Generation History and Improvement Rationale in SIA
The context.md file in SIA serves as a comprehensive, chronological log that records every agent generation, its quantitative performance metrics, and the specific rationale for code changes, managed entirely by the ContextManager class in sia/context_manager.py.
The hexo-ai/sia repository implements a self-documenting agent evolution system where context.md functions as the central audit trail. This markdown file captures not only what changed between successive generations of an agent, but why those changes were proposed and how they impacted performance outcomes.
The Three-Phase Logging Architecture
The ContextManager orchestrates the creation and maintenance of context.md through three distinct phases, each implemented as a specific method call in sia/context_manager.py.
Phase 1: Header Initialization
When a run begins, ContextManager.initialize() establishes the document structure by writing a header section that identifies the experimental run. According to the source code (lines 48-64), this header includes the task directory, meta-model and task-model names, the agent implementation type, the start timestamp, and the maximum number of generations allowed.
This initialization creates the foundation for the chronological narrative that follows.
Phase 2: Per-Generation Entry Logging
After each generation completes, ContextManager.add_generation() appends a detailed markdown section to context.md. This method aggregates multiple data sources to create a comprehensive snapshot:
- Agent Statistics: File size and line count are captured via
_get_agent_stats()(lines 27-33), providing baseline metrics for code complexity. - Delta Calculations: The system calculates percentage changes in file size and line count compared to the previous generation (lines 28-36), quantifying the scope of modifications.
- Performance Metrics: Raw results are extracted from
results.json,detailed_results.json, or log files through_extract_metrics()(lines 38-84), capturing accuracy, execution time, and task-specific outcomes. - Improvement Rationale: The system parses
improvement.mdvia_extract_insights()(lines 25-52) to capture the LLM-generated reasoning behind proposed changes. - Evolution Summary: A concise narrative comparing previous and current agent code is generated by
_generate_llm_summary()(lines 66-78, 114-140), synthesizing the improvement plan and metric changes into human-readable insights.
All components are assembled by _format_generation_entry() (lines 54-59) into a standardized markdown block separated by --- delimiters, ensuring clear visual separation between generations.
Phase 3: Final Summary Statistics
When the run concludes, ContextManager.finalize() appends a Summary Statistics section (lines 67-108) that aggregates the entire evolution history. This final block reports total generations executed, success counts, identification of the best-performing generation, overall accuracy evolution trends, and code-size growth patterns across the entire run.
Key Technical Components
Tracking Metric Deltas
The logging system automatically calculates and displays Changes vs Previous Generation tables (lines 28-55). These tables highlight metric deltas, making it immediately apparent whether a code modification improved or degraded performance, and by what magnitude.
Linking Code to Rationale
By extracting content from improvement.md through _extract_insights(), context.md maintains a tight coupling between quantitative results and qualitative reasoning. This allows developers to trace specific performance changes back to the specific improvement hypotheses that generated them.
Implementation Example
The following implementation demonstrates the complete lifecycle of context.md generation:
from sia.context_manager import ContextManager
from sia.config import Config
# Initialize the manager for a run directory
run_dir = "runs/run_1"
run_cfg = {
"task_dir": "tasks/example",
"meta_model": "claude-2",
"task_model": "gpt-4",
"agent_impl": "claude",
"max_gen": 5,
}
ctx = ContextManager(run_dir, run_cfg, Config())
ctx.initialize() # Writes the header to context.md
# After each generation, add an entry
gen_data = {
"success": True,
"timestamp": "2026-06-12 12:34:56",
"duration": 42.7,
"agent_path": f"{run_dir}/gen_1/target_agent.py",
"gen_dir": f"{run_dir}/gen_1",
"improvement_path": f"{run_dir}/gen_1/improvement.md",
"execution_type": "Single",
}
ctx.add_generation(gen_num=1, gen_data=gen_data) # Appends Generation-1 block
# When the run finishes
ctx.finalize() # Appends Summary Statistics block
This workflow mirrors the patterns validated in tests/test_context_manager.py (lines 38-84), which verifies the full lifecycle from initialization through finalization.
Summary
context.mdserves as the central chronological record of agent evolution in the SIA framework, combining quantitative metrics with qualitative improvement rationale.- The
ContextManagerclass insia/context_manager.pyorchestrates three distinct phases: header initialization, per-generation entry logging, and final summary statistics. - Each generation entry captures agent statistics, performance deltas, LLM-generated improvement insights, and evolution summaries to create a self-documenting audit trail.
- The system automatically extracts metrics from JSON result files and improvement rationale from
improvement.md, linking code changes to their intended outcomes. - Final summary statistics provide aggregate views of success rates, best-performing generations, and code growth patterns across the entire experimental run.
Frequently Asked Questions
What file format does context.md use?
The context.md file uses standard markdown format with ATX-style headers, bullet lists, and tables. This choice ensures the generation history remains human-readable while being parseable by standard markdown processors. Each generation entry is separated by --- horizontal rules to demarcate distinct iterations clearly.
How does SIA extract improvement rationale for the log?
The ContextManager extracts improvement rationale through the _extract_insights() method (lines 25-52), which parses the optional improvement.md file generated during each iteration. This file contains the LLM's reasoning for proposed code changes, which the context manager then embeds directly into the corresponding generation entry, creating a direct link between the justification and the resulting metrics.
Where is the ContextManager implementation tested?
The core functionality is validated in tests/test_context_manager.py (lines 38-84), which exercises the complete lifecycle including initialization, generation entry addition, and finalization. Additionally, tests/golden/feedback_context_success_single.txt provides a reference example of a fully rendered context.md entry showing the expected format for metrics, insights, and LLM summaries.
Can context.md track multiple generations with different execution types?
Yes, the add_generation() method accepts a gen_data dictionary that includes an execution_type field (e.g., "Single"), allowing the context manager to log heterogeneous execution modes within the same run. Each entry maintains its own timestamp, duration, and success status, enabling the final summary to aggregate statistics across varying execution strategies.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →