# How context.md Tracks Generation History and Improvement Rationale in SIA

> Discover how SIA's context.md tracks agent generation history, performance metrics, and code improvement rationale. Learn about the ContextManager class.

- Repository: [Hexo Labs/sia](https://github.com/hexo-ai/sia)
- Tags: deep-dive
- Published: 2026-06-12

---

**The [`context.md`](https://github.com/hexo-ai/sia/blob/main/context.md) file in SIA serves as a comprehensive, chronological log that records every agent generation, its quantitative performance metrics, and the specific rationale for code changes, managed entirely by the `ContextManager` class in [`sia/context_manager.py`](https://github.com/hexo-ai/sia/blob/main/sia/context_manager.py).**

The `hexo-ai/sia` repository implements a self-documenting agent evolution system where [`context.md`](https://github.com/hexo-ai/sia/blob/main/context.md) functions as the central audit trail. This markdown file captures not only what changed between successive generations of an agent, but why those changes were proposed and how they impacted performance outcomes.

## The Three-Phase Logging Architecture

The `ContextManager` orchestrates the creation and maintenance of [`context.md`](https://github.com/hexo-ai/sia/blob/main/context.md) through three distinct phases, each implemented as a specific method call in [`sia/context_manager.py`](https://github.com/hexo-ai/sia/blob/main/sia/context_manager.py).

### Phase 1: Header Initialization

When a run begins, `ContextManager.initialize()` establishes the document structure by writing a header section that identifies the experimental run. According to the source code (lines 48-64), this header includes the task directory, meta-model and task-model names, the agent implementation type, the start timestamp, and the maximum number of generations allowed.

This initialization creates the foundation for the chronological narrative that follows.

### Phase 2: Per-Generation Entry Logging

After each generation completes, `ContextManager.add_generation()` appends a detailed markdown section to [`context.md`](https://github.com/hexo-ai/sia/blob/main/context.md). This method aggregates multiple data sources to create a comprehensive snapshot:

- **Agent Statistics**: File size and line count are captured via `_get_agent_stats()` (lines 27-33), providing baseline metrics for code complexity.
- **Delta Calculations**: The system calculates percentage changes in file size and line count compared to the previous generation (lines 28-36), quantifying the scope of modifications.
- **Performance Metrics**: Raw results are extracted from [`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json), [`detailed_results.json`](https://github.com/hexo-ai/sia/blob/main/detailed_results.json), or log files through `_extract_metrics()` (lines 38-84), capturing accuracy, execution time, and task-specific outcomes.
- **Improvement Rationale**: The system parses [`improvement.md`](https://github.com/hexo-ai/sia/blob/main/improvement.md) via `_extract_insights()` (lines 25-52) to capture the LLM-generated reasoning behind proposed changes.
- **Evolution Summary**: A concise narrative comparing previous and current agent code is generated by `_generate_llm_summary()` (lines 66-78, 114-140), synthesizing the improvement plan and metric changes into human-readable insights.

All components are assembled by `_format_generation_entry()` (lines 54-59) into a standardized markdown block separated by `---` delimiters, ensuring clear visual separation between generations.

### Phase 3: Final Summary Statistics

When the run concludes, `ContextManager.finalize()` appends a **Summary Statistics** section (lines 67-108) that aggregates the entire evolution history. This final block reports total generations executed, success counts, identification of the best-performing generation, overall accuracy evolution trends, and code-size growth patterns across the entire run.

## Key Technical Components

### Tracking Metric Deltas

The logging system automatically calculates and displays **Changes vs Previous Generation** tables (lines 28-55). These tables highlight metric deltas, making it immediately apparent whether a code modification improved or degraded performance, and by what magnitude.

### Linking Code to Rationale

By extracting content from [`improvement.md`](https://github.com/hexo-ai/sia/blob/main/improvement.md) through `_extract_insights()`, [`context.md`](https://github.com/hexo-ai/sia/blob/main/context.md) maintains a tight coupling between quantitative results and qualitative reasoning. This allows developers to trace specific performance changes back to the specific improvement hypotheses that generated them.

## Implementation Example

The following implementation demonstrates the complete lifecycle of [`context.md`](https://github.com/hexo-ai/sia/blob/main/context.md) generation:

```python
from sia.context_manager import ContextManager
from sia.config import Config

# Initialize the manager for a run directory

run_dir = "runs/run_1"
run_cfg = {
    "task_dir": "tasks/example",
    "meta_model": "claude-2",
    "task_model": "gpt-4",
    "agent_impl": "claude",
    "max_gen": 5,
}
ctx = ContextManager(run_dir, run_cfg, Config())
ctx.initialize()                     # Writes the header to context.md

# After each generation, add an entry

gen_data = {
    "success": True,
    "timestamp": "2026-06-12 12:34:56",
    "duration": 42.7,
    "agent_path": f"{run_dir}/gen_1/target_agent.py",
    "gen_dir": f"{run_dir}/gen_1",
    "improvement_path": f"{run_dir}/gen_1/improvement.md",
    "execution_type": "Single",
}
ctx.add_generation(gen_num=1, gen_data=gen_data)   # Appends Generation-1 block

# When the run finishes

ctx.finalize()                       # Appends Summary Statistics block

```

This workflow mirrors the patterns validated in [`tests/test_context_manager.py`](https://github.com/hexo-ai/sia/blob/main/tests/test_context_manager.py) (lines 38-84), which verifies the full lifecycle from initialization through finalization.

## Summary

- **[`context.md`](https://github.com/hexo-ai/sia/blob/main/context.md)** serves as the central chronological record of agent evolution in the SIA framework, combining quantitative metrics with qualitative improvement rationale.
- The **`ContextManager`** class in [`sia/context_manager.py`](https://github.com/hexo-ai/sia/blob/main/sia/context_manager.py) orchestrates three distinct phases: header initialization, per-generation entry logging, and final summary statistics.
- Each generation entry captures **agent statistics**, **performance deltas**, **LLM-generated improvement insights**, and **evolution summaries** to create a self-documenting audit trail.
- The system automatically extracts metrics from JSON result files and improvement rationale from [`improvement.md`](https://github.com/hexo-ai/sia/blob/main/improvement.md), linking code changes to their intended outcomes.
- **Final summary statistics** provide aggregate views of success rates, best-performing generations, and code growth patterns across the entire experimental run.

## Frequently Asked Questions

### What file format does context.md use?

The [`context.md`](https://github.com/hexo-ai/sia/blob/main/context.md) file uses standard **markdown format** with ATX-style headers, bullet lists, and tables. This choice ensures the generation history remains human-readable while being parseable by standard markdown processors. Each generation entry is separated by `---` horizontal rules to demarcate distinct iterations clearly.

### How does SIA extract improvement rationale for the log?

The `ContextManager` extracts improvement rationale through the `_extract_insights()` method (lines 25-52), which parses the optional [`improvement.md`](https://github.com/hexo-ai/sia/blob/main/improvement.md) file generated during each iteration. This file contains the LLM's reasoning for proposed code changes, which the context manager then embeds directly into the corresponding generation entry, creating a direct link between the justification and the resulting metrics.

### Where is the ContextManager implementation tested?

The core functionality is validated in **[`tests/test_context_manager.py`](https://github.com/hexo-ai/sia/blob/main/tests/test_context_manager.py)** (lines 38-84), which exercises the complete lifecycle including initialization, generation entry addition, and finalization. Additionally, [`tests/golden/feedback_context_success_single.txt`](https://github.com/hexo-ai/sia/blob/main/tests/golden/feedback_context_success_single.txt) provides a reference example of a fully rendered [`context.md`](https://github.com/hexo-ai/sia/blob/main/context.md) entry showing the expected format for metrics, insights, and LLM summaries.

### Can context.md track multiple generations with different execution types?

Yes, the `add_generation()` method accepts a `gen_data` dictionary that includes an `execution_type` field (e.g., "Single"), allowing the context manager to log heterogeneous execution modes within the same run. Each entry maintains its own timestamp, duration, and success status, enabling the final summary to aggregate statistics across varying execution strategies.