How Ablation Studies Are Performed in Stage 4 of AI Scientist v2
Ablation studies in Stage 4 of AI Scientist v2 are orchestrated by the ParallelAgent class, which systematically generates, executes, and tracks experiments that remove or modify components from the best Stage 3 implementation to isolate their impact.
Stage 4 serves as the experimental validation phase of the SakanaAI/AI-Scientist-v2 pipeline. Unlike earlier stages that focus on ideation and implementation, this stage rigorously tests the contribution of individual architectural choices using an LLM-guided workflow that prevents duplicate experiments and maintains scientific rigor.
Initialization of the ParallelAgent
When the pipeline transitions to Stage 4, the AgentManager._create_agent_for_stage method in ai_scientist/treesearch/agent_manager.py constructs a dedicated ParallelAgent instance. This agent receives the best implementation from Stage 3 via the best_stage3_node parameter, establishing the baseline code that will undergo systematic ablation.
The initialization logic specifically checks for main_stage == 4 to trigger this specialized agent configuration:
# Inside AgentManager._create_agent_for_stage
if main_stage == 4:
# Retrieve the best node from Stage 3
best_stage3_node = self._get_best_implementation(last_substage.name)
agent = ParallelAgent(
task_desc=task_desc,
cfg=stage_cfg,
journal=self.journals[stage.name],
stage_name=stage.name, # e.g. "4_ablation"
best_stage3_node=best_stage3_node, # baseline code to ablate
)
Baseline Selection and Node Handling
The ParallelAgent._select_parallel_nodes method detects Stage 4 execution through a simple prefix check on self.stage_name. When the stage name starts with "4_", the agent automatically adds self.best_stage3_node to the processing queue according to lines 2006-2012 in ai_scientist/treesearch/parallel_agent.py.
This design ensures that every parallel worker operates on the identical baseline implementation, preventing the redundant ablation attempts that would occur if workers selected different starting points.
Generating Ablation Ideas with LLM Prompting
The core of Stage 4's scientific methodology resides in _generate_ablation_idea, a private method that constructs targeted prompts for the LLM. This method builds a structured context containing:
- The current implementation code (
self.best_stage3_node.code) - A history of completed ablations (
self._ablation_state["completed_ablations"]) - Specific instructions to propose one new, distinct ablation (such as disabling a specific component or swapping a dataset)
The prompt engineering explicitly requires the LLM to identify a single component to ablate and ensure it differs from previous attempts:
def _generate_ablation_idea(self) -> Optional[AblationIdea]:
completed = list(self._ablation_state["completed_ablations"])
prompt = {
"Introduction": "You are an AI researcher conducting ablation studies...",
"Base code you are working on": wrap_code(self.best_stage3_node.code),
"Previous Ablations": {"Has been tried": completed or "Nothing has been tried yet."},
"Instructions": {"Requirements": [
"1. Identify ONE specific component/feature to ablate",
"2. Ensure the ablation is different from previous attempts",
"3. ..."]},
"Response format": "ABLATION NAME: ...\\nABLATION DESCRIPTION: ..."
}
response = query(system_message=prompt,
model=self.cfg.agent.code.model,
temperature=self.cfg.agent.code.temp)
name, desc = _parse_keyword_prefix_response(
response, "ABLATION NAME:", "ABLATION DESCRIPTION:")
return AblationIdea(name=name, description=desc) if name and desc else None
If parsing fails after configured retries, the system implements a fallback mechanism that returns a default ablation idea (adding a layer) to maintain workflow continuity.
Tracking Ablation State and Preventing Duplicates
Stage 4 maintains rigorous experiment tracking through the _ablation_state dictionary, which uses a "completed_ablations" set to ensure scientific validity. The workflow implements immediate reservation of ablation ideas:
When a new ablation is generated, it is immediately added to self._ablation_state["completed_ablations"] (lines 2111-2113) before execution begins. This prevents other parallel workers from proposing identical experiments even while the first worker is still running the code.
After execution completes, _update_ablation_state validates the result:
def _update_ablation_state(self, result_node: Node):
if not self.stage_name or not self.stage_name.startswith("4_"):
return
if result_node.ablation_name and not result_node.is_buggy:
self._ablation_state["completed_ablations"].add(result_node.ablation_name)
logger.info(f"Ablation {result_node.ablation_name} completed successfully")
Successful runs (non-buggy nodes) are recorded permanently, while failed experiments are logged but excluded from the completion set, allowing the system to retry or propose alternatives.
Parallel Execution Architecture
The ParallelAgent spawns a process pool (self.executor) that maps workers across available GPUs or CPU cores. Each worker receives:
- The baseline node data from Stage 3
- The generated ablation idea (name and description)
- Auxiliary context such as plot code from prior stages
Workers execute the modified code, capture performance metrics, and return result nodes to the main process, which inserts them into the stage journal for analysis.
Completion Criteria and Configuration
Stage 4 terminates based on criteria defined in bfts_config.yaml. The workflow respects the stage4_max_iters: 18 limit by default, or stops when no further improvements are observed against the evaluation metric defined in _define_global_metrics.
Upon completion, the system aggregates results through perform_experiments_bfts_with_agentmanager.py, incorporating ablation summaries into the final research report.
Summary
- ParallelAgent orchestration: Stage 4 uses a specialized agent that receives the best Stage 3 implementation via
best_stage3_nodeinagent_manager.py. - Systematic idea generation: The
_generate_ablation_ideamethod inparallel_agent.pyuses structured LLM prompting to propose unique component removals. - Duplicate prevention: The
_ablation_statedictionary trackscompleted_ablationsimmediately upon idea generation and confirms them after successful execution. - Parallel execution: Process pools distribute ablation experiments across hardware resources while maintaining shared state.
- Configured limits: Execution respects
stage4_max_itersfrombfts_config.yamland global metric thresholds.
Frequently Asked Questions
What is the purpose of Stage 4 in AI Scientist v2?
Stage 4 conducts ablation studies to isolate and measure the contribution of specific architectural components or features within the best implementation discovered in Stage 3. This stage transforms promising implementations into rigorously validated scientific findings by systematically removing elements and measuring performance deltas.
How does the system prevent running the same ablation twice?
The system implements a reservation pattern through _ablation_state["completed_ablations"]. When an ablation idea is generated, it is immediately added to this set before execution begins. The _update_ablation_state method only permanently records successful runs, while failed attempts remain available for retry or replacement with alternative ablations.
What happens if the LLM fails to generate a valid ablation idea?
The _generate_ablation_idea method includes a fallback mechanism (lines 1910-1919 in parallel_agent.py) that activates after multiple parsing retries fail. When the LLM response cannot be parsed for "ABLATION NAME" and "ABLATION DESCRIPTION", the system defaults to a generic fallback idea—typically adding a layer—to ensure the workflow continues rather than stalling.
Where is the baseline code for ablation studies stored?
The baseline code originates from Stage 3's best implementation, passed through the best_stage3_node parameter during ParallelAgent initialization in agent_manager.py. This node contains the complete code that workers modify according to each ablation specification, ensuring all experiments share a consistent starting point.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →