AI-Scientist-v2 Hyperparameter Tuning in Stage 2: Closed-Loop Baseline Optimization

The AI-Scientist-v2 system handles hyperparameter tuning in Stage 2 through a closed feedback loop that maintains a global state of tested configurations, prompts an LLM to propose novel hyperparameters, executes experiments in the main process, and updates the state to prevent duplicate trials until the configured trial limit is reached.

Stage 2 of the AI-Scientist-v2 pipeline focuses exclusively on baseline hyperparameter tuning, where the system autonomously explores configuration space to optimize foundational model performance. This stage operates through the ParallelAgent class in ai_scientist/treesearch/parallel_agent.py, which orchestrates a systematic search while avoiding redundant experiments through persistent state tracking.

Initializing the Global Tuning State

When ParallelAgent is instantiated, it creates a dedicated data structure to track every hyperparameter that has been attempted during Stage 2. This initialization occurs in parallel_agent.py at lines 1190–1192:

self._hyperparam_tuning_state = {"tried_hyperparams": set()}

This dictionary serves as the source of truth for the entire tuning lifecycle. The "tried_hyperparams" set ensures that the system maintains a unique record of all configurations that have already been executed, enabling the deduplication logic that prevents the LLM from proposing identical settings in subsequent iterations.

Generating the Next Hyperparameter Idea

The core intelligence of Stage 2 resides in _generate_hyperparam_tuning_idea (lines 1798–1846 of parallel_agent.py). This method constructs a detailed prompt that instructs the LLM to act as an AI researcher conducting baseline hyperparameter tuning.

The prompt explicitly commands: “You are an AI researcher conducting baseline hyper‑parameter tuning. Propose ONE new hyper‑parameter that has not been tried yet.”

The system injects three critical context elements into the prompt:

  • The current baseline code
  • The complete list of previously tried hyperparameters from self._hyperparam_tuning_state["tried_hyperparams"]
  • Strict formatting instructions requiring HYPERPARAM NAME: and DESCRIPTION: fields

If parsing the LLM response fails, the system implements a retry mechanism with up to five attempts before falling back to a default "increase learning rate" idea. This robust error handling ensures that the tuning loop continues even when the LLM produces malformed output.

Creating and Executing Tuning Nodes

Once a novel hyperparameter idea is generated, the system calls _generate_hyperparam_tuning_node (lines 557–606) to construct an executable experiment. This method creates a Node object that encapsulates:

return Node(
    plan="Hyperparam tuning name: " + hyperparam_idea.name + ".\n" + plan,
    code=code,
    parent=parent_node,
    hyperparam_name=hyperparam_idea.name,
)

The hyperparam_name attribute (defined in ai_scientist/treesearch/journal.py) links the node directly to the specific tuning idea.

Stage 2 executes these nodes in the main process rather than distributed workers. The generated node is handed to the worker pool, executed, and its results are parsed to determine success or failure.

Updating State and Preventing Duplicates

After execution completes, _update_hyperparam_tuning_state (lines 2192–2208) processes the results:

if not result_node.is_buggy:
    self._hyperparam_tuning_state["tried_hyperparams"].add(hyperparam_name)

Only successful experiments (where result_node.is_buggy is False) are added to the tried set. Failed runs are logged as warnings but not recorded as completed attempts, allowing the system to retry similar configurations if the failure was transient.

When the loop requests the next idea, the updated state feeds back into _generate_hyperparam_tuning_idea, which includes the current "tried_hyperparams" set in the "Previous Hyperparam Tuning Attempts" field of the prompt. This creates a closed feedback loop: state → LLM prompt → LLM idea → node generation → execution → state update.

Configuration and Stopping Criteria

The tuning process continues until exhaustion of the configured trial budget. The limit is controlled by search_cfg.num_trials_stage2, defined in bfts_config.yaml. This configuration parameter determines how many distinct hyperparameter configurations the system will attempt before transitioning to the next stage of the AI-Scientist pipeline.

Summary

  • Global state tracking: The system initializes self._hyperparam_tuning_state with a "tried_hyperparams" set to record all attempted configurations.
  • LLM-driven exploration: _generate_hyperparam_tuning_idea prompts the LLM to propose novel hyperparameters while injecting the history of previous attempts to prevent duplicates.
  • Robust node generation: _generate_hyperparam_tuning_node packages the idea into an executable Node with the hyperparam_name attribute for traceability.
  • State mutation: _update_hyperparam_tuning_state adds successful experiments to the tried set, maintaining the deduplication invariant.
  • Configurable limits: The loop terminates after exhausting num_trials_stage2 as specified in the YAML configuration.

Frequently Asked Questions

How does the system prevent testing the same hyperparameter twice?

The system maintains self._hyperparam_tuning_state["tried_hyperparams"], a set that records every successfully executed hyperparameter name. This set is injected into the LLM prompt via the "Previous Hyperparam Tuning Attempts" field, explicitly instructing the model to propose only novel configurations. When _update_hyperparam_tuning_state runs after each experiment, it adds the hyperparam_name to this set only if the node executed without bugs.

What happens if the LLM fails to generate a valid hyperparameter idea?

The _generate_hyperparam_tuning_idea method implements a retry mechanism that attempts parsing up to five times. If the LLM response fails to match the required HYPERPARAM NAME: and DESCRIPTION: format after all retries, the system falls back to a default "increase learning rate" idea. This ensures the tuning loop never stalls due to malformed LLM output.

Where is the hyperparameter tuning state stored?

The tuning state resides in memory as the self._hyperparam_tuning_state dictionary within the ParallelAgent instance, initialized at lines 1190–1192 of ai_scientist/treesearch/parallel_agent.py. The state persists only for the duration of the Stage 2 execution and is not serialized to disk between stages.

How many trials does Stage 2 run by default?

The number of trials is controlled by the search_cfg.num_trials_stage2 parameter in bfts_config.yaml. The agent continues requesting new hyperparameter ideas and executing nodes until this configured limit is reached, at which point Stage 2 concludes and the system proceeds to the next phase of the AI-Scientist workflow.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →