How to Specify Different LLMs for Different Tasks in AI Scientist v2
AI Scientist v2 decouples model selection from task execution, allowing you to assign specific LLMs to distinct stages—code generation, feedback, write-ups, citations, and plotting—via YAML configuration files and command-line flags that route through the central create_client dispatcher in ai_scientist/llm.py.
AI Scientist v2 by SakanaAI introduces a modular architecture where each research stage can leverage the most suitable language model. Whether you need Claude for complex code generation, GPT-4o for rapid feedback, or a local Ollama instance for cost-effective plotting, the system routes requests through a unified dispatch layer. This article explains how to configure these per-task assignments using the repository's configuration system and CLI arguments.
Configuration-Driven Model Selection via YAML
The default experiment configuration lives in bfts_config.yaml, which separates model assignments by functional role. During startup, launch_scientist_bfts.py loads this configuration via load_cfg in ai_scientist/treesearch/utils/config.py, exposing the AgentConfig fields through cfg.agent to the rest of the system.
Code Generation and Feedback Models
The agent section in bfts_config.yaml defines distinct models for writing and evaluating code:
agent.code.model— The LLM that generates Python experiment code (default:anthropic.claude-3-5-sonnet-20241022-v2:0)agent.feedback.model— The LLM that evaluates execution results and traces errors (default:gpt-4o-2024-11-20)agent.vlm_feedback.model— The vision-language model for image captioning and figure analysis
Each field accepts any model string recognized by the AVAILABLE_LLMS list in ai_scientist/llm.py (lines 13–33). For example:
agent:
code:
model: anthropic.claude-3-5-sonnet-20241022-v2:0
temp: 1.0
feedback:
model: gpt-4o-2024-11-20
temp: 0.5
vlm_feedback:
model: gemini-2.5-flash
temp: 0.7
The agent constructs separate clients for each role by passing these strings to create_client at runtime.
Command-Line Overrides for Research Stages
For stages occurring outside the core agent loop—specifically write-up generation, citation gathering, and plot aggregation—AI Scientist v2 exposes dedicated CLI flags in launch_scientist_bfts.py (lines 86–102). These override the configuration when invoking helper functions like perform_writeup, gather_citations, and aggregate_plots.
Write-Up Generation Flags
The write-up stage supports a two-tier model strategy:
--model_writeup— The primary LLM for generating the research paper (default:o1-preview-2024-09-12)--model_writeup_small— A lightweight model for subsidiary tasks (default:gpt-4o-2024-05-13)
These values flow into perform_writeup and perform_icbinb_writeup, which instantiate clients via create_client.
Citation and Plot Aggregation Models
Additional flags control auxiliary intellectual tasks:
--model_citation— LLM for literature search and citation formatting (default:gpt-4o-2024-11-20)--model_agg_plots— LLM for summarizing and aggregating visual results (default:o3-mini-2025-01-31)
Example execution mixing multiple providers:
python -m launch_scientist_bfts \
--model_writeup claude-3-5-sonnet-20241022 \
--model_citation gpt-4o-mini \
--model_agg_plots ollama/gpt-oss:20b
The Model Dispatch Architecture
All model strings route through create_client in ai_scientist/llm.py (lines 80–104), which acts as a factory returning the appropriate API wrapper and canonical model name. Subsequent calls like get_response_from_llm, make_llm_call, and get_batch_responses_from_llm use that client to issue requests.
Supported Model Prefixes and Backends
The dispatcher recognizes provider prefixes and keywords to instantiate the correct client:
claude-…→ Anthropic clientbedrock/…→ Amazon Bedrock (Claude)vertex_ai/…→ Google Vertex AI (Claude)ollama/…→ Ollama local endpoint (OpenAI-compatible)- Contains
gpt,o1, oro3→ OpenAI client deepseek-coder-v2-0724→ DeepSeek OpenAI-compatible endpointdeepcoder-14b→ HuggingFace inferencellama3.1-405b→ OpenRoutergemini→ Google Gemini wrapper
Because the model parameter is a plain string, you can freely mix providers—using Claude for reasoning, GPT-4o for structured output, and local models for cost-sensitive batch tasks—without modifying core logic in perform_writeup, perform_icbinb_writeup, or perform_llm_review.py.
Practical Configuration Examples
Customizing bfts_config.yaml for Local Development
To reduce API costs during development, configure local or smaller models for iterative tasks while reserving powerful models for final write-ups:
agent:
code:
model: ollama/deepcoder:14b
temp: 0.8
max_tokens: 12000
feedback:
model: gpt-4o-mini
temp: 0.5
vlm_feedback:
model: gemini-2.5-flash-preview
Save the file and execute with CLI overrides for the write-up stage:
python -m launch_scientist_bfts \
--model_writeup o1-preview-2024-09-12 \
--model_citation gpt-4o-2024-08-06
Dynamic Model Selection in Python
For programmatic control, import the dispatch utilities directly from ai_scientist.llm:
from ai_scientist.llm import create_client, get_response_from_llm
def ask_llm(prompt: str, model_name: str) -> str:
client, canonical = create_client(model_name)
response, _ = get_response_from_llm(
prompt=prompt,
client=client,
model=canonical,
system_message="You are a helpful AI researcher.",
temperature=0.7,
)
return response
# Use Claude for reasoning, then GPT-4o for summarization
reasoning = ask_llm("Explain why this loss function fails.", "claude-3-5-sonnet-20241022")
summary = ask_llm(f"Summarize in 3 sentences: {reasoning}", "gpt-4o-2024-11-20")
Summary
- AI Scientist v2 separates LLM selection from execution logic, enabling per-task model assignment through
bfts_config.yamland CLI flags. - Code generation, feedback, and vision tasks are configured in
bfts_config.yamlunderagent.code.model,agent.feedback.model, andagent.vlm_feedback.model. - Write-ups, citations, and plot aggregation accept model overrides via
--model_writeup,--model_writeup_small,--model_citation, and--model_agg_plotsinlaunch_scientist_bfts.py. - The
create_clientfunction inai_scientist/llm.py(lines 80–104) dispatches model strings to the appropriate backend based on prefix matching. - Because model identifiers are strings, you can mix proprietary APIs, local endpoints, and specialized models within a single experiment without code changes.
Frequently Asked Questions
Can I use local models like Ollama for specific tasks while using cloud APIs for others?
Yes. Prefix the model name with ollama/ in bfts_config.yaml or CLI flags (e.g., ollama/gpt-oss:20b). The create_client dispatcher routes these to your local Ollama endpoint while routing other tasks to Anthropic, OpenAI, or Google APIs based on their respective prefixes.
What happens if I don't specify a model for a specific task?
Each task has a default defined in the codebase. The YAML configuration defaults to Claude for code generation and GPT-4o for feedback, while launch_scientist_bfts.py provides defaults like o1-preview-2024-09-12 for write-ups. If a required model string is missing, the system raises a configuration error during client initialization.
How does the system handle model name canonicalization?
create_client returns a tuple of (client, canonical_model_name). The canonical name normalizes the identifier to the format expected by the specific backend API. For example, Bedrock model IDs are transformed to the format Anthropic's API expects, ensuring compatibility when the client actually calls get_response_from_llm or get_batch_responses_from_llm.
Can I mix proprietary and open-source models in the same experiment?
Absolutely. The architecture encourages mixing models by capability and cost. You might use claude-3-5-sonnet for code generation in bfts_config.yaml, override with --model_writeup o1-preview for the final paper, and use ollama/llama3.1 for --model_agg_plots to summarize figures cheaply. The dispatcher instantiates the appropriate client for each call without cross-contamination.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →