How to Specify Different LLMs for Different Tasks in AI Scientist v2

AI Scientist v2 decouples model selection from task execution, allowing you to assign specific LLMs to distinct stages—code generation, feedback, write-ups, citations, and plotting—via YAML configuration files and command-line flags that route through the central create_client dispatcher in ai_scientist/llm.py.

AI Scientist v2 by SakanaAI introduces a modular architecture where each research stage can leverage the most suitable language model. Whether you need Claude for complex code generation, GPT-4o for rapid feedback, or a local Ollama instance for cost-effective plotting, the system routes requests through a unified dispatch layer. This article explains how to configure these per-task assignments using the repository's configuration system and CLI arguments.

Configuration-Driven Model Selection via YAML

The default experiment configuration lives in bfts_config.yaml, which separates model assignments by functional role. During startup, launch_scientist_bfts.py loads this configuration via load_cfg in ai_scientist/treesearch/utils/config.py, exposing the AgentConfig fields through cfg.agent to the rest of the system.

Code Generation and Feedback Models

The agent section in bfts_config.yaml defines distinct models for writing and evaluating code:

  • agent.code.model — The LLM that generates Python experiment code (default: anthropic.claude-3-5-sonnet-20241022-v2:0)
  • agent.feedback.model — The LLM that evaluates execution results and traces errors (default: gpt-4o-2024-11-20)
  • agent.vlm_feedback.model — The vision-language model for image captioning and figure analysis

Each field accepts any model string recognized by the AVAILABLE_LLMS list in ai_scientist/llm.py (lines 13–33). For example:

agent:
  code:
    model: anthropic.claude-3-5-sonnet-20241022-v2:0
    temp: 1.0
  feedback:
    model: gpt-4o-2024-11-20
    temp: 0.5
  vlm_feedback:
    model: gemini-2.5-flash
    temp: 0.7

The agent constructs separate clients for each role by passing these strings to create_client at runtime.

Command-Line Overrides for Research Stages

For stages occurring outside the core agent loop—specifically write-up generation, citation gathering, and plot aggregation—AI Scientist v2 exposes dedicated CLI flags in launch_scientist_bfts.py (lines 86–102). These override the configuration when invoking helper functions like perform_writeup, gather_citations, and aggregate_plots.

Write-Up Generation Flags

The write-up stage supports a two-tier model strategy:

  • --model_writeup — The primary LLM for generating the research paper (default: o1-preview-2024-09-12)
  • --model_writeup_small — A lightweight model for subsidiary tasks (default: gpt-4o-2024-05-13)

These values flow into perform_writeup and perform_icbinb_writeup, which instantiate clients via create_client.

Citation and Plot Aggregation Models

Additional flags control auxiliary intellectual tasks:

  • --model_citation — LLM for literature search and citation formatting (default: gpt-4o-2024-11-20)
  • --model_agg_plots — LLM for summarizing and aggregating visual results (default: o3-mini-2025-01-31)

Example execution mixing multiple providers:

python -m launch_scientist_bfts \
  --model_writeup claude-3-5-sonnet-20241022 \
  --model_citation gpt-4o-mini \
  --model_agg_plots ollama/gpt-oss:20b

The Model Dispatch Architecture

All model strings route through create_client in ai_scientist/llm.py (lines 80–104), which acts as a factory returning the appropriate API wrapper and canonical model name. Subsequent calls like get_response_from_llm, make_llm_call, and get_batch_responses_from_llm use that client to issue requests.

Supported Model Prefixes and Backends

The dispatcher recognizes provider prefixes and keywords to instantiate the correct client:

  • claude-… → Anthropic client
  • bedrock/… → Amazon Bedrock (Claude)
  • vertex_ai/… → Google Vertex AI (Claude)
  • ollama/… → Ollama local endpoint (OpenAI-compatible)
  • Contains gpt, o1, or o3 → OpenAI client
  • deepseek-coder-v2-0724 → DeepSeek OpenAI-compatible endpoint
  • deepcoder-14b → HuggingFace inference
  • llama3.1-405b → OpenRouter
  • gemini → Google Gemini wrapper

Because the model parameter is a plain string, you can freely mix providers—using Claude for reasoning, GPT-4o for structured output, and local models for cost-sensitive batch tasks—without modifying core logic in perform_writeup, perform_icbinb_writeup, or perform_llm_review.py.

Practical Configuration Examples

Customizing bfts_config.yaml for Local Development

To reduce API costs during development, configure local or smaller models for iterative tasks while reserving powerful models for final write-ups:

agent:
  code:
    model: ollama/deepcoder:14b
    temp: 0.8
    max_tokens: 12000
  feedback:
    model: gpt-4o-mini
    temp: 0.5
  vlm_feedback:
    model: gemini-2.5-flash-preview

Save the file and execute with CLI overrides for the write-up stage:

python -m launch_scientist_bfts \
  --model_writeup o1-preview-2024-09-12 \
  --model_citation gpt-4o-2024-08-06

Dynamic Model Selection in Python

For programmatic control, import the dispatch utilities directly from ai_scientist.llm:

from ai_scientist.llm import create_client, get_response_from_llm

def ask_llm(prompt: str, model_name: str) -> str:
    client, canonical = create_client(model_name)
    response, _ = get_response_from_llm(
        prompt=prompt,
        client=client,
        model=canonical,
        system_message="You are a helpful AI researcher.",
        temperature=0.7,
    )
    return response

# Use Claude for reasoning, then GPT-4o for summarization

reasoning = ask_llm("Explain why this loss function fails.", "claude-3-5-sonnet-20241022")
summary = ask_llm(f"Summarize in 3 sentences: {reasoning}", "gpt-4o-2024-11-20")

Summary

  • AI Scientist v2 separates LLM selection from execution logic, enabling per-task model assignment through bfts_config.yaml and CLI flags.
  • Code generation, feedback, and vision tasks are configured in bfts_config.yaml under agent.code.model, agent.feedback.model, and agent.vlm_feedback.model.
  • Write-ups, citations, and plot aggregation accept model overrides via --model_writeup, --model_writeup_small, --model_citation, and --model_agg_plots in launch_scientist_bfts.py.
  • The create_client function in ai_scientist/llm.py (lines 80–104) dispatches model strings to the appropriate backend based on prefix matching.
  • Because model identifiers are strings, you can mix proprietary APIs, local endpoints, and specialized models within a single experiment without code changes.

Frequently Asked Questions

Can I use local models like Ollama for specific tasks while using cloud APIs for others?

Yes. Prefix the model name with ollama/ in bfts_config.yaml or CLI flags (e.g., ollama/gpt-oss:20b). The create_client dispatcher routes these to your local Ollama endpoint while routing other tasks to Anthropic, OpenAI, or Google APIs based on their respective prefixes.

What happens if I don't specify a model for a specific task?

Each task has a default defined in the codebase. The YAML configuration defaults to Claude for code generation and GPT-4o for feedback, while launch_scientist_bfts.py provides defaults like o1-preview-2024-09-12 for write-ups. If a required model string is missing, the system raises a configuration error during client initialization.

How does the system handle model name canonicalization?

create_client returns a tuple of (client, canonical_model_name). The canonical name normalizes the identifier to the format expected by the specific backend API. For example, Bedrock model IDs are transformed to the format Anthropic's API expects, ensuring compatibility when the client actually calls get_response_from_llm or get_batch_responses_from_llm.

Can I mix proprietary and open-source models in the same experiment?

Absolutely. The architecture encourages mixing models by capability and cost. You might use claude-3-5-sonnet for code generation in bfts_config.yaml, override with --model_writeup o1-preview for the final paper, and use ollama/llama3.1 for --model_agg_plots to summarize figures cheaply. The dispatcher instantiates the appropriate client for each call without cross-contamination.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →