What LLM Models Are Supported by AI Scientist v2: Complete Model Catalog
AI Scientist v2 supports over 40 LLM identifiers including Anthropic Claude, OpenAI GPT-4/o1/o3 series, DeepSeek, Google Gemini, Meta Llama 3, and various Ollama-hosted models, all centrally defined in the AVAILABLE_LLMS constant located in ai_scientist/llm.py.
The SakanaAI/AI-Scientist-v2 repository provides a unified interface for automated scientific research, and its flexibility starts with broad LLM compatibility. Understanding which models are supported—and how the system validates and routes requests to each provider—is essential for configuring experiments and optimizing research workflows.
The Complete AI Scientist v2 Model Catalog
All supported model identifiers live in ai_scientist/llm.py (lines 13–73) as the AVAILABLE_LLMS tuple. The system uses this catalog to validate model names before instantiating the appropriate client.
Anthropic Claude (Direct API)
The following Claude models connect directly via the Anthropic API:
claude-3-5-sonnet-20240620claude-3-5-sonnet-20241022
These map to the anthropic.Anthropic SDK wrapper in create_client().
OpenAI GPT-4 and Reasoning Models
OpenAI support spans standard GPT-4 variants and the o1/o3 reasoning series:
GPT-4o family:
gpt-4o-mini,gpt-4o-mini-2024-07-18gpt-4o,gpt-4o-2024-05-13,gpt-4o-2024-08-06gpt-4.1,gpt-4.1-2025-04-14gpt-4.1-mini,gpt-4.1-mini-2025-04-14
Reasoning models (o1/o3):
o1,o1-2024-12-17,o1-preview-2024-09-12o1-mini,o1-mini-2024-09-12o3-mini,o3-mini-2025-01-31
All use the openai.OpenAI client with specialized payload handling for reasoning parameters.
DeepSeek Coder
For DeepSeek integration:
deepseek-coder-v2-0724deepcoder-14b
Meta Llama 3
The catalog includes Meta's largest open model:
llama3.1-405b
Cloud Provider Integrations (Bedrock & Vertex AI)
AI Scientist v2 supports Anthropic Claude through enterprise cloud APIs.
Amazon Bedrock:
bedrock/anthropic.claude-3-sonnet-20240229-v1:0bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0bedrock/anthropic.claude-3-haiku-20240307-v1:0bedrock/anthropic.claude-3-opus-20240229-v1:0
Google Vertex AI:
vertex_ai/claude-3-opus@20240229vertex_ai/claude-3-5-sonnet@20240620vertex_ai/claude-3-5-sonnet@20241022vertex_ai/claude-3-sonnet@20240229vertex_ai/claude-3-haiku@20240307
Google Gemini
Native Gemini support includes:
gemini-2.0-flashgemini-2.5-flash-preview-04-17gemini-2.5-pro-preview-03-25
Ollama Local Models (GPT-OSS, Qwen, DeepSeek)
For local inference via Ollama (localhost:11434), the system recognizes:
GPT-OSS:
ollama/gpt-oss:20bollama/gpt-oss:120b
Qwen 3 and 2.5 VL:
ollama/qwen3:8b,ollama/qwen3:32b,ollama/qwen3:235bollama/qwen2.5vl:8b,ollama/qwen2.5vl:32bollama/qwen3-coder:70b,ollama/qwen3-coder:480b
DeepSeek R1:
ollama/deepseek-r1:8b,ollama/deepseek-r1:32bollama/deepseek-r1:70b,ollama/deepseek-r1:671b
How Model Selection Works Under the Hood
The architecture abstracts provider-specific details behind three key components in ai_scientist/llm.py.
Client Creation via create_client()
When you request a model, create_client(model) checks the identifier against AVAILABLE_LLMS and returns the appropriate SDK client:
- Anthropic models →
anthropic.Anthropic(or Bedrock/Vertex variants) - OpenAI-style models →
openai.OpenAI(including o1/o3 reasoning) - Ollama-hosted models →
openai.OpenAIpointing athttp://localhost:11434/v1 - DeepSeek and Gemini use specialized endpoint configurations with distinct API key handling
Request Dispatch and Retry Logic
The generic helpers get_response_from_llm() and get_batch_responses_from_llm() handle provider-specific payload shapes, system message formatting, and token limits. All calls are wrapped with backoff decorators for automatic retry on rate limits.
Token Usage Tracking
Every LLM invocation passes through track_token_usage (defined in ai_scientist/utils/token_tracker.py), which records consumption metrics regardless of provider. This enables unified cost monitoring across heterogeneous model deployments.
Working with Supported Models in Code
You can inspect the catalog and instantiate clients programmatically using the internal API.
List All Available Models
from ai_scientist.llm import AVAILABLE_LLMS
print("AI Scientist v2 supports these LLMs:")
for model in AVAILABLE_LLMS:
print(f" • {model}")
Initialize a Client and Generate Text
from ai_scientist.llm import create_client, get_response_from_llm
# Any identifier from AVAILABLE_LLMS
model_name = "gpt-4o-mini"
client, resolved_name = create_client(model_name)
system_msg = "You are a research assistant specialized in machine learning."
user_msg = "Explain the concept of inductive bias in neural networks."
output, history = get_response_from_llm(
prompt=user_msg,
client=client,
model=resolved_name,
system_message=system_msg,
temperature=0.0,
)
print(output)
Both snippets execute immediately within the repository environment because they rely solely on the internal model catalogue and abstracted client builder.
Summary
- AI Scientist v2 supports 40+ model identifiers spanning Anthropic, OpenAI, DeepSeek, Meta, Google, and Ollama, defined in
ai_scientist/llm.pyasAVAILABLE_LLMS. - Validation happens at initialization—the system rejects unsupported model strings before attempting API calls.
- Unified client architecture maps model names to the correct SDK (Anthropic, OpenAI, or Ollama-compatible) via
create_client(). - Automatic retries and token tracking wrap every request through
get_response_from_llm()andtrack_token_usage, ensuring robust handling across providers. - Enterprise options include Bedrock and Vertex AI prefixes for secure cloud deployments of Claude models.
Frequently Asked Questions
How do I add a new custom model to AI Scientist v2?
To add a custom model, append its identifier string to the AVAILABLE_LLMS tuple in ai_scientist/llm.py (lines 13–73), then update the create_client() function to handle the new provider's authentication and endpoint configuration. If the model uses an OpenAI-compatible API, you can route it through the existing Ollama client logic by pointing it at your custom base URL.
Does AI Scientist v2 support local inference without cloud APIs?
Yes. Any model identifier prefixed with ollama/ routes to a local Ollama instance at http://localhost:11434/v1. This includes GPT-OSS, Qwen, and DeepSeek R1 variants, allowing fully offline operation for sensitive research workflows.
What is the difference between gpt-4o and o1 model identifiers in the codebase?
gpt-4o models use standard chat completions with immediate token generation, while o1 and o3 identifiers trigger reasoning-specific payload handling in get_response_from_llm(). The reasoning models require different parameter structures (no temperature setting, extended timeout handling) which the backend manages transparently based on the model name prefix.
Where is the token usage tracked for cost monitoring?
Token consumption is recorded in ai_scientist/utils/token_tracker.py via the track_token_usage decorator. This wraps all LLM calls in llm.py, capturing input and output counts regardless of whether the underlying provider is Anthropic, OpenAI, or a local Ollama instance, enabling unified budget tracking across heterogeneous model deployments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →