How to Integrate Custom LLMs with AI Scientist v2: The Complete Developer Guide
Integrating custom LLMs with AI Scientist v2 requires registering your model identifier in the AVAILABLE_LLMS list within ai_scientist/llm.py, extending the create_client() function to instantiate your specific SDK client, and optionally adjusting request payload handling in get_response_from_llm().
AI Scientist v2 from SakanaAI provides a modular architecture for automating scientific research workflows using language models. To integrate custom LLMs with AI Scientist v2—whether hosted locally, via a private API, or through specialized providers—you must extend the thin abstraction layer defined in the core LLM module. This guide walks through the exact source code modifications needed to connect proprietary or self-hosted models to the ideation, review, and writeup pipelines.
Understanding the LLM Abstraction Architecture
The integration point lives in ai_scientist/llm.py, which implements a three-stage workflow:
-
Model Declaration – The
AVAILABLE_LLMSlist (defined around lines 13–74) contains every supported model identifier, such as"gpt-4o-mini"or"bedrock/anthropic.claude-3-sonnet-20240229-v1:0". These strings are validated against CLI arguments across all entry-point scripts. -
Client Instantiation – The
create_client(model)function maps identifiers to SDK instances. It branches by prefix (e.g.,"ollama/","gpt","claude") to return the appropriate client object and normalized model name. -
Request Execution – The
get_response_from_llm()function (spanning roughly lines 80–440) constructs payloads, manages message history, and parses responses. It contains provider-specific logic for OpenAI, Anthropic, Gemini, and other families.
Step-by-Step Guide to Custom LLM Integration
Step 1: Register the Model Identifier
Append your custom model string to AVAILABLE_LLMS in ai_scientist/llm.py. Use a descriptive prefix to avoid collisions:
# In ai_scientist/llm.py
AVAILABLE_LLMS.append("myco/custom-llm")
This registration enables the --model CLI argument in scripts like perform_ideation_temp_free.py and perform_writeup.py to recognize your model.
Step 2: Create a Client Branch
Add a conditional branch in create_client() to instantiate your HTTP client. For OpenAI-compatible endpoints, reuse the existing openai library:
elif model.startswith("myco/"):
endpoint = os.getenv("MYCO_ENDPOINT", "https://api.myco.com/v1")
api_key = os.getenv("MYCO_API_KEY")
print(f"Using MyCo endpoint {endpoint} for model {model}.")
return openai.OpenAI(api_key=api_key, base_url=endpoint), model
Return a tuple of (client_instance, model_string) to maintain compatibility with downstream callers.
Step 3: Adjust Request Handling (Optional)
If your LLM service uses a custom request/response schema, add a dedicated processing branch inside get_response_from_llm(). For standard OpenAI chat completions, the existing logic handles your client automatically without modification.
Step 4: Configure Environment Variables
Store credentials and endpoints outside source control. Define variables like MYCO_ENDPOINT and MYCO_API_KEY, then document them in your .env.example or README. This aligns with the repository's security patterns used for OpenAI and Anthropic keys.
Step 5: Test the Integration
Validate connectivity by running an ideation script:
python -m ai_scientist.perform_ideation_temp_free \
--model myco/custom-llm \
--max-num-generations 1
The @track_token_usage decorator in ai_scientist/utils/token_tracker.py automatically records usage statistics without requiring code changes.
Practical Example: Local OpenAI-Compatible Server
To connect a local Llama 3 instance running on http://localhost:8000/v1:
# Add to AVAILABLE_LLMS in ai_scientist/llm.py
AVAILABLE_LLMS.append("local/llama-3-8b")
# Extend create_client with a local branch
elif model.startswith("local/"):
endpoint = os.getenv("LOCAL_LLM_ENDPOINT", "http://localhost:8000/v1")
print(f"Using local LLM at {endpoint}")
return openai.OpenAI(api_key="not-needed", base_url=endpoint), model
Invoke the writeup pipeline:
python -m ai_scientist.perform_writeup \
--folder my_project \
--model local/llama-3-8b \
--big-model gpt-4o-2024-05-13
Code Examples for Production Integration
Minimal Reproducible Integration Snippet
import os
from ai_scientist.llm import create_client, get_response_from_llm, AVAILABLE_LLMS
# 1️⃣ Register model (do this before CLI parsing occurs)
AVAILABLE_LLMS.append("myco/custom-llm")
# 2️⃣ (Inject the elif branch into ai_scientist/llm.py as shown above)
# 3️⃣ Use the model programmatically
client, client_model = create_client("myco/custom-llm")
system_msg = "You are a helpful research assistant."
prompt = "Summarize the key contributions of the paper 'Attention Is All You Need'."
answer, history = get_response_from_llm(
prompt=prompt,
client=client,
model=client_model,
system_message=system_msg,
temperature=0.0,
)
print("LLM answer:", answer)
Running Built-in Scripts with Custom Models
export MYCO_ENDPOINT="https://api.myco.com/v1"
export MYCO_API_KEY="sk-****************"
python -m ai_scientist.perform_ideation_temp_free \
--model myco/custom-llm \
--max-num-generations 2 \
--workshop-file ideas/workshop.md
Summary
- AI Scientist v2 centralizes LLM integration in
ai_scientist/llm.pythrough theAVAILABLE_LLMSregistry andcreate_client()factory. - To integrate custom LLMs with AI Scientist v2, append your model identifier to
AVAILABLE_LLMS, add anelifbranch increate_client()to instantiate your client, and use environment variables for secrets. - The
get_response_from_llm()function handles request routing; standard OpenAI-compatible endpoints require no additional payload modifications. - Token tracking via
ai_scientist/utils/token_tracker.pyworks automatically for any properly configured client. - Entry-point scripts like
perform_writeup.pyandperform_ideation_temp_free.pyaccept custom models through the--modelCLI argument.
Frequently Asked Questions
Do I need to modify token tracking code to use a custom LLM?
No. The @track_token_usage decorator in ai_scientist/utils/token_tracker.py automatically intercepts calls from any client that returns a standard response object. As long as your create_client() implementation returns a compatible client (such as an openai.OpenAI instance), usage statistics are recorded without additional code changes.
Can I use different custom models for the small and big model roles in the writeup pipeline?
Yes. Scripts like perform_writeup.py accept both --model (for citation gathering) and --big-model (for final generation) arguments. Register both custom identifiers in AVAILABLE_LLMS and ensure each has a corresponding client branch in create_client(). The pipeline instantiates separate clients for each role automatically.
How do I integrate a vision-language model (VLM) instead of a text-only LLM?
For multimodal models, extend ai_scientist/vlm.py following the same pattern: add your model to the VLM availability list, extend the VLM client factory, and handle image inputs in the VLM-specific request function. The architecture mirrors llm.py but manages base64-encoded images and vision-specific message formats.
What if my custom LLM uses a non-OpenAI API schema?
If your service uses a custom payload structure (e.g., different field names or authentication headers), add a dedicated processing branch inside get_response_from_llm() in ai_scientist/llm.py. Inspect the model parameter to identify your custom type, construct the exact JSON payload required by your endpoint, and extract the response text according to your schema before returning it to the caller.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →