How to Integrate Custom LLMs with AI Scientist v2: The Complete Developer Guide

Integrating custom LLMs with AI Scientist v2 requires registering your model identifier in the AVAILABLE_LLMS list within ai_scientist/llm.py, extending the create_client() function to instantiate your specific SDK client, and optionally adjusting request payload handling in get_response_from_llm().

AI Scientist v2 from SakanaAI provides a modular architecture for automating scientific research workflows using language models. To integrate custom LLMs with AI Scientist v2—whether hosted locally, via a private API, or through specialized providers—you must extend the thin abstraction layer defined in the core LLM module. This guide walks through the exact source code modifications needed to connect proprietary or self-hosted models to the ideation, review, and writeup pipelines.

Understanding the LLM Abstraction Architecture

The integration point lives in ai_scientist/llm.py, which implements a three-stage workflow:

  • Model Declaration – The AVAILABLE_LLMS list (defined around lines 13–74) contains every supported model identifier, such as "gpt-4o-mini" or "bedrock/anthropic.claude-3-sonnet-20240229-v1:0". These strings are validated against CLI arguments across all entry-point scripts.

  • Client Instantiation – The create_client(model) function maps identifiers to SDK instances. It branches by prefix (e.g., "ollama/", "gpt", "claude") to return the appropriate client object and normalized model name.

  • Request Execution – The get_response_from_llm() function (spanning roughly lines 80–440) constructs payloads, manages message history, and parses responses. It contains provider-specific logic for OpenAI, Anthropic, Gemini, and other families.

Step-by-Step Guide to Custom LLM Integration

Step 1: Register the Model Identifier

Append your custom model string to AVAILABLE_LLMS in ai_scientist/llm.py. Use a descriptive prefix to avoid collisions:


# In ai_scientist/llm.py

AVAILABLE_LLMS.append("myco/custom-llm")

This registration enables the --model CLI argument in scripts like perform_ideation_temp_free.py and perform_writeup.py to recognize your model.

Step 2: Create a Client Branch

Add a conditional branch in create_client() to instantiate your HTTP client. For OpenAI-compatible endpoints, reuse the existing openai library:

elif model.startswith("myco/"):
    endpoint = os.getenv("MYCO_ENDPOINT", "https://api.myco.com/v1")
    api_key = os.getenv("MYCO_API_KEY")
    print(f"Using MyCo endpoint {endpoint} for model {model}.")
    return openai.OpenAI(api_key=api_key, base_url=endpoint), model

Return a tuple of (client_instance, model_string) to maintain compatibility with downstream callers.

Step 3: Adjust Request Handling (Optional)

If your LLM service uses a custom request/response schema, add a dedicated processing branch inside get_response_from_llm(). For standard OpenAI chat completions, the existing logic handles your client automatically without modification.

Step 4: Configure Environment Variables

Store credentials and endpoints outside source control. Define variables like MYCO_ENDPOINT and MYCO_API_KEY, then document them in your .env.example or README. This aligns with the repository's security patterns used for OpenAI and Anthropic keys.

Step 5: Test the Integration

Validate connectivity by running an ideation script:

python -m ai_scientist.perform_ideation_temp_free \
    --model myco/custom-llm \
    --max-num-generations 1

The @track_token_usage decorator in ai_scientist/utils/token_tracker.py automatically records usage statistics without requiring code changes.

Practical Example: Local OpenAI-Compatible Server

To connect a local Llama 3 instance running on http://localhost:8000/v1:


# Add to AVAILABLE_LLMS in ai_scientist/llm.py

AVAILABLE_LLMS.append("local/llama-3-8b")

# Extend create_client with a local branch

elif model.startswith("local/"):
    endpoint = os.getenv("LOCAL_LLM_ENDPOINT", "http://localhost:8000/v1")
    print(f"Using local LLM at {endpoint}")
    return openai.OpenAI(api_key="not-needed", base_url=endpoint), model

Invoke the writeup pipeline:

python -m ai_scientist.perform_writeup \
    --folder my_project \
    --model local/llama-3-8b \
    --big-model gpt-4o-2024-05-13

Code Examples for Production Integration

Minimal Reproducible Integration Snippet

import os
from ai_scientist.llm import create_client, get_response_from_llm, AVAILABLE_LLMS

# 1️⃣ Register model (do this before CLI parsing occurs)

AVAILABLE_LLMS.append("myco/custom-llm")

# 2️⃣ (Inject the elif branch into ai_scientist/llm.py as shown above)

# 3️⃣ Use the model programmatically

client, client_model = create_client("myco/custom-llm")
system_msg = "You are a helpful research assistant."
prompt = "Summarize the key contributions of the paper 'Attention Is All You Need'."

answer, history = get_response_from_llm(
    prompt=prompt,
    client=client,
    model=client_model,
    system_message=system_msg,
    temperature=0.0,
)
print("LLM answer:", answer)

Running Built-in Scripts with Custom Models

export MYCO_ENDPOINT="https://api.myco.com/v1"
export MYCO_API_KEY="sk-****************"

python -m ai_scientist.perform_ideation_temp_free \
    --model myco/custom-llm \
    --max-num-generations 2 \
    --workshop-file ideas/workshop.md

Summary

  • AI Scientist v2 centralizes LLM integration in ai_scientist/llm.py through the AVAILABLE_LLMS registry and create_client() factory.
  • To integrate custom LLMs with AI Scientist v2, append your model identifier to AVAILABLE_LLMS, add an elif branch in create_client() to instantiate your client, and use environment variables for secrets.
  • The get_response_from_llm() function handles request routing; standard OpenAI-compatible endpoints require no additional payload modifications.
  • Token tracking via ai_scientist/utils/token_tracker.py works automatically for any properly configured client.
  • Entry-point scripts like perform_writeup.py and perform_ideation_temp_free.py accept custom models through the --model CLI argument.

Frequently Asked Questions

Do I need to modify token tracking code to use a custom LLM?

No. The @track_token_usage decorator in ai_scientist/utils/token_tracker.py automatically intercepts calls from any client that returns a standard response object. As long as your create_client() implementation returns a compatible client (such as an openai.OpenAI instance), usage statistics are recorded without additional code changes.

Can I use different custom models for the small and big model roles in the writeup pipeline?

Yes. Scripts like perform_writeup.py accept both --model (for citation gathering) and --big-model (for final generation) arguments. Register both custom identifiers in AVAILABLE_LLMS and ensure each has a corresponding client branch in create_client(). The pipeline instantiates separate clients for each role automatically.

How do I integrate a vision-language model (VLM) instead of a text-only LLM?

For multimodal models, extend ai_scientist/vlm.py following the same pattern: add your model to the VLM availability list, extend the VLM client factory, and handle image inputs in the VLM-specific request function. The architecture mirrors llm.py but manages base64-encoded images and vision-specific message formats.

What if my custom LLM uses a non-OpenAI API schema?

If your service uses a custom payload structure (e.g., different field names or authentication headers), add a dedicated processing branch inside get_response_from_llm() in ai_scientist/llm.py. Inspect the model parameter to identify your custom type, construct the exact JSON payload required by your endpoint, and extract the response text according to your schema before returning it to the caller.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →