# Understanding the Purpose of ai_scientist/llm.py in AI-Scientist-v2

> Discover the central role of ai_scientist/llm.py in AI-Scientist-v2. This module unifies LLM interactions, manages client creation, routing, retries, and output parsing for seamless AI development.

- Repository: [Sakana AI/AI-Scientist-v2](https://github.com/SakanaAI/AI-Scientist-v2)
- Tags: deep-dive
- Published: 2026-03-28

---

**The [`ai_scientist/llm.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/llm.py) module serves as the central abstraction layer that unifies all Large Language Model interactions across the AI-Scientist-v2 codebase, handling provider-specific client creation, request routing, automatic retries, token tracking, and structured output parsing.**

The SakanaAI/AI-Scientist-v2 project relies on [`ai_scientist/llm.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/llm.py) to eliminate provider fragmentation when interacting with models from Anthropic, OpenAI, Google, and local Ollama instances. This file consolidates SDK differences, authentication logic, and error handling into a single location, allowing the scientific workflow modules to invoke LLMs without managing vendor-specific quirks or retry logic.

## Core Responsibilities of ai_scientist/llm.py

### Centralized Model Registry and Client Factory

At lines 13–73, the `AVAILABLE_LLMS` list enumerates every supported model, from Anthropic Claude variants to OpenAI GPT-4o, Gemini, DeepSeek, Llama, and Ollama endpoints. The `create_client()` function (lines 480–524) acts as a factory that instantiates the correct SDK client—whether `anthropic.Anthropic()`, `openai.OpenAI()`, or a custom HTTP handler—and resolves the exact model identifier each provider expects.

### Unified Request API

The module exposes three primary functions that present a single interface regardless of backend:

- **`make_llm_call()`** (lines 151–162) – Low-level dispatcher that routes to the appropriate SDK method.
- **`get_response_from_llm()`** (lines 267–399) – High-level wrapper for single-turn conversations with history, temperature, and token limit management.
- **`get_batch_responses_from_llm()`** (lines 76–124) – Parallel interface for requesting multiple completions (ensembling) in one call.

All three handle conversation history serialization, parameter validation, and response parsing uniformly.

### Automatic Retry and Backoff Mechanism

Long-running scientific pipelines cannot afford to crash on transient API failures. The `@backoff.on_exception` decorator (applied at lines 76–85 and 267–275) automatically retries calls when rate limits, timeouts, or server errors occur. This decorator wraps both `get_batch_responses_from_llm` and `get_response_from_llm`, implementing exponential backoff without cluttering the business logic.

### Token Usage Tracking

Every public call function is wrapped with `@track_token_usage` imported from `ai_scientist.utils.token_tracker` (visible at lines 86, 151, and 267). This decorator intercepts responses to record prompt and completion token counts, feeding real-time usage metrics back to the orchestrator for cost monitoring and budget enforcement.

### Model-Agnostic JSON Extraction

LLMs often return structured data wrapped in markdown code blocks or explanatory text. The `extract_json_between_markers()` utility (lines 452–477) safely extracts JSON snippets from raw strings, handling malformed outputs and missing delimiters gracefully. This ensures that downstream modules receive valid Python dictionaries regardless of the model’s formatting inconsistencies.

## Practical Code Examples

### Creating a Client for Any Supported Model

```python
from ai_scientist.llm import create_client

# Instantiate an Anthropic client for Claude 3.5 Sonnet

client, model_id = create_client("claude-3-5-sonnet-20240620")

# client → anthropic.Anthropic() instance

# model_id → "claude-3-5-sonnet-20240620"

```

### Sending Single Prompts with Conversation History

```python
from ai_scientist.llm import create_client, get_response_from_llm

client, model = create_client("gpt-4o-mini")
system_msg = "You are an AI research assistant."
prompt = "Summarize the key contributions of the paper 'Attention Is All You Need'."

answer, history = get_response_from_llm(
    prompt=prompt,
    client=client,
    model=model,
    system_message=system_msg,
    temperature=0.0,
    print_debug=True,
)
print(answer)

```

*Behind the scenes*, the function selects `client.chat.completions.create` for OpenAI, injects the system message, enforces token limits, and records usage via the tracking decorator.

### Generating Multiple Responses for Ensembling

```python
from ai_scientist.llm import create_client, get_batch_responses_from_llm

client, model = create_client("ollama/qwen3")
system_msg = "You are a concise summarizer."
prompt = "Explain the concept of reinforcement learning in two sentences."

responses, histories = get_batch_responses_from_llm(
    prompt=prompt,
    client=client,
    model=model,
    system_message=system_msg,
    temperature=0.7,
    n_responses=3,
)
for i, r in enumerate(responses, 1):
    print(f"=== Response {i} ===\n{r}\n")

```

### Extracting Structured JSON from LLM Output

```python
from ai_scientist.llm import extract_json_between_markers

raw_output = """
Here is the JSON you asked for:

```json
{
  "title": "AI Scientist",
  "version": "v2"
}

```

Enjoy!
"""
data = extract_json_between_markers(raw_output)
print(data)  # {'title': 'AI Scientist', 'version': 'v2'}

```

## Integration with the AI-Scientist Workflow

The abstraction provided by [`ai_scientist/llm.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/llm.py) allows higher-level modules to remain provider-agnostic:

- **[`ai_scientist/perform_llm_review.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/perform_llm_review.py)** – Uses `get_batch_responses_from_llm` to generate multiple review drafts simultaneously, then extracts structured metadata via `extract_json_between_markers`.
- **[`ai_scientist/perform_writeup.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/perform_writeup.py)** – Calls `get_response_from_llm` iteratively to flesh out individual sections of a scientific paper while relying on the built-in retry logic for long generation tasks.
- **[`ai_scientist/treesearch/log_summarization.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/log_summarization.py)** – Demonstrates how log analysis pipelines retrieve and parse LLM outputs without handling raw SDK responses.
- **[`ai_scientist/utils/token_tracker.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/utils/token_tracker.py)** – Supplies the `@track_token_usage` decorator that enables cost accounting across every module listed above.

## Summary

- **[`ai_scientist/llm.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/llm.py)** centralizes all LLM interactions for the SakanaAI/AI-Scientist-v2 repository, exposing a unified interface over disparate provider SDKs.
- The **`create_client()`** factory (lines 480–524) and **`AVAILABLE_LLMS`** registry (lines 13–73) support Anthropic, OpenAI, Gemini, Ollama, and other backends.
- **`get_response_from_llm()`** and **`get_batch_responses_from_llm()`** provide single and batched inference with automatic retry logic via the `@backoff.on_exception` decorator.
- The **`@track_token_usage`** decorator (lines 86, 151, 267) integrates with [`ai_scientist/utils/token_tracker.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/utils/token_tracker.py) to record consumption for every call.
- **`extract_json_between_markers()`** (lines 452–477) ensures robust parsing of structured model outputs regardless of formatting variations.

## Frequently Asked Questions

### Which LLM providers are supported by ai_scientist/llm.py?

According to the `AVAILABLE_LLMS` definition at lines 13–73, the module supports Anthropic Claude (all variants), OpenAI GPT-4o and GPT-4o-mini, Google Gemini, DeepSeek, Llama models via Ollama, and additional local or remote endpoints. The `create_client()` function maps these friendly names to their respective SDK implementations.

### How does the retry mechanism handle rate limiting?

The `@backoff.on_exception` decorator (visible at lines 76–85 and 267–275) wraps the main call functions and catches rate-limit, timeout, and server errors. It implements exponential backoff automatically, ensuring that transient API failures do not terminate long-running scientific experiments while respecting provider rate limits.

### What is the difference between get_response_from_llm and get_batch_responses_from_llm?

**`get_response_from_llm()`** (lines 267–399) returns a single completion along with the updated conversation history, suitable for sequential dialogue or single-shot generation. **`get_batch_responses_from_llm()`** (lines 76–124) accepts an `n_responses` parameter and returns a list of independent completions, enabling ensemble methods or majority-vote scoring within the review pipeline.

### How is token usage tracked across the system?

Every public entry point—`get_batch_responses_from_llm`, `make_llm_call`, and `get_response_from_llm`—is decorated with `@track_token_usage` (applied at lines 86, 151, and 267). This decorator, defined in [`ai_scientist/utils/token_tracker.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/utils/token_tracker.py), intercepts API responses to extract `prompt_tokens` and `completion_tokens`, then logs these metrics to a central tracker for real-time cost monitoring and budget enforcement.