# How AI Scientist v2 Gathers Citations for Research Papers: A Deep Dive into Semantic Scholar Integration

> Discover how AI Scientist v2 gathers research paper citations using Semantic Scholar integration. Learn about metadata retrieval, citation sorting, and reference formatting for LLM ideation.

- Repository: [Sakana AI/AI-Scientist-v2](https://github.com/SakanaAI/AI-Scientist-v2)
- Tags: deep-dive
- Published: 2026-03-28

---

**AI Scientist v2 gathers citations by querying the Semantic Scholar Graph API through the `SemanticScholarSearchTool`, which retrieves paper metadata, sorts results by citation count, and injects formatted references into the LLM's ideation loop.**

The SakanaAI/AI-Scientist-v2 repository automates research ideation by enabling large language models to discover and cite relevant academic work. To accomplish this, the system implements a dedicated citation gathering mechanism that retrieves real-time paper metadata and citation metrics from external scholarly databases.

## The SemanticScholarSearchTool Architecture

### Tool Definition and Interface

Located in [`ai_scientist/tools/semantic_scholar.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/tools/semantic_scholar.py), the `SemanticScholarSearchTool` class inherits from `BaseTool` (defined in [`ai_scientist/tools/base_tool.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/tools/base_tool.py)). It exposes a single required parameter, **`query`**, which the LLM populates with search terms relevant to the research topic.

### API Authentication and Rate Limits

The tool reads an optional **`S2_API_KEY`** environment variable. When present, the key is transmitted as the `X-API-KEY` header to the Semantic Scholar Graph API, lifting default rate restrictions and enabling higher-volume citation gathering without throttling.

## How the Citation Retrieval Pipeline Works

### Querying the Semantic Scholar Graph API

The `search_for_papers` method constructs a GET request to `https://api.semanticscholar.org/graph/v1/paper/search`. The request explicitly requests the fields **title**, **authors**, **venue**, **year**, **abstract**, and **citationCount** to ensure comprehensive metadata capture for every retrieved paper.

### Resilience with Exponential Backoff

To handle transient network failures, the API call is wrapped with **`backoff.on_exception`**, implementing exponential back-off for automatic retries on HTTP or connection errors. This ensures reliable citation gathering even under intermittent connectivity or temporary API unavailability.

### Ranking by Citation Impact

After retrieving the JSON response, the tool extracts the `data` array and sorts papers in descending order of **`citationCount`**:

```python
papers.sort(key=lambda x: x.get("citationCount", 0), reverse=True)

```

This prioritizes highly-cited works, ensuring the LLM references influential research rather than obscure publications.

### Formatting Results for LLM Consumption

The **`format_papers`** method converts each entry into a structured string containing the title, authors, venue, year, citation count, and abstract. This formatted text is optimized for token efficiency and contextual relevance in downstream LLM prompts.

## Integrating Citations into the Ideation Workflow

In the free-form ideation pipeline ([`ai_scientist/perform_ideation_temp_free.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/perform_ideation_temp_free.py)), the tool is instantiated at line 17:

```python
semantic_scholar_tool = SemanticScholarSearchTool()

```

When the LLM emits the action **`SearchSemanticScholar`**, the framework retrieves the tool from `tools_dict` (around lines 41-45) and executes `use_tool(query)`. The formatted citation list is stored in `last_tool_results` and injected into subsequent reflection prompts, allowing the model to incorporate up-to-date scholarly references into its research proposals.

## Practical Code Examples

Direct tool usage:

```python
from ai_scientist.tools.semantic_scholar import SemanticScholarSearchTool

# Initialize tool (reads S2_API_KEY from environment automatically)

scholar = SemanticScholarSearchTool(max_results=5)

# Retrieve papers on graph neural networks

result_text = scholar.use_tool(query="graph neural networks")
print(result_text)

```

Integration within the ideation loop:

```python

# Inside perform_ideation_temp_free.py

semantic_scholar_tool = SemanticScholarSearchTool()

# LLM decides to search

action = "SearchSemanticScholar"
arguments = {"query": "self-supervised learning"}

# Framework executes tool lookup and invocation

tool = tools_dict[action]
tool_output = tool.use_tool(**arguments)

# tool_output contains formatted citations for the next LLM prompt

```

## Summary

- **`SemanticScholarSearchTool`** handles all citation gathering in [`ai_scientist/tools/semantic_scholar.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/tools/semantic_scholar.py).
- The tool queries `https://api.semanticscholar.org/graph/v1/paper/search` with optional API key authentication via **`S2_API_KEY`**.
- Results are automatically sorted by **`citationCount`** in descending order to prioritize high-impact papers.
- Exponential backoff via **`backoff.on_exception`** ensures robust API interaction during network instability.
- Citations are formatted by **`format_papers`** and injected into the ideation loop in [`perform_ideation_temp_free.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/perform_ideation_temp_free.py) to inform LLM reasoning.

## Frequently Asked Questions

### Does AI Scientist v2 require a Semantic Scholar API key to gather citations?

No, the system functions without an API key, but setting the **`S2_API_KEY`** environment variable removes rate limits and prevents throttling during high-volume citation searches. Without the key, requests proceed anonymously but may face restrictive rate limits.

### How does the system prioritize which papers to cite?

The `search_for_papers` method sorts retrieved papers by their **`citationCount`** field in descending order, ensuring the LLM references the most influential and highly-cited works first. This ranking happens server-side before formatting and injection into prompts.

### Can the citation tool handle API failures or network interruptions?

Yes, the implementation uses **`backoff.on_exception`** to automatically retry failed requests with exponential backoff, making the citation gathering process resilient to temporary HTTP errors, timeouts, or connection failures.

### Where is the citation data integrated into the research workflow?

The formatted citation output is injected into the reflection prompts during ideation in **[`ai_scientist/perform_ideation_temp_free.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/perform_ideation_temp_free.py)**, specifically around lines 41-45, where `last_tool_results` feeds the LLM's next reasoning step with retrieved scholarly context.