How AI Scientist v2 Gathers Citations for Research Papers: A Deep Dive into Semantic Scholar Integration

AI Scientist v2 gathers citations by querying the Semantic Scholar Graph API through the SemanticScholarSearchTool, which retrieves paper metadata, sorts results by citation count, and injects formatted references into the LLM's ideation loop.

The SakanaAI/AI-Scientist-v2 repository automates research ideation by enabling large language models to discover and cite relevant academic work. To accomplish this, the system implements a dedicated citation gathering mechanism that retrieves real-time paper metadata and citation metrics from external scholarly databases.

The SemanticScholarSearchTool Architecture

Tool Definition and Interface

Located in ai_scientist/tools/semantic_scholar.py, the SemanticScholarSearchTool class inherits from BaseTool (defined in ai_scientist/tools/base_tool.py). It exposes a single required parameter, query, which the LLM populates with search terms relevant to the research topic.

API Authentication and Rate Limits

The tool reads an optional S2_API_KEY environment variable. When present, the key is transmitted as the X-API-KEY header to the Semantic Scholar Graph API, lifting default rate restrictions and enabling higher-volume citation gathering without throttling.

How the Citation Retrieval Pipeline Works

Querying the Semantic Scholar Graph API

The search_for_papers method constructs a GET request to https://api.semanticscholar.org/graph/v1/paper/search. The request explicitly requests the fields title, authors, venue, year, abstract, and citationCount to ensure comprehensive metadata capture for every retrieved paper.

Resilience with Exponential Backoff

To handle transient network failures, the API call is wrapped with backoff.on_exception, implementing exponential back-off for automatic retries on HTTP or connection errors. This ensures reliable citation gathering even under intermittent connectivity or temporary API unavailability.

Ranking by Citation Impact

After retrieving the JSON response, the tool extracts the data array and sorts papers in descending order of citationCount:

papers.sort(key=lambda x: x.get("citationCount", 0), reverse=True)

This prioritizes highly-cited works, ensuring the LLM references influential research rather than obscure publications.

Formatting Results for LLM Consumption

The format_papers method converts each entry into a structured string containing the title, authors, venue, year, citation count, and abstract. This formatted text is optimized for token efficiency and contextual relevance in downstream LLM prompts.

Integrating Citations into the Ideation Workflow

In the free-form ideation pipeline (ai_scientist/perform_ideation_temp_free.py), the tool is instantiated at line 17:

semantic_scholar_tool = SemanticScholarSearchTool()

When the LLM emits the action SearchSemanticScholar, the framework retrieves the tool from tools_dict (around lines 41-45) and executes use_tool(query). The formatted citation list is stored in last_tool_results and injected into subsequent reflection prompts, allowing the model to incorporate up-to-date scholarly references into its research proposals.

Practical Code Examples

Direct tool usage:

from ai_scientist.tools.semantic_scholar import SemanticScholarSearchTool

# Initialize tool (reads S2_API_KEY from environment automatically)

scholar = SemanticScholarSearchTool(max_results=5)

# Retrieve papers on graph neural networks

result_text = scholar.use_tool(query="graph neural networks")
print(result_text)

Integration within the ideation loop:


# Inside perform_ideation_temp_free.py

semantic_scholar_tool = SemanticScholarSearchTool()

# LLM decides to search

action = "SearchSemanticScholar"
arguments = {"query": "self-supervised learning"}

# Framework executes tool lookup and invocation

tool = tools_dict[action]
tool_output = tool.use_tool(**arguments)

# tool_output contains formatted citations for the next LLM prompt

Summary

  • SemanticScholarSearchTool handles all citation gathering in ai_scientist/tools/semantic_scholar.py.
  • The tool queries https://api.semanticscholar.org/graph/v1/paper/search with optional API key authentication via S2_API_KEY.
  • Results are automatically sorted by citationCount in descending order to prioritize high-impact papers.
  • Exponential backoff via backoff.on_exception ensures robust API interaction during network instability.
  • Citations are formatted by format_papers and injected into the ideation loop in perform_ideation_temp_free.py to inform LLM reasoning.

Frequently Asked Questions

Does AI Scientist v2 require a Semantic Scholar API key to gather citations?

No, the system functions without an API key, but setting the S2_API_KEY environment variable removes rate limits and prevents throttling during high-volume citation searches. Without the key, requests proceed anonymously but may face restrictive rate limits.

How does the system prioritize which papers to cite?

The search_for_papers method sorts retrieved papers by their citationCount field in descending order, ensuring the LLM references the most influential and highly-cited works first. This ranking happens server-side before formatting and injection into prompts.

Can the citation tool handle API failures or network interruptions?

Yes, the implementation uses backoff.on_exception to automatically retry failed requests with exponential backoff, making the citation gathering process resilient to temporary HTTP errors, timeouts, or connection failures.

Where is the citation data integrated into the research workflow?

The formatted citation output is injected into the reflection prompts during ideation in ai_scientist/perform_ideation_temp_free.py, specifically around lines 41-45, where last_tool_results feeds the LLM's next reasoning step with retrieved scholarly context.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →