How AI Scientist v2 Gathers Citations for Research Papers: A Deep Dive into Semantic Scholar Integration
AI Scientist v2 gathers citations by querying the Semantic Scholar Graph API through the SemanticScholarSearchTool, which retrieves paper metadata, sorts results by citation count, and injects formatted references into the LLM's ideation loop.
The SakanaAI/AI-Scientist-v2 repository automates research ideation by enabling large language models to discover and cite relevant academic work. To accomplish this, the system implements a dedicated citation gathering mechanism that retrieves real-time paper metadata and citation metrics from external scholarly databases.
The SemanticScholarSearchTool Architecture
Tool Definition and Interface
Located in ai_scientist/tools/semantic_scholar.py, the SemanticScholarSearchTool class inherits from BaseTool (defined in ai_scientist/tools/base_tool.py). It exposes a single required parameter, query, which the LLM populates with search terms relevant to the research topic.
API Authentication and Rate Limits
The tool reads an optional S2_API_KEY environment variable. When present, the key is transmitted as the X-API-KEY header to the Semantic Scholar Graph API, lifting default rate restrictions and enabling higher-volume citation gathering without throttling.
How the Citation Retrieval Pipeline Works
Querying the Semantic Scholar Graph API
The search_for_papers method constructs a GET request to https://api.semanticscholar.org/graph/v1/paper/search. The request explicitly requests the fields title, authors, venue, year, abstract, and citationCount to ensure comprehensive metadata capture for every retrieved paper.
Resilience with Exponential Backoff
To handle transient network failures, the API call is wrapped with backoff.on_exception, implementing exponential back-off for automatic retries on HTTP or connection errors. This ensures reliable citation gathering even under intermittent connectivity or temporary API unavailability.
Ranking by Citation Impact
After retrieving the JSON response, the tool extracts the data array and sorts papers in descending order of citationCount:
papers.sort(key=lambda x: x.get("citationCount", 0), reverse=True)
This prioritizes highly-cited works, ensuring the LLM references influential research rather than obscure publications.
Formatting Results for LLM Consumption
The format_papers method converts each entry into a structured string containing the title, authors, venue, year, citation count, and abstract. This formatted text is optimized for token efficiency and contextual relevance in downstream LLM prompts.
Integrating Citations into the Ideation Workflow
In the free-form ideation pipeline (ai_scientist/perform_ideation_temp_free.py), the tool is instantiated at line 17:
semantic_scholar_tool = SemanticScholarSearchTool()
When the LLM emits the action SearchSemanticScholar, the framework retrieves the tool from tools_dict (around lines 41-45) and executes use_tool(query). The formatted citation list is stored in last_tool_results and injected into subsequent reflection prompts, allowing the model to incorporate up-to-date scholarly references into its research proposals.
Practical Code Examples
Direct tool usage:
from ai_scientist.tools.semantic_scholar import SemanticScholarSearchTool
# Initialize tool (reads S2_API_KEY from environment automatically)
scholar = SemanticScholarSearchTool(max_results=5)
# Retrieve papers on graph neural networks
result_text = scholar.use_tool(query="graph neural networks")
print(result_text)
Integration within the ideation loop:
# Inside perform_ideation_temp_free.py
semantic_scholar_tool = SemanticScholarSearchTool()
# LLM decides to search
action = "SearchSemanticScholar"
arguments = {"query": "self-supervised learning"}
# Framework executes tool lookup and invocation
tool = tools_dict[action]
tool_output = tool.use_tool(**arguments)
# tool_output contains formatted citations for the next LLM prompt
Summary
SemanticScholarSearchToolhandles all citation gathering inai_scientist/tools/semantic_scholar.py.- The tool queries
https://api.semanticscholar.org/graph/v1/paper/searchwith optional API key authentication viaS2_API_KEY. - Results are automatically sorted by
citationCountin descending order to prioritize high-impact papers. - Exponential backoff via
backoff.on_exceptionensures robust API interaction during network instability. - Citations are formatted by
format_papersand injected into the ideation loop inperform_ideation_temp_free.pyto inform LLM reasoning.
Frequently Asked Questions
Does AI Scientist v2 require a Semantic Scholar API key to gather citations?
No, the system functions without an API key, but setting the S2_API_KEY environment variable removes rate limits and prevents throttling during high-volume citation searches. Without the key, requests proceed anonymously but may face restrictive rate limits.
How does the system prioritize which papers to cite?
The search_for_papers method sorts retrieved papers by their citationCount field in descending order, ensuring the LLM references the most influential and highly-cited works first. This ranking happens server-side before formatting and injection into prompts.
Can the citation tool handle API failures or network interruptions?
Yes, the implementation uses backoff.on_exception to automatically retry failed requests with exponential backoff, making the citation gathering process resilient to temporary HTTP errors, timeouts, or connection failures.
Where is the citation data integrated into the research workflow?
The formatted citation output is injected into the reflection prompts during ideation in ai_scientist/perform_ideation_temp_free.py, specifically around lines 41-45, where last_tool_results feeds the LLM's next reasoning step with retrieved scholarly context.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →