How AI Scientist v2 Uses Semantic Scholar for Literature Search and Automated Citations
AI Scientist v2 integrates Semantic Scholar through a dedicated LLM tool and a standalone wrapper function, enabling autonomous literature discovery during ideation and automated citation harvesting during manuscript generation.
The SakanaAI/AI-Scientist-v2 repository implements an end-to-end autonomous research system that generates scientific hypotheses and produces LaTeX manuscripts. Central to its literature discovery pipeline is a robust integration with the Semantic Scholar Graph API, implemented in ai_scientist/tools/semantic_scholar.py, which provides both tool-based LLM interactions and direct programmatic access for citation management.
Core Tool Implementation
The integration centers on the SemanticScholarSearchTool class, which exposes Semantic Scholar's academic search capabilities as a callable instrument that the LLM can invoke during reasoning.
The SemanticScholarSearchTool Class
In ai_scientist/tools/semantic_scholar.py (lines 19-84), the SemanticScholarSearchTool inherits from BaseTool and registers under the name SearchSemanticScholar (lines 19-27). The class constructor reads the optional S2_API_KEY environment variable (line 38), emitting a warning when absent to signal reliance on stricter public rate limits.
The primary method search_for_papers(self, query: str) (lines 57-85) executes GET requests against the Semantic Scholar Graph API endpoint https://api.semanticscholar.org/graph/v1/paper/search. Results are automatically sorted by citationCount and returned as structured dictionaries (line 84). The class also provides format_papers (lines 87-98), which renders retrieved papers into human-readable strings containing titles, authors, venues, years, citation counts, and abstracts for LLM consumption.
Resilient API Communication
To handle transient network failures, the tool wraps HTTP requests with the backoff decorator (lines 52-56). This automatically retries failed connections with exponential backoff, printing wait times to stderr (lines 12-16) before re-attempting the API call.
from ai_scientist.tools.semantic_scholar import SemanticScholarSearchTool
# Instantiate with optional result limiting
scholar = SemanticScholarSearchTool(max_results=5)
# Execute search - returns formatted string for LLM display
result = scholar.use_tool(query="transformer architectures")
print(result)
Standalone Wrapper for Direct Queries
Beyond the tool interface, the module exposes a convenience function search_for_papers(query, result_limit=10) (lines 101-138) for direct programmatic access. This function bypasses the tool abstraction, returning raw JSON dictionaries suitable for automated citation processing.
Like the class-based implementation, the wrapper respects the S2_API_KEY environment variable (line 105) and employs identical backoff retry logic for resilience against rate limiting.
from ai_scientist.tools.semantic_scholar import search_for_papers
# Fetch raw paper metadata
papers = search_for_papers("diffusion models", result_limit=3)
for paper in papers:
print(f"{paper['title']} ({paper['year']})")
Pipeline Integration
Semantic Scholar serves dual purposes within the AI Scientist v2 workflow: supporting exploratory literature searches during idea generation and enabling automated citation collection during manuscript write-up.
Literature Discovery During Ideation
In ai_scientist/perform_ideation_temp_free.py (lines 20-26), the system instantiates SemanticScholarSearchTool() and appends it to the available tools list. The LLM system prompt enumerates the tool's name and description (lines 65-68), enabling the model to issue ACTION: "SearchSemanticScholar" commands with JSON payloads containing "query" fields.
When the LLM invokes the tool, the framework parses the response, executes tool.use_tool(**arguments_json), and feeds the returned paper list back as last_tool_results (lines 98-104, 108-112). This creates a feedback loop where the LLM can iteratively refine literature searches while brainstorming research directions.
Automated Citation Collection
During the write-up stage, ai_scientist/perform_writeup.py directly imports the wrapper function (line 19). After the LLM proposes a citation query, the system calls papers = search_for_papers(query) (line 61), then formats the retrieved metadata into BibTeX entries for insertion into the LaTeX references.bib block (lines 71-84).
This automation eliminates manual citation gathering, ensuring the generated manuscript includes relevant, verifiable academic sources without human intervention.
Configuration and Environment Setup
Both the tool class and wrapper function source authentication credentials from the S2_API_KEY environment variable (lines 38 and 105). When this variable is absent, the system emits a runtime warning but continues execution against the public API endpoint, subject to Semantic Scholar's unauthenticated rate limits.
The optional API key enables higher throughput for intensive literature searches during large-scale idea generation campaigns.
Summary
- AI Scientist v2 treats Semantic Scholar as a first-class external knowledge source through the
SemanticScholarSearchToolclass inai_scientist/tools/semantic_scholar.py. - The
search_for_papersmethod (lines 57-85) queries the Graph API with automatic retry logic via thebackoffdecorator. - During ideation, the LLM calls the tool via
ACTION: "SearchSemanticScholar"to discover relevant literature iteratively. - During write-up, the standalone wrapper function harvests citations automatically for BibTeX generation.
- Authentication relies on the
S2_API_KEYenvironment variable, with graceful degradation to public rate limits when unspecified.
Frequently Asked Questions
How does AI Scientist v2 authenticate with Semantic Scholar?
The system reads the S2_API_KEY environment variable in both SemanticScholarSearchTool (line 38) and the search_for_papers wrapper (line 105). If absent, it emits a warning and proceeds with unauthenticated requests against the public API endpoint, which imposes stricter rate limits.
What happens if the Semantic Scholar API returns an error?
Both the tool and wrapper implement exponential backoff retry logic using the backoff decorator (lines 52-56). This automatically retries failed HTTP or connection errors, printing wait times to stderr, ensuring transient failures do not halt the research pipeline.
Can I use the Semantic Scholar integration outside of the AI Scientist pipeline?
Yes. The module-level search_for_papers(query, result_limit=10) function (lines 101-138) provides a standalone interface that returns raw paper dictionaries without requiring LLM tool instantiation, suitable for any Python script requiring academic literature metadata.
How does the LLM know when to search Semantic Scholar?
During ideation in perform_ideation_temp_free.py, the LLM receives a system prompt (lines 65-68) describing the SearchSemanticScholar tool. The model can then output ACTION: "SearchSemanticScholar" with a query parameter, which the framework detects and routes to SemanticScholarSearchTool.use_tool() (lines 98-112).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →