How to Perform Ideation with AI Scientist v2: Automated Research Generation Guide
AI Scientist v2 performs fully automated ideation through a template-free, LLM-driven loop in ai_scientist/perform_ideation_temp_free.py that generates research proposals using literature search tools and multi-round reflection before finalizing structured JSON outputs.
AI Scientist v2 implements a comprehensive research ideation pipeline that autonomously generates, refines, and finalizes scientific proposals. To perform ideation with AI Scientist v2, you interact with the generate_temp_free_idea function, which orchestrates a multi-step workflow combining LLM reasoning with real-time literature search capabilities via Semantic Scholar.
CLI Configuration and Entry Points
The ideation process starts through the command-line interface defined in ai_scientist/perform_ideation_temp_free.py. The entry point accepts several key parameters that control the generation behavior:
if __name__ == "__main__":
parser = argparse.ArgumentParser(...)
parser.add_argument("--model", default="gpt-4o-2024-05-13", choices=AVAILABLE_LLMS)
parser.add_argument("--workshop-file", default="ideas/i_cant_believe_its_not_better.md")
parser.add_argument("--max-num-generations", type=int, default=1)
parser.add_argument("--num-reflections", type=int, default=5)
args = parser.parse_args()
You must provide a workshop description as a markdown file that defines your research domain and constraints. The --max-num-generations parameter controls how many distinct proposals the system attempts to create, while --num-reflections determines how many refinement rounds each proposal undergoes before finalization.
The Ideation Engine Architecture
LLM Client Initialization and Token Tracking
Before generation begins, the system creates an LLM client via ai_scientist/llm.create_client, which abstracts over multiple providers including OpenAI, Anthropic, Ollama, and Gemini. All LLM calls are wrapped with the @track_token_usage decorator from ai_scientist/utils/token_tracker.py, providing transparent cost monitoring throughout the ideation process.
The Generation Loop: generate_temp_free_idea
The core ideation logic resides in the generate_temp_free_idea function:
def generate_temp_free_idea(
idea_fname, client, model, workshop_description,
max_num_generations=20, num_reflections=5, reload_ideas=True,
) -> List[Dict]:
This function implements an iterative generation strategy with the following characteristics:
- Idea archival: If
idea_fnamepoints to an existing JSON file, previous proposals load automatically to provide context for new generations - Multi-generation support: The loop runs up to
max_num_generationstimes, creating distinct research directions - Reflection rounds: Each proposal undergoes
num_reflectionsiterative refinement cycles before finalization
Prompt Construction and LLM Interaction
The system constructs three distinct prompt types for each generation cycle:
- System prompt: A static instruction defining the LLM as an "experienced AI researcher" and listing available tools
- Idea generation prompt: Combines the workshop description with previously generated ideas (
prev_ideas_string) to ensure novelty - Reflection prompt: Incorporates tool outputs (
last_tool_results) and reflection round counters to guide iterative improvement
All prompts route through get_response_from_llm in ai_scientist/llm.py, which handles model-specific request formatting and response parsing.
Tool-Augmented Reasoning: Search and Finalize
AI Scientist v2 exposes two specialized tools to the LLM during ideation, both inheriting from the abstract BaseTool class in ai_scientist/tools/base_tool.py:
-
SearchSemanticScholar: Implemented in
ai_scientist/tools/semantic_scholar.py, this tool queries the Semantic Scholar Graph API viasearch_for_papersto retrieve relevant literature. Results sort by citation count and format into human-readable lists for the LLM's next reflection round. -
FinalizeIdea: A signaling mechanism (not a concrete class) that triggers when the LLM returns an
action == "FinalizeIdea"response, indicating the proposal meets quality thresholds.
The LLM must format tool calls using a specific syntax:
ACTION: SearchSemanticScholar
ARGUMENTS:
{
"query": "transformer architecture efficiency"
}
The engine extracts ACTION and ARGUMENTS using regular expressions (lines 181-190 in perform_ideation_temp_free.py), validates against registered tools, and invokes tool.use_tool(**arguments_json) for valid actions.
Parsing and Persistence
When the LLM returns FinalizeIdea, the system parses the accompanying JSON payload containing structured research proposal fields:
{
"idea": {
"Name": "efficient_attention_mechanism",
"Title": "Reducing Computational Complexity in Transformer Attention",
"Short Hypothesis": "If we apply dynamic sparsity patterns to attention matrices, then we can reduce FLOPs by 40% without accuracy loss",
"Related Work": "Prior work by Smith et al. (2023) explored static sparsity...",
"Abstract": "A 250-word summary of the proposed approach...",
"Experiments": "1. Benchmark on GLUE tasks. 2. Compare against dense baseline...",
"Risk Factors and Limitations": "Potential issues with implementation complexity..."
}
}
The engine appends finalized proposals to idea_str_archive, which serializes to a pretty-printed JSON file at idea_fname upon completion. This file serves as the canonical record of all generated research ideas.
End-to-End Workflow
The complete ideation process follows this sequence:
- Load the workshop description markdown file
- Initialize the LLM client for the specified model (default:
gpt-4o-2024-05-13) - Iterate through generations:
- Build context-aware prompts → LLM call → parse ACTION/ARGUMENTS
- Execute
SearchSemanticScholarqueries when requested, feeding literature results back into the reflection loop - Capture
FinalizeIdeapayloads and validate JSON schema compliance - Perform
num_reflectionsrefinement rounds per proposal
- Serialize all finalized ideas to the output JSON file
Practical Examples
Command-Line Execution
Run ideation directly from the terminal to generate three research proposals with five reflection rounds each:
python -m ai_scientist.perform_ideation_temp_free \
--model gpt-4o-2024-05-13 \
--workshop-file ideas/i_cant_believe_its_not_better.md \
--max-num-generations 3 \
--num-reflections 5
This creates ideas/i_cant_believe_its_not_better.json containing up to three complete research proposals.
Programmatic Python Usage
Integrate ideation into custom pipelines by importing the core functions directly:
from ai_scientist.llm import create_client
from ai_scientist.perform_ideation_temp_free import generate_temp_free_idea
# Load workshop description
with open("ideas/i_cant_believe_its_not_better.md") as f:
workshop = f.read()
# Initialize LLM client
client, model = create_client("gpt-4o-2024-05-13")
# Generate ideas with custom parameters
ideas = generate_temp_free_idea(
idea_fname="ideas/output.json",
client=client,
model=model,
workshop_description=workshop,
max_num_generations=2,
num_reflections=4,
)
print(f"Generated {len(ideas)} research proposals")
Expected Output Schema
The finalized JSON structure follows a strict schema required for downstream processing in AI Scientist v2:
{
"idea": {
"Name": "unique_identifier",
"Title": "Descriptive Research Title",
"Short Hypothesis": "Testable causal statement",
"Related Work": "Literature context and gaps",
"Abstract": "Comprehensive summary",
"Experiments": "Numbered experimental plan",
"Risk Factors and Limitations": "Critical analysis of constraints"
}
}
Summary
- AI Scientist v2 implements template-free ideation through
ai_scientist/perform_ideation_temp_free.py, utilizing a multi-generational, reflection-based workflow. - The LLM client factory in
ai_scientist/llm.pysupports multiple providers (OpenAI, Anthropic, Ollama, Gemini) with integrated token tracking viaai_scientist/utils/token_tracker.py. - Tool augmentation enables literature-aware proposals through
SearchSemanticScholarinai_scientist/tools/semantic_scholar.py, which queries the Semantic Scholar Graph API and returns citation-ranked results. - The system accepts workshop descriptions as markdown files and outputs structured JSON proposals containing standardized research fields (Name, Title, Hypothesis, Abstract, Experiments, Risks).
- Control parameters include
max_num_generations(total proposals) andnum_reflections(refinement depth per proposal).
Frequently Asked Questions
What is the difference between max_num_generations and num_reflections?
The max_num_generations parameter controls how many distinct research proposals the system creates in a single run, while num_reflections specifies how many iterative improvement cycles each individual proposal undergoes before finalization. For example, with max_num_generations=3 and num_reflections=5, the system generates three separate ideas, each refined through five rounds of literature search and self-critique.
Which workshop file format does AI Scientist v2 expect?
AI Scientist v2 expects workshop descriptions as markdown files (.md) containing domain context, research constraints, and preferred methodological approaches. The default location is ideas/i_cant_believe_its_not_better.md, though you can specify any path via the --workshop-file argument. The content should describe the research area sufficiently to guide the LLM toward relevant literature and feasible experimental designs.
How does the Semantic Scholar integration work during ideation?
The SearchSemanticScholar tool, implemented in ai_scientist/tools/semantic_scholar.py, activates when the LLM outputs ACTION: SearchSemanticScholar with query arguments. The tool calls search_for_papers against the Semantic Scholar Graph API, optionally using an S2_API_KEY for higher rate limits. Results sort by citation count and format into a human-readable list that feeds back into the next reflection round, enabling literature-aware hypothesis generation.
Can I use local models instead of commercial APIs for ideation?
Yes, the create_client function in ai_scientist/llm.py supports multiple providers including Ollama for local model hosting. Specify your local model identifier (e.g., --model ollama/llama3) to route requests to your local inference server instead of OpenAI or Anthropic endpoints, though you must ensure your local model supports the tool-calling format required for SearchSemanticScholar and FinalizeIdea actions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →