How to Perform Ideation with AI Scientist v2: Automated Research Generation Guide

AI Scientist v2 performs fully automated ideation through a template-free, LLM-driven loop in ai_scientist/perform_ideation_temp_free.py that generates research proposals using literature search tools and multi-round reflection before finalizing structured JSON outputs.

AI Scientist v2 implements a comprehensive research ideation pipeline that autonomously generates, refines, and finalizes scientific proposals. To perform ideation with AI Scientist v2, you interact with the generate_temp_free_idea function, which orchestrates a multi-step workflow combining LLM reasoning with real-time literature search capabilities via Semantic Scholar.

CLI Configuration and Entry Points

The ideation process starts through the command-line interface defined in ai_scientist/perform_ideation_temp_free.py. The entry point accepts several key parameters that control the generation behavior:

if __name__ == "__main__":
    parser = argparse.ArgumentParser(...)
    parser.add_argument("--model", default="gpt-4o-2024-05-13", choices=AVAILABLE_LLMS)
    parser.add_argument("--workshop-file", default="ideas/i_cant_believe_its_not_better.md")
    parser.add_argument("--max-num-generations", type=int, default=1)
    parser.add_argument("--num-reflections", type=int, default=5)
    args = parser.parse_args()

You must provide a workshop description as a markdown file that defines your research domain and constraints. The --max-num-generations parameter controls how many distinct proposals the system attempts to create, while --num-reflections determines how many refinement rounds each proposal undergoes before finalization.

The Ideation Engine Architecture

LLM Client Initialization and Token Tracking

Before generation begins, the system creates an LLM client via ai_scientist/llm.create_client, which abstracts over multiple providers including OpenAI, Anthropic, Ollama, and Gemini. All LLM calls are wrapped with the @track_token_usage decorator from ai_scientist/utils/token_tracker.py, providing transparent cost monitoring throughout the ideation process.

The Generation Loop: generate_temp_free_idea

The core ideation logic resides in the generate_temp_free_idea function:

def generate_temp_free_idea(
    idea_fname, client, model, workshop_description,
    max_num_generations=20, num_reflections=5, reload_ideas=True,
) -> List[Dict]:

This function implements an iterative generation strategy with the following characteristics:

  • Idea archival: If idea_fname points to an existing JSON file, previous proposals load automatically to provide context for new generations
  • Multi-generation support: The loop runs up to max_num_generations times, creating distinct research directions
  • Reflection rounds: Each proposal undergoes num_reflections iterative refinement cycles before finalization

Prompt Construction and LLM Interaction

The system constructs three distinct prompt types for each generation cycle:

  1. System prompt: A static instruction defining the LLM as an "experienced AI researcher" and listing available tools
  2. Idea generation prompt: Combines the workshop description with previously generated ideas (prev_ideas_string) to ensure novelty
  3. Reflection prompt: Incorporates tool outputs (last_tool_results) and reflection round counters to guide iterative improvement

All prompts route through get_response_from_llm in ai_scientist/llm.py, which handles model-specific request formatting and response parsing.

Tool-Augmented Reasoning: Search and Finalize

AI Scientist v2 exposes two specialized tools to the LLM during ideation, both inheriting from the abstract BaseTool class in ai_scientist/tools/base_tool.py:

  • SearchSemanticScholar: Implemented in ai_scientist/tools/semantic_scholar.py, this tool queries the Semantic Scholar Graph API via search_for_papers to retrieve relevant literature. Results sort by citation count and format into human-readable lists for the LLM's next reflection round.

  • FinalizeIdea: A signaling mechanism (not a concrete class) that triggers when the LLM returns an action == "FinalizeIdea" response, indicating the proposal meets quality thresholds.

The LLM must format tool calls using a specific syntax:


ACTION: SearchSemanticScholar
ARGUMENTS:
{
  "query": "transformer architecture efficiency"
}

The engine extracts ACTION and ARGUMENTS using regular expressions (lines 181-190 in perform_ideation_temp_free.py), validates against registered tools, and invokes tool.use_tool(**arguments_json) for valid actions.

Parsing and Persistence

When the LLM returns FinalizeIdea, the system parses the accompanying JSON payload containing structured research proposal fields:

{
  "idea": {
    "Name": "efficient_attention_mechanism",
    "Title": "Reducing Computational Complexity in Transformer Attention",
    "Short Hypothesis": "If we apply dynamic sparsity patterns to attention matrices, then we can reduce FLOPs by 40% without accuracy loss",
    "Related Work": "Prior work by Smith et al. (2023) explored static sparsity...",
    "Abstract": "A 250-word summary of the proposed approach...",
    "Experiments": "1. Benchmark on GLUE tasks. 2. Compare against dense baseline...",
    "Risk Factors and Limitations": "Potential issues with implementation complexity..."
  }
}

The engine appends finalized proposals to idea_str_archive, which serializes to a pretty-printed JSON file at idea_fname upon completion. This file serves as the canonical record of all generated research ideas.

End-to-End Workflow

The complete ideation process follows this sequence:

  1. Load the workshop description markdown file
  2. Initialize the LLM client for the specified model (default: gpt-4o-2024-05-13)
  3. Iterate through generations:
    • Build context-aware prompts → LLM call → parse ACTION/ARGUMENTS
    • Execute SearchSemanticScholar queries when requested, feeding literature results back into the reflection loop
    • Capture FinalizeIdea payloads and validate JSON schema compliance
    • Perform num_reflections refinement rounds per proposal
  4. Serialize all finalized ideas to the output JSON file

Practical Examples

Command-Line Execution

Run ideation directly from the terminal to generate three research proposals with five reflection rounds each:

python -m ai_scientist.perform_ideation_temp_free \
    --model gpt-4o-2024-05-13 \
    --workshop-file ideas/i_cant_believe_its_not_better.md \
    --max-num-generations 3 \
    --num-reflections 5

This creates ideas/i_cant_believe_its_not_better.json containing up to three complete research proposals.

Programmatic Python Usage

Integrate ideation into custom pipelines by importing the core functions directly:

from ai_scientist.llm import create_client
from ai_scientist.perform_ideation_temp_free import generate_temp_free_idea

# Load workshop description

with open("ideas/i_cant_believe_its_not_better.md") as f:
    workshop = f.read()

# Initialize LLM client

client, model = create_client("gpt-4o-2024-05-13")

# Generate ideas with custom parameters

ideas = generate_temp_free_idea(
    idea_fname="ideas/output.json",
    client=client,
    model=model,
    workshop_description=workshop,
    max_num_generations=2,
    num_reflections=4,
)

print(f"Generated {len(ideas)} research proposals")

Expected Output Schema

The finalized JSON structure follows a strict schema required for downstream processing in AI Scientist v2:

{
  "idea": {
    "Name": "unique_identifier",
    "Title": "Descriptive Research Title",
    "Short Hypothesis": "Testable causal statement",
    "Related Work": "Literature context and gaps",
    "Abstract": "Comprehensive summary",
    "Experiments": "Numbered experimental plan",
    "Risk Factors and Limitations": "Critical analysis of constraints"
  }
}

Summary

  • AI Scientist v2 implements template-free ideation through ai_scientist/perform_ideation_temp_free.py, utilizing a multi-generational, reflection-based workflow.
  • The LLM client factory in ai_scientist/llm.py supports multiple providers (OpenAI, Anthropic, Ollama, Gemini) with integrated token tracking via ai_scientist/utils/token_tracker.py.
  • Tool augmentation enables literature-aware proposals through SearchSemanticScholar in ai_scientist/tools/semantic_scholar.py, which queries the Semantic Scholar Graph API and returns citation-ranked results.
  • The system accepts workshop descriptions as markdown files and outputs structured JSON proposals containing standardized research fields (Name, Title, Hypothesis, Abstract, Experiments, Risks).
  • Control parameters include max_num_generations (total proposals) and num_reflections (refinement depth per proposal).

Frequently Asked Questions

What is the difference between max_num_generations and num_reflections?

The max_num_generations parameter controls how many distinct research proposals the system creates in a single run, while num_reflections specifies how many iterative improvement cycles each individual proposal undergoes before finalization. For example, with max_num_generations=3 and num_reflections=5, the system generates three separate ideas, each refined through five rounds of literature search and self-critique.

Which workshop file format does AI Scientist v2 expect?

AI Scientist v2 expects workshop descriptions as markdown files (.md) containing domain context, research constraints, and preferred methodological approaches. The default location is ideas/i_cant_believe_its_not_better.md, though you can specify any path via the --workshop-file argument. The content should describe the research area sufficiently to guide the LLM toward relevant literature and feasible experimental designs.

How does the Semantic Scholar integration work during ideation?

The SearchSemanticScholar tool, implemented in ai_scientist/tools/semantic_scholar.py, activates when the LLM outputs ACTION: SearchSemanticScholar with query arguments. The tool calls search_for_papers against the Semantic Scholar Graph API, optionally using an S2_API_KEY for higher rate limits. Results sort by citation count and format into a human-readable list that feeds back into the next reflection round, enabling literature-aware hypothesis generation.

Can I use local models instead of commercial APIs for ideation?

Yes, the create_client function in ai_scientist/llm.py supports multiple providers including Ollama for local model hosting. Specify your local model identifier (e.g., --model ollama/llama3) to route requests to your local inference server instead of OpenAI or Anthropic endpoints, though you must ensure your local model supports the tool-calling format required for SearchSemanticScholar and FinalizeIdea actions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →