How to Use the MetaGPT Research Agent for Automated Web Scraping and Report Generation

The MetaGPT Research agent automates end-to-end research by chaining three specialized actions—CollectLinks, WebBrowseAndSummarize, and ConductResearch—to transform any topic into a structured markdown report with APA-style references.

The Research agent in the MetaGPT repository provides a production-ready framework for autonomous information gathering. By orchestrating web search, content extraction, and synthesis through a deterministic react-loop, this agent eliminates manual scraping while delivering publication-quality research reports.

How the MetaGPT Research Agent Works

The Research agent operates as a three-stage pipeline defined in metagpt/actions/research.py. Each stage is implemented as a discrete action class, allowing for modular execution and easy extension.

The CollectLinks class (line 80 in metagpt/actions/research.py) initiates the research workflow by generating searchable keywords and retrieving candidate URLs.

This action executes the SEARCH_TOPIC_PROMPT to derive up to two search keywords from the research topic. It then queries the configured search engine via the SearchEngine abstraction (metagpt/tools/search_engine.py), which supports backends including Serper, SERP-API, DuckDuckGo, and custom implementations.

The action ranks retrieved URLs using the COLLECT_AND_RANKURLS_PROMPT, returning only the top-N most relevant links to minimize noise in subsequent stages.

Step 2: Content Extraction via WebBrowseAndSummarize

The WebBrowseAndSummarize class (line 97 in metagpt/actions/research.py) handles the actual web scraping and content distillation.

This action utilizes the WebBrowserEngine (metagpt/tools/web_browser_engine.py) to fetch page content, supporting Chrome/Playwright for JavaScript-heavy sites or a lightweight curl-based fetcher for static pages. Large documents are automatically chunked using generate_prompt_chunk to respect context window limits, with each chunk summarized individually before merging into a coherent overview.

The action maintains query-context awareness, ensuring summaries focus on aspects relevant to the original research topic rather than generic page descriptions.

Step 3: Report Generation with ConductResearch

The ConductResearch class (line 5 in metagpt/actions/research.py) performs the final synthesis, transforming collected summaries into a structured research report.

This action aggregates all URL summaries into a single context and invokes the CONDUCT_RESEARCH_PROMPT to generate approximately 2,000 words of analysis. The output adheres to markdown formatting with APA-style references, proper citation of sources, and logical section organization.

The resulting report is persisted to RESEARCH_PATH (defined in metagpt/const.py, defaulting to ./research) with a filename sanitized to remove illegal filesystem characters.

Core Implementation Details

The Researcher role (metagpt/roles/researcher.py) orchestrates the three actions through a deterministic react-loop. It configures the execution order using:

self.set_actions([CollectLinks, WebBrowseAndSummarize, ConductResearch])
self._set_react_mode(RoleReactMode.BY_ORDER.value, len(self.actions))

This BY_ORDER mode ensures sequential execution: link collection must complete before browsing begins, and all summaries must be gathered before the final report generation.

System prompts are constructed dynamically via get_research_system_text in metagpt/actions/research.py, combining RESEARCH_TOPIC_SYSTEM (embedding the user's topic), RESEARCH_BASE_SYSTEM (framing the agent as a critical thinker), and LANG_PROMPT (for localization).

Practical Code Examples

Basic Research Report Generation

The minimal implementation requires only a topic string and the Researcher role:

import asyncio
from metagpt.roles.researcher import Researcher, RESEARCH_PATH

async def basic_report():
    topic = "dataiku vs. datarobot"
    researcher = Researcher()
    await researcher.run(topic)
    print(f"Report saved to: {RESEARCH_PATH / f'{topic}.md'}")

asyncio.run(basic_report())

This executes the full pipeline: search, scrape, summarize, and generate the markdown report saved to ./research/dataiku vs. datarobot.md.

Customizing Language and Concurrency

Control output language and parallel processing behavior via constructor parameters:

from metagpt.roles.researcher import Researcher

async def chinese_report():
    researcher = Researcher(language="zh-cn", enable_concurrency=False)
    await researcher.run("生成式 AI 在教育中的应用")
    # Report saved with Chinese content and sequential processing

Setting enable_concurrency=True (default) allows WebBrowseAndSummarize to process multiple URLs simultaneously, significantly reducing execution time for broad research topics.

Integrating Custom Search Engines

Replace the default search backend with internal document stores or proprietary APIs:

from metagpt.tools.search_engine import SearchEngine

async def custom_search(query: str, max_results: int, as_string: bool):
    # Example: query a private Elasticsearch cluster

    results = await my_elastic_client.search(query, size=max_results)
    return results if not as_string else str(results)

# Build a SearchEngine that uses the custom function

engine = SearchEngine.from_search_func(custom_search)

# Inject it into the CollectLinks action

from metagpt.actions.research import CollectLinks
collect_links_action = CollectLinks()
collect_links_action.search_engine = engine

This pattern enables the Research agent to operate on enterprise knowledge bases while maintaining the same summarization and reporting pipeline.

Direct Action Invocation (Advanced)

Execute individual pipeline stages manually for granular control or debugging:

from metagpt.actions.research import CollectLinks, WebBrowseAndSummarize, ConductResearch
from metagpt.roles.researcher import RESEARCH_PATH

topic = "edge computing security"

# Step 1 – collect & rank URLs

links = await CollectLinks().run(topic)

# Step 2 – summarize each URL

summaries = await WebBrowseAndSummarize().run(
    *sum([[url] for url in links.values()], []),
    query="What are the main security challenges?",
)

# Step 3 – generate final report

content = "\n---\n".join(f"url: {u}\nsummary: {s}" for u, s in summaries.items())
report = await ConductResearch().run(topic, content)

# Persist manually

(RESEARCH_PATH / f"{topic}.md").write_text(report)

This approach bypasses the Researcher role's react-loop, allowing custom preprocessing of URLs or post-processing of summaries before report generation.

Key Configuration Options

The Research agent behavior is controlled through several configuration vectors:

  • Search Engine Backend: Configure via config.yaml or environment variables to select between Serper, SERP-API, DuckDuckGo, or custom implementations in metagpt/tools/search_engine.py.
  • Browser Engine: Choose between Playwright (for JavaScript rendering), Selenium, or lightweight curl fetching in metagpt/tools/web_browser_engine.py.
  • System Prompts: Modify RESEARCH_TOPIC_SYSTEM, RESEARCH_BASE_SYSTEM, and LANG_PROMPT in metagpt/actions/research.py to adjust tone, citation style, or output language.
  • Output Path: Override RESEARCH_PATH in metagpt/const.py to change the default ./research directory for report persistence.

Summary

  • The MetaGPT Research agent automates academic-grade research through a three-stage pipeline implemented in metagpt/actions/research.py.
  • CollectLinks generates keywords, queries search engines via SearchEngine, and ranks URLs using LLM-based relevance scoring.
  • WebBrowseAndSummarize fetches content via WebBrowserEngine and produces chunked, context-aware summaries.
  • ConductResearch synthesizes summaries into structured markdown reports with APA citations, persisted to RESEARCH_PATH.
  • The Researcher role (metagpt/roles/researcher.py) orchestrates execution via a BY_ORDER react-loop, with configurable concurrency, language localization, and pluggable search backends.

Frequently Asked Questions

How does the Research agent handle large web pages that exceed LLM context limits?

The WebBrowseAndSummarize action in metagpt/actions/research.py automatically splits large documents into manageable chunks using the generate_prompt_chunk method. Each chunk is summarized individually against the research query, and the partial summaries are merged into a coherent final summary before being passed to the report generation stage.

Can I use the Research agent with internal corporate databases instead of public search engines?

Yes. The SearchEngine class in metagpt/tools/search_engine.py supports custom backends through the from_search_func factory method. You can implement an async function that queries Elasticsearch, SharePoint, or any internal API, then inject the resulting engine instance into the CollectLinks action before execution.

What is the difference between using the Researcher role versus invoking actions directly?

The Researcher role (metagpt/roles/researcher.py) provides a high-level orchestration layer that manages the react-loop, state persistence, and sequential execution of CollectLinks, WebBrowseAndSummarize, and ConductResearch. Invoking actions directly (as shown in the advanced example) bypasses this orchestration, giving you granular control over intermediate outputs, custom URL filtering, or modified summarization logic, but requires manual handling of data flow between stages.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →