How to Use the MetaGPT Research Agent for Automated Web Scraping and Report Generation
The MetaGPT Research agent automates end-to-end research by chaining three specialized actions—CollectLinks, WebBrowseAndSummarize, and ConductResearch—to transform any topic into a structured markdown report with APA-style references.
The Research agent in the MetaGPT repository provides a production-ready framework for autonomous information gathering. By orchestrating web search, content extraction, and synthesis through a deterministic react-loop, this agent eliminates manual scraping while delivering publication-quality research reports.
How the MetaGPT Research Agent Works
The Research agent operates as a three-stage pipeline defined in metagpt/actions/research.py. Each stage is implemented as a discrete action class, allowing for modular execution and easy extension.
Step 1: Link Collection with CollectLinks
The CollectLinks class (line 80 in metagpt/actions/research.py) initiates the research workflow by generating searchable keywords and retrieving candidate URLs.
This action executes the SEARCH_TOPIC_PROMPT to derive up to two search keywords from the research topic. It then queries the configured search engine via the SearchEngine abstraction (metagpt/tools/search_engine.py), which supports backends including Serper, SERP-API, DuckDuckGo, and custom implementations.
The action ranks retrieved URLs using the COLLECT_AND_RANKURLS_PROMPT, returning only the top-N most relevant links to minimize noise in subsequent stages.
Step 2: Content Extraction via WebBrowseAndSummarize
The WebBrowseAndSummarize class (line 97 in metagpt/actions/research.py) handles the actual web scraping and content distillation.
This action utilizes the WebBrowserEngine (metagpt/tools/web_browser_engine.py) to fetch page content, supporting Chrome/Playwright for JavaScript-heavy sites or a lightweight curl-based fetcher for static pages. Large documents are automatically chunked using generate_prompt_chunk to respect context window limits, with each chunk summarized individually before merging into a coherent overview.
The action maintains query-context awareness, ensuring summaries focus on aspects relevant to the original research topic rather than generic page descriptions.
Step 3: Report Generation with ConductResearch
The ConductResearch class (line 5 in metagpt/actions/research.py) performs the final synthesis, transforming collected summaries into a structured research report.
This action aggregates all URL summaries into a single context and invokes the CONDUCT_RESEARCH_PROMPT to generate approximately 2,000 words of analysis. The output adheres to markdown formatting with APA-style references, proper citation of sources, and logical section organization.
The resulting report is persisted to RESEARCH_PATH (defined in metagpt/const.py, defaulting to ./research) with a filename sanitized to remove illegal filesystem characters.
Core Implementation Details
The Researcher role (metagpt/roles/researcher.py) orchestrates the three actions through a deterministic react-loop. It configures the execution order using:
self.set_actions([CollectLinks, WebBrowseAndSummarize, ConductResearch])
self._set_react_mode(RoleReactMode.BY_ORDER.value, len(self.actions))
This BY_ORDER mode ensures sequential execution: link collection must complete before browsing begins, and all summaries must be gathered before the final report generation.
System prompts are constructed dynamically via get_research_system_text in metagpt/actions/research.py, combining RESEARCH_TOPIC_SYSTEM (embedding the user's topic), RESEARCH_BASE_SYSTEM (framing the agent as a critical thinker), and LANG_PROMPT (for localization).
Practical Code Examples
Basic Research Report Generation
The minimal implementation requires only a topic string and the Researcher role:
import asyncio
from metagpt.roles.researcher import Researcher, RESEARCH_PATH
async def basic_report():
topic = "dataiku vs. datarobot"
researcher = Researcher()
await researcher.run(topic)
print(f"Report saved to: {RESEARCH_PATH / f'{topic}.md'}")
asyncio.run(basic_report())
This executes the full pipeline: search, scrape, summarize, and generate the markdown report saved to ./research/dataiku vs. datarobot.md.
Customizing Language and Concurrency
Control output language and parallel processing behavior via constructor parameters:
from metagpt.roles.researcher import Researcher
async def chinese_report():
researcher = Researcher(language="zh-cn", enable_concurrency=False)
await researcher.run("生成式 AI 在教育中的应用")
# Report saved with Chinese content and sequential processing
Setting enable_concurrency=True (default) allows WebBrowseAndSummarize to process multiple URLs simultaneously, significantly reducing execution time for broad research topics.
Integrating Custom Search Engines
Replace the default search backend with internal document stores or proprietary APIs:
from metagpt.tools.search_engine import SearchEngine
async def custom_search(query: str, max_results: int, as_string: bool):
# Example: query a private Elasticsearch cluster
results = await my_elastic_client.search(query, size=max_results)
return results if not as_string else str(results)
# Build a SearchEngine that uses the custom function
engine = SearchEngine.from_search_func(custom_search)
# Inject it into the CollectLinks action
from metagpt.actions.research import CollectLinks
collect_links_action = CollectLinks()
collect_links_action.search_engine = engine
This pattern enables the Research agent to operate on enterprise knowledge bases while maintaining the same summarization and reporting pipeline.
Direct Action Invocation (Advanced)
Execute individual pipeline stages manually for granular control or debugging:
from metagpt.actions.research import CollectLinks, WebBrowseAndSummarize, ConductResearch
from metagpt.roles.researcher import RESEARCH_PATH
topic = "edge computing security"
# Step 1 – collect & rank URLs
links = await CollectLinks().run(topic)
# Step 2 – summarize each URL
summaries = await WebBrowseAndSummarize().run(
*sum([[url] for url in links.values()], []),
query="What are the main security challenges?",
)
# Step 3 – generate final report
content = "\n---\n".join(f"url: {u}\nsummary: {s}" for u, s in summaries.items())
report = await ConductResearch().run(topic, content)
# Persist manually
(RESEARCH_PATH / f"{topic}.md").write_text(report)
This approach bypasses the Researcher role's react-loop, allowing custom preprocessing of URLs or post-processing of summaries before report generation.
Key Configuration Options
The Research agent behavior is controlled through several configuration vectors:
- Search Engine Backend: Configure via
config.yamlor environment variables to select between Serper, SERP-API, DuckDuckGo, or custom implementations inmetagpt/tools/search_engine.py. - Browser Engine: Choose between Playwright (for JavaScript rendering), Selenium, or lightweight curl fetching in
metagpt/tools/web_browser_engine.py. - System Prompts: Modify
RESEARCH_TOPIC_SYSTEM,RESEARCH_BASE_SYSTEM, andLANG_PROMPTinmetagpt/actions/research.pyto adjust tone, citation style, or output language. - Output Path: Override
RESEARCH_PATHinmetagpt/const.pyto change the default./researchdirectory for report persistence.
Summary
- The MetaGPT Research agent automates academic-grade research through a three-stage pipeline implemented in
metagpt/actions/research.py. - CollectLinks generates keywords, queries search engines via
SearchEngine, and ranks URLs using LLM-based relevance scoring. - WebBrowseAndSummarize fetches content via
WebBrowserEngineand produces chunked, context-aware summaries. - ConductResearch synthesizes summaries into structured markdown reports with APA citations, persisted to
RESEARCH_PATH. - The Researcher role (
metagpt/roles/researcher.py) orchestrates execution via aBY_ORDERreact-loop, with configurable concurrency, language localization, and pluggable search backends.
Frequently Asked Questions
How does the Research agent handle large web pages that exceed LLM context limits?
The WebBrowseAndSummarize action in metagpt/actions/research.py automatically splits large documents into manageable chunks using the generate_prompt_chunk method. Each chunk is summarized individually against the research query, and the partial summaries are merged into a coherent final summary before being passed to the report generation stage.
Can I use the Research agent with internal corporate databases instead of public search engines?
Yes. The SearchEngine class in metagpt/tools/search_engine.py supports custom backends through the from_search_func factory method. You can implement an async function that queries Elasticsearch, SharePoint, or any internal API, then inject the resulting engine instance into the CollectLinks action before execution.
What is the difference between using the Researcher role versus invoking actions directly?
The Researcher role (metagpt/roles/researcher.py) provides a high-level orchestration layer that manages the react-loop, state persistence, and sequential execution of CollectLinks, WebBrowseAndSummarize, and ConductResearch. Invoking actions directly (as shown in the advanced example) bypasses this orchestration, giving you granular control over intermediate outputs, custom URL filtering, or modified summarization logic, but requires manual handling of data flow between stages.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →