# How to Use the MetaGPT Research Agent for Automated Web Scraping and Report Generation

> Learn how to use the MetaGPT Research agent for automated web scraping and report generation. Transform topics into structured markdown reports with APA references.

- Repository: [FoundationAgents/MetaGPT](https://github.com/FoundationAgents/MetaGPT)
- Tags: how-to-guide
- Published: 2026-03-04

---

**The MetaGPT Research agent automates end-to-end research by chaining three specialized actions—CollectLinks, WebBrowseAndSummarize, and ConductResearch—to transform any topic into a structured markdown report with APA-style references.**

The **Research agent** in the [MetaGPT](https://github.com/FoundationAgents/MetaGPT) repository provides a production-ready framework for autonomous information gathering. By orchestrating web search, content extraction, and synthesis through a deterministic react-loop, this agent eliminates manual scraping while delivering publication-quality research reports.

## How the MetaGPT Research Agent Works

The Research agent operates as a **three-stage pipeline** defined in [`metagpt/actions/research.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/actions/research.py). Each stage is implemented as a discrete action class, allowing for modular execution and easy extension.

### Step 1: Link Collection with CollectLinks

The `CollectLinks` class (line 80 in [`metagpt/actions/research.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/actions/research.py)) initiates the research workflow by generating searchable keywords and retrieving candidate URLs.

This action executes the `SEARCH_TOPIC_PROMPT` to derive up to two search keywords from the research topic. It then queries the configured search engine via the `SearchEngine` abstraction ([`metagpt/tools/search_engine.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/tools/search_engine.py)), which supports backends including Serper, SERP-API, DuckDuckGo, and custom implementations.

The action ranks retrieved URLs using the `COLLECT_AND_RANKURLS_PROMPT`, returning only the top-N most relevant links to minimize noise in subsequent stages.

### Step 2: Content Extraction via WebBrowseAndSummarize

The `WebBrowseAndSummarize` class (line 97 in [`metagpt/actions/research.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/actions/research.py)) handles the actual web scraping and content distillation.

This action utilizes the `WebBrowserEngine` ([`metagpt/tools/web_browser_engine.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/tools/web_browser_engine.py)) to fetch page content, supporting Chrome/Playwright for JavaScript-heavy sites or a lightweight curl-based fetcher for static pages. Large documents are automatically chunked using `generate_prompt_chunk` to respect context window limits, with each chunk summarized individually before merging into a coherent overview.

The action maintains query-context awareness, ensuring summaries focus on aspects relevant to the original research topic rather than generic page descriptions.

### Step 3: Report Generation with ConductResearch

The `ConductResearch` class (line 5 in [`metagpt/actions/research.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/actions/research.py)) performs the final synthesis, transforming collected summaries into a structured research report.

This action aggregates all URL summaries into a single context and invokes the `CONDUCT_RESEARCH_PROMPT` to generate approximately 2,000 words of analysis. The output adheres to markdown formatting with APA-style references, proper citation of sources, and logical section organization.

The resulting report is persisted to `RESEARCH_PATH` (defined in [`metagpt/const.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/const.py), defaulting to `./research`) with a filename sanitized to remove illegal filesystem characters.

## Core Implementation Details

The **Researcher** role ([`metagpt/roles/researcher.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/roles/researcher.py)) orchestrates the three actions through a deterministic react-loop. It configures the execution order using:

```python
self.set_actions([CollectLinks, WebBrowseAndSummarize, ConductResearch])
self._set_react_mode(RoleReactMode.BY_ORDER.value, len(self.actions))

```

This `BY_ORDER` mode ensures sequential execution: link collection must complete before browsing begins, and all summaries must be gathered before the final report generation.

System prompts are constructed dynamically via `get_research_system_text` in [`metagpt/actions/research.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/actions/research.py), combining `RESEARCH_TOPIC_SYSTEM` (embedding the user's topic), `RESEARCH_BASE_SYSTEM` (framing the agent as a critical thinker), and `LANG_PROMPT` (for localization).

## Practical Code Examples

### Basic Research Report Generation

The minimal implementation requires only a topic string and the `Researcher` role:

```python
import asyncio
from metagpt.roles.researcher import Researcher, RESEARCH_PATH

async def basic_report():
    topic = "dataiku vs. datarobot"
    researcher = Researcher()
    await researcher.run(topic)
    print(f"Report saved to: {RESEARCH_PATH / f'{topic}.md'}")

asyncio.run(basic_report())

```

This executes the full pipeline: search, scrape, summarize, and generate the markdown report saved to `./research/dataiku vs. datarobot.md`.

### Customizing Language and Concurrency

Control output language and parallel processing behavior via constructor parameters:

```python
from metagpt.roles.researcher import Researcher

async def chinese_report():
    researcher = Researcher(language="zh-cn", enable_concurrency=False)
    await researcher.run("生成式 AI 在教育中的应用")
    # Report saved with Chinese content and sequential processing

```

Setting `enable_concurrency=True` (default) allows `WebBrowseAndSummarize` to process multiple URLs simultaneously, significantly reducing execution time for broad research topics.

### Integrating Custom Search Engines

Replace the default search backend with internal document stores or proprietary APIs:

```python
from metagpt.tools.search_engine import SearchEngine

async def custom_search(query: str, max_results: int, as_string: bool):
    # Example: query a private Elasticsearch cluster

    results = await my_elastic_client.search(query, size=max_results)
    return results if not as_string else str(results)

# Build a SearchEngine that uses the custom function

engine = SearchEngine.from_search_func(custom_search)

# Inject it into the CollectLinks action

from metagpt.actions.research import CollectLinks
collect_links_action = CollectLinks()
collect_links_action.search_engine = engine

```

This pattern enables the Research agent to operate on enterprise knowledge bases while maintaining the same summarization and reporting pipeline.

### Direct Action Invocation (Advanced)

Execute individual pipeline stages manually for granular control or debugging:

```python
from metagpt.actions.research import CollectLinks, WebBrowseAndSummarize, ConductResearch
from metagpt.roles.researcher import RESEARCH_PATH

topic = "edge computing security"

# Step 1 – collect & rank URLs

links = await CollectLinks().run(topic)

# Step 2 – summarize each URL

summaries = await WebBrowseAndSummarize().run(
    *sum([[url] for url in links.values()], []),
    query="What are the main security challenges?",
)

# Step 3 – generate final report

content = "\n---\n".join(f"url: {u}\nsummary: {s}" for u, s in summaries.items())
report = await ConductResearch().run(topic, content)

# Persist manually

(RESEARCH_PATH / f"{topic}.md").write_text(report)

```

This approach bypasses the `Researcher` role's react-loop, allowing custom preprocessing of URLs or post-processing of summaries before report generation.

## Key Configuration Options

The Research agent behavior is controlled through several configuration vectors:

- **Search Engine Backend**: Configure via [`config.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config.yaml) or environment variables to select between Serper, SERP-API, DuckDuckGo, or custom implementations in [`metagpt/tools/search_engine.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/tools/search_engine.py).
- **Browser Engine**: Choose between Playwright (for JavaScript rendering), Selenium, or lightweight curl fetching in [`metagpt/tools/web_browser_engine.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/tools/web_browser_engine.py).
- **System Prompts**: Modify `RESEARCH_TOPIC_SYSTEM`, `RESEARCH_BASE_SYSTEM`, and `LANG_PROMPT` in [`metagpt/actions/research.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/actions/research.py) to adjust tone, citation style, or output language.
- **Output Path**: Override `RESEARCH_PATH` in [`metagpt/const.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/const.py) to change the default `./research` directory for report persistence.

## Summary

- The **MetaGPT Research agent** automates academic-grade research through a three-stage pipeline implemented in [`metagpt/actions/research.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/actions/research.py).
- **CollectLinks** generates keywords, queries search engines via `SearchEngine`, and ranks URLs using LLM-based relevance scoring.
- **WebBrowseAndSummarize** fetches content via `WebBrowserEngine` and produces chunked, context-aware summaries.
- **ConductResearch** synthesizes summaries into structured markdown reports with APA citations, persisted to `RESEARCH_PATH`.
- The **Researcher** role ([`metagpt/roles/researcher.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/roles/researcher.py)) orchestrates execution via a `BY_ORDER` react-loop, with configurable concurrency, language localization, and pluggable search backends.

## Frequently Asked Questions

### How does the Research agent handle large web pages that exceed LLM context limits?

The `WebBrowseAndSummarize` action in [`metagpt/actions/research.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/actions/research.py) automatically splits large documents into manageable chunks using the `generate_prompt_chunk` method. Each chunk is summarized individually against the research query, and the partial summaries are merged into a coherent final summary before being passed to the report generation stage.

### Can I use the Research agent with internal corporate databases instead of public search engines?

Yes. The `SearchEngine` class in [`metagpt/tools/search_engine.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/tools/search_engine.py) supports custom backends through the `from_search_func` factory method. You can implement an async function that queries Elasticsearch, SharePoint, or any internal API, then inject the resulting engine instance into the `CollectLinks` action before execution.

### What is the difference between using the Researcher role versus invoking actions directly?

The `Researcher` role ([`metagpt/roles/researcher.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/roles/researcher.py)) provides a high-level orchestration layer that manages the react-loop, state persistence, and sequential execution of `CollectLinks`, `WebBrowseAndSummarize`, and `ConductResearch`. Invoking actions directly (as shown in the advanced example) bypasses this orchestration, giving you granular control over intermediate outputs, custom URL filtering, or modified summarization logic, but requires manual handling of data flow between stages.