Leveraging the Data Interpreter for Advanced Data Analysis Tasks in MetaGPT
The MetaGPT Data Interpreter autonomously generates, executes, and refines Python analysis code through a REACT-style loop with integrated tool recommendation and reflection capabilities.
MetaGPT's Data Interpreter transforms natural-language data analysis requests into self-healing Python workflows. By combining a planning engine with notebook-style code execution, this role automates complex pipelines—from preprocessing and feature engineering to model training—without requiring manual boilerplate. This guide examines the interpreter's architecture, implementation patterns, and extension points based on the FoundationAgents/MetaGPT source code.
What Is the MetaGPT Data Interpreter?
The Data Interpreter is a specialized Role subclass that orchestrates autonomous data analysis through an iterative think-act-execute cycle. Unlike static code generators, it maintains working memory across iterations, validates intermediate results, and can retry failed executions with reflected improvements.
Key capabilities include:
- Autonomous planning with task decomposition
- Tool recommendation via BM25 retrieval for context-aware tool selection
- Interactive code execution in isolated notebook kernels
- Data validation with automatic sanity checks on transformed datasets
- Reflection-based debugging for self-correcting failed attempts
Core Architecture and Components
The interpreter's functionality is distributed across specialized actions and roles defined in the metagpt/roles/di/ and metagpt/actions/di/ directories.
DataInterpreter Role
The central orchestrator resides in metagpt/roles/di/data_interpreter.py. This class extends the base Role to implement a REACT-style loop with configurable modes:
class DataInterpreter(Role):
name: str = "David"
profile: str = "DataInterpreter"
auto_run: bool = True
use_plan: bool = True
react_mode: Literal["plan_and_act", "react"] = "plan_and_act"
max_react_loop: int = 10
Source: lines 36-46 of [metagpt/roles/di/data_interpreter.py](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/roles/di/data_interpreter.py#L36-L46)
WriteAnalysisCode Action
Located in metagpt/actions/di/write_analysis_code.py, this action handles LLM prompt construction and code generation. It supports an optional reflection step for debugging failed executions through the _debug_with_reflection method (lines 23-34).
The action receives:
- User requirements
- Current plan status
- Tool hints from the recommender
- Previous execution history from working memory
CheckData and ExecuteNbCode Actions
CheckData performs post-execution validation on datasets after preprocessing or model training tasks. It injects dataset summaries (DATA_INFO) back into memory to inform subsequent planning iterations.
ExecuteNbCode manages the Jupyter-like kernel environment, executing generated Python code in isolation and returning output strings with success flags.
The REACT Loop: How the Data Interpreter Works
The interpreter operates through a structured cognitive cycle that mimics human data analysis workflows.
Planning and Tool Recommendation
When initialized with tools, the interpreter creates a BM25ToolRecommender (lines 55-58 of data_interpreter.py) that suggests relevant tools based on the current task context. This enables dynamic tool selection without hardcoded logic.
Think and Act Phases
In react mode, the _think method (lines 65-84) queries the LLM with REACT_THINK_PROMPT to generate structured thoughts and determine if additional actions are required. The _act phase then invokes _write_and_exec_code (lines 107-148), which:
- Generates Python via
WriteAnalysisCode.run - Executes via
ExecuteNbCode.run - Stores results in working memory
- Returns structured status messages
Data Validation and Reflection
Before each iteration, _check_data (lines 70-90) validates datasets when the current task involves preprocessing, feature engineering, or model training. If execution fails and use_reflection is enabled, _debug_with_reflection generates improved code using the previous implementation and error context.
Practical Implementation: Using the DataAnalyst Role
For end-users, MetaGPT provides the DataAnalyst concrete role in metagpt/roles/di/data_analyst.py. This wrapper pre-configures the interpreter with the write_and_exec_code tool, enabling immediate use without manual setup.
import fire
from metagpt.logs import logger
from metagpt.roles.di.data_analyst import DataAnalyst
async def main():
# ① Create the role
analyst = DataAnalyst()
# ② Define the high-level goal for the planner
analyst.planner.plan.goal = "construct a two-dimensional array"
# ③ Append a concrete DATA_ANALYSIS task
analyst.planner.plan.append_task(
task_id="1",
dependent_task_ids=[],
instruction="construct a two-dimensional array",
assignee="David",
task_type="DATA_ANALYSIS",
)
# ④ Let the interpreter write and execute the code
result = await analyst.write_and_exec_code("construct a two-dimensional array")
# ⑤ Log the final structured output
logger.info(result)
if __name__ == "__main__":
fire.Fire(main)
Source: [examples/di/data_analyst_write_code.py](https://github.com/FoundationAgents/MetaGPT/blob/main/examples/di/data_analyst_write_code.py)
Execution flow:
write_and_exec_codeconstructs a plan status string and retrieves tool hintsWriteAnalysisCode.runprompts the LLM to generate Python (e.g.,import numpy as np; arr = np.arange(9).reshape(3,3))ExecuteNbCode.runevaluates the code in an isolated kernel and returns the matrix output- The interpreter formats the result with
CODE_STATUS(e.g.,"Success ✅")
Key Source Files and Their Responsibilities
| File | Purpose |
|---|---|
metagpt/roles/di/data_interpreter.py |
Core REACT loop, tool recommendation, memory handling, and error retry logic |
metagpt/actions/di/write_analysis_code.py |
LLM prompt templates, code generation, and optional reflection debugging |
metagpt/roles/di/data_analyst.py |
Concrete user-facing role exposing write_and_exec_code |
metagpt/prompts/di/write_analysis_code.py |
System messages (INTERPRETER_SYSTEM_MSG, STRUCTUAL_PROMPT) guiding LLM output |
metagpt/actions/di/execute_nb_code.py |
Jupyter-like kernel execution for generated Python |
examples/di/data_analyst_write_code.py |
End-to-end demonstration of the DataAnalyst workflow |
Extending the Data Interpreter
The modular architecture enables customization for specialized analysis workflows.
Adding Custom Tools
Populate the custom_tools list when instantiating DataAnalyst (e.g., "SQL", "PandasProfile"). The built-in BM25ToolRecommender automatically surfaces relevant tool-usage guidance based on the current task context, enabling dynamic tool selection without hardcoded logic.
Enabling Reflection and Debugging
Set use_reflection = True (enabled by default) to activate the _debug_with_reflection method in WriteAnalysisCode. When execution fails, this triggers a debugging LLM prompt that receives the previous implementation and error context to generate improved code automatically.
Summary
- The Data Interpreter in MetaGPT provides an autonomous REACT-style loop for advanced data analysis, combining planning, code generation, and execution in
metagpt/roles/di/data_interpreter.py. - WriteAnalysisCode and ExecuteNbCode handle LLM prompt construction and isolated notebook execution, respectively, enabling safe, iterative code refinement.
- The DataAnalyst role offers a production-ready wrapper that exposes
write_and_exec_codefor immediate use without manual configuration. - Built-in reflection and BM25 tool recommendation support self-healing workflows and dynamic tool selection for complex preprocessing, feature engineering, and modeling tasks.
Frequently Asked Questions
How does the Data Interpreter handle execution errors?
When code execution fails, the interpreter checks the use_reflection flag (enabled by default). If active, it invokes _debug_with_reflection in WriteAnalysisCode, which sends the failed implementation and error traceback to the LLM with a debugging prompt. The LLM returns a corrected code snippet, and the interpreter retries execution up to the configured maximum loop count.
What is the difference between DataInterpreter and DataAnalyst?
DataInterpreter is the abstract base class in metagpt/roles/di/data_interpreter.py that implements the core REACT loop, memory management, and tool recommendation. DataAnalyst in metagpt/roles/di/data_analyst.py is a concrete subclass that pre-configures the interpreter with the write_and_exec_code tool and default settings, providing a ready-to-use interface for end-users without requiring manual action registration.
Can I use custom Python libraries in the Data Interpreter?
Yes. The ExecuteNbCode action runs generated code in a Jupyter-like kernel environment, which supports any library installed in the Python environment. When extending the interpreter, you can guide the LLM to import specific packages by modifying the system prompts in metagpt/prompts/di/write_analysis_code.py or by providing tool hints through the custom_tools parameter to influence code generation patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →