How to Define Custom Agents and Their Capabilities in MetaGPT: A Complete Guide
To define a custom agent in MetaGPT, subclass metagpt.roles.role.Role, implement one or more Action subclasses with async run() methods that encapsulate specific capabilities, and register them using set_actions() in the Role's __init__ method.
MetaGPT is a multi-agent framework that treats every autonomous AI entity as a Role. When you want to define custom agents and their capabilities in MetaGPT, you work within an architecture that cleanly separates an agent's orchestration logic from its executable skills. This guide demonstrates how to create specialized agents by extending the base Role class and attaching custom Action implementations that leverage large language models (LLMs).
Core Architecture Components
Understanding the relationship between Roles and Actions is essential before implementing custom agents.
Role: The Agent Foundation
The Role class, defined in metagpt/roles/role.py, serves as the base for every agent. It holds the agent's identity (name, profile), goal, constraints, and a list of available actions. The Role drives the think-act loop through _think and _act methods and manages a private message buffer via rc.msg_buffer. Each Role maintains a RoleContext (rc) that stores runtime state, including working memory, the current todo action, and the react mode.
Action: Encapsulating Capabilities
Actions are the atomic units of capability. To create a custom capability, subclass Action (from metagpt.actions.action_node.py) and implement an async run method. This method typically calls self._aask(prompt) to interact with the LLM and returns a processed result. Each Action should define a PROMPT_TEMPLATE class attribute that instructs the LLM on how to perform the specific task.
React Modes and Execution Flow
MetaGPT supports multiple strategies for action selection:
- ReAct (
react_mode = "react"): The default mode where the LLM decides which action to take next based on the current context. - BY_ORDER: Sequential execution where actions run in the order they were registered.
- PLAN_AND_ACT: The agent plans all steps first, then executes them.
Configure the mode using self._set_react_mode() inside your Role's __init__ method.
Step-by-Step: Building a Custom Agent
Follow this pattern to implement functional agents with specific capabilities.
Step 1: Create Action Subclasses
Define what your agent can do by implementing Actions. Each Action needs a prompt template and a parser for the LLM response.
import asyncio
import re
import subprocess
from metagpt.actions import Action
from metagpt.logs import logger
from metagpt.roles.role import Role, RoleReactMode
from metagpt.schema import Message
class SimpleWriteCode(Action):
"""Generate Python code from natural language instructions."""
PROMPT_TEMPLATE = """
Write a python function that can {instruction} and provide two runnable test cases.
Return ```python your_code_here ``` with NO other texts,
your code:
"""
name = "SimpleWriteCode"
async def run(self, instruction: str):
rsp = await self._aask(self.PROMPT_TEMPLATE.format(instruction=instruction))
return self.parse_code(rsp)
@staticmethod
def parse_code(rsp):
match = re.search(r"```python(.*)```", rsp, re.DOTALL)
return match.group(1) if match else rsp
class SimpleRunCode(Action):
"""Execute Python code and capture output."""
name = "SimpleRunCode"
async def run(self, code_text: str):
result = subprocess.run(
["python3", "-c", code_text], capture_output=True, text=True
)
logger.info(f"{result.stdout=}")
return result.stdout
Step 2: Subclass Role and Register Actions
Create the agent by subclassing Role, setting its identity, and attaching the Actions in __init__.
class RunnableCoder(Role):
"""An agent that writes code and then executes it."""
name = "Alice"
profile = "RunnableCoder"
def __init__(self, **kwargs):
super().__init__(**kwargs)
# Register capabilities
self.set_actions([SimpleWriteCode, SimpleRunCode])
# Enforce sequential execution: write first, then run
self._set_react_mode(react_mode=RoleReactMode.BY_ORDER.value)
async def _act(self) -> Message:
logger.info(f"{self._setting}: executing {self.rc.todo}")
# Retrieve the most recent message from memory
msg = self.get_memories(k=1)[0]
# Execute the current todo action
result = await self.rc.todo.run(msg.content)
# Wrap result in Message and store in memory
msg = Message(content=result, role=self.profile, cause_by=type(self.rc.todo))
self.rc.memory.add(msg)
return msg
Step 3: Run the Agent
Instantiate your Role and call run() from an async context. The framework automatically manages the think-act loop until no actions remain.
async def main():
role = RunnableCoder()
message = "write a function that calculates the product of a list and run it"
result = await role.run(message)
logger.info(result)
if __name__ == "__main__":
asyncio.run(main())
The complete working example is available in examples/build_customized_agent.py within the MetaGPT repository.
Advanced Pattern: Dynamic Agent Generation
MetaGPT includes a meta-programming capability where an agent can generate new agent code. The AgentCreator Role (found in examples/agent_creator.py) uses the CreateAgent Action to receive natural language descriptions and output Python class definitions.
class CreateAgent(Action):
# Prompts the LLM to generate agent code based on requirements
...
class AgentCreator(Role):
name = "Matrix"
profile = "AgentCreator"
def __init__(self, **kwargs):
super().__init__(**kwargs)
self.set_actions([CreateAgent])
async def _act(self) -> Message:
instruction = self.rc.memory.get()[-1].content
code = await CreateAgent().run(
example=self.agent_template,
instruction=instruction
)
return Message(content=code, role=self.profile, cause_by=self.rc.todo)
This pattern allows you to create agents that write other agents, enabling rapid prototyping of multi-agent systems from high-level descriptions.
Key Source Files Reference
| File | Purpose |
|---|---|
metagpt/roles/role.py |
Core Role class, think-act loop, and RoleContext implementation. |
metagpt/actions/action_node.py |
Base Action class and LLM interaction utilities (_aask). |
examples/build_customized_agent.py |
Complete example of a multi-action agent (write → run code). |
examples/agent_creator.py |
Meta-agent that synthesizes new agent classes from prompts. |
metagpt/configs/role_custom_config.py |
YAML configuration schema for external role definitions. |
metagpt/schema.py |
Message and memory structures for inter-agent communication. |
Summary
- Define capabilities by subclassing
Actionand implementing the asyncrun()method inmetagpt/actions/action_node.py. - Create agents by extending
Rolefrommetagpt/roles/role.py, settingnameandprofile, and registering Actions viaset_actions(). - Control execution flow using
_set_react_mode()to choose between ReAct, BY_ORDER, or PLAN_AND_ACT strategies. - Interact with LLMs inside Actions using
self._aask()with structuredPROMPT_TEMPLATEstrings. - Store state using the Role's memory system (
rc.memory) andMessageobjects defined inmetagpt/schema.py.
Frequently Asked Questions
What is the difference between a Role and an Action in MetaGPT?
A Role represents the agent itself—its identity, goals, constraints, and orchestration logic—while an Action represents a single, specific capability the agent can perform (like writing code or searching the web). The Role manages which Action to execute and when, whereas the Action contains the implementation details and LLM prompts needed to complete its specific task.
How do I control the order in which my agent executes actions?
Call self._set_react_mode(react_mode=RoleReactMode.BY_ORDER.value) inside your Role's __init__ method after set_actions(). This switches from the default ReAct mode (where the LLM chooses the next action) to sequential execution, guaranteeing that Actions run in the exact order they appear in the list passed to set_actions().
Can I create an agent that generates other agents automatically?
Yes. MetaGPT provides the AgentCreator pattern in examples/agent_creator.py. This Role uses a specialized CreateAgent Action that prompts the LLM to write Python code defining new Role and Action classes. You can feed it natural language requirements like "Create a data analyst agent that reads CSV files and generates matplotlib charts," and it will output runnable agent code.
Where does the agent store its memory and conversation history?
Each Role maintains a Memory object accessible via self.rc.memory (part of RoleContext). When an Action completes, its output should be wrapped in a Message object (from metagpt/schema.py) and added to memory using self.rc.memory.add(msg). Subsequent Actions can retrieve this history using methods like self.get_memories(k=1) to access recent context.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →