OpenDerisk Multi-Agent Architecture for Root Cause Analysis: How Five Specialized Agents Collaborate
OpenDerisk's multi-agent architecture orchestrates five specialized agents—SRE-Agent, Code-Agent, Data-Agent, Report-Agent, and Vis-Agent—through structured LLM tool calls to execute autonomous root cause analysis workflows.
OpenDerisk (also referred to as Open DeRisk) from the derisk-ai/openderisk repository implements a sophisticated multi-agent architecture for root cause analysis (RCA) that mimics human SRE workflows. By dividing responsibilities across five specialized agents, the system transforms raw observability data into actionable incident reports through a coordinated chain of analysis, execution, and visualization. This design leverages a master-agent orchestration pattern where the SRE-Agent delegates tasks to domain-specific sub-agents via deterministic tool calls defined in the core agent framework.
The Five Specialized Agents in OpenDerisk
OpenDerisk assigns distinct capabilities to each agent, with implementation details spread across the derisk-core and derisk-ext packages.
SRE-Agent (Master Orchestrator)
The SRE-Agent acts as the central brain of the RCA operation. Implemented as the ReActMasterAgent class in derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py, this agent manages the entire ReAct (Reasoning and Acting) loop. Its system prompt, defined in derisk-core/src/derisk/agent/expand/tool_agent/prompt_v0.py, enforces an SRE-only policy that prevents query drift outside the operations domain. The SRE-Agent plans tool-call graphs, enumerates available skills, and decides when to invoke sub-agents via the agent_start tool.
Code-Agent
The Code-Agent (formally CodeAssistantAgent) handles all analytical code generation and execution. Located in derisk-core/src/derisk/agent/expand/code_agent/agent.py, this agent generates Python, JavaScript, Bash, or SQL code to perform correlation and aggregation tasks. It executes code safely inside a sandbox environment using sandbox.shell.exec_command, storing results in the agent file system for downstream consumption.
Data-Agent (DataExpert)
The Data-Agent specializes in ingesting and structuring raw observability data. While conceptually described as the "DataExpert" in docs/OpenDerisk_v0.2.md (lines 42-44), the concrete implementation utilizes data-ingestion helpers in the derisk-ext package. This agent loads log files, metrics, Excel sheets, and trace data, transforming them into structured formats that the Code-Agent can consume for analysis.
Report-Agent
The Report-Agent assembles human-readable RCA documents from intermediate analysis artifacts. Defined in derisk-core/src/derisk/agent/expand/react_master_agent/report_generator.py (class starting at line 825), this agent consumes code outputs, logs, and statistics to generate Markdown, HTML, or JSON reports. It structures findings into sections including executive summaries, timelines, and root-cause hypotheses.
Vis-Agent
The Vis-Agent renders the entire reasoning chain as an interactive visual flow. Utilizing derisk-core/src/derisk/vis/vis_converter.py and concrete tag implementations under derisk-ext/src/derisk_ext/vis/gptvis/tags/, this agent builds GPT-Vis component trees (e.g., VisAgentPlans, VisAgentMessages, VisCode). It visualizes skill usage, tool calls, and sub-agent interactions, pushing the evidence chain to the front-end via the vis protocol.
The Root Cause Analysis Workflow: Step-by-Step Collaboration
The OpenDerisk multi-agent root cause analysis follows a deterministic pipeline orchestrated by the SRE-Agent's ReAct loop:
-
Query Intake and Planning – The SRE-Agent receives a user query (e.g., "Why did service X experience a latency spike at 10:42 AM?"). It evaluates
<available_skills>,<available_knowledges>, and<available_agents>to construct a tool-call execution plan. -
Data Ingestion – The SRE-Agent invokes the Data-Agent via the
agent_starttool with a sub-task to load specific resources (e.g.,svc_x.log). The Data-Agent uses system tools likeread_fileto ingest and structure the data, returning a structured blob to the SRE-Agent. -
Analytical Execution – The SRE-Agent delegates computational tasks to the Code-Agent, passing the structured data and analysis requirements. The Code-Agent generates code (e.g., Python for log correlation), executes it in a sandbox, and returns execution results (stdout, exit codes) saved to the agent file system.
-
Report Generation – With raw analysis results available, the SRE-Agent triggers the Report-Agent to synthesize findings. The Report-Agent formats the evidence into a comprehensive RCA document with sections for summaries and root-cause hypotheses.
-
Evidence Visualization – Finally, the SRE-Agent invokes the Vis-Agent to construct the evidence chain. The Vis-Agent creates a visual trace of all steps—data loading, code execution, and report generation—rendering it as interactive GPT-Vis components for the operator.
-
Unified Output – The composite result (structured report + visual evidence) returns to the user, completing the RCA loop.
All inter-agent communication occurs through LLM-structured tool calls (agent_start, knowledge_search, read_file, etc.) defined in derisk.agent.core.tools, making the workflow observable and auditable.
Code Implementation: How Agents Collaborate
Initiating the RCA Workflow
The following Python script demonstrates how to initiate a root cause analysis using the SRE-Agent as the entry point:
from derisk.agent import AgentContext, ConversableAgent
from derisk.agent.core.action.agent_action import AgentStart
from derisk.agent.util.llm.llm_client import AIWrapper
async def run_rca(query: str):
# Initialize the SRE-Agent (ReActMasterAgent)
ctx = AgentContext(conv_id="rca_demo")
sre_agent = ConversableAgent.from_name("ReActMasterV2", agent_context=ctx)
# The ReAct loop automatically invokes Data-Agent, Code-Agent, etc.
response = await AIWrapper(sre_agent).chat(query)
# Output contains final report and visualization payload
print("=== RCA Report ===")
print(response.message) # markdown report
print("\n=== Visualisation ===")
print(response.vis_payload) # JSON for front-end
The underlying orchestrator logic resides in derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py.
Delegating to the Code-Agent
When the SRE-Agent determines analytical code is required, it generates a structured tool call:
{
"action": "agent_start",
"action_input": {
"agent_name": "CodeAssistant",
"params": {
"task": "Analyze latency logs",
"data_key": "svc_x_log"
}
}
}
The agent_start tool schema is defined in derisk/agent/core/tools/agent_start_tool.py.
Sandbox Execution in Code-Agent
Inside derisk-core/src/derisk/agent/expand/code_agent/agent.py (lines 58-66), the Code-Agent executes generated code safely:
# Inside CodeAssistantAgent.execute_code
if language.lower() in ["python", "python3"]:
result = await sandbox.shell.exec_command(
command=f"python3 -c {repr(code)}",
timeout=timeout or self.execution_timeout,
work_dir=work_dir,
)
This ensures isolated execution of analysis scripts with configurable timeouts.
Report Generation
The Report-Agent constructs structured documents as shown in report_generator.py (lines 71-82):
report = Report(
metadata=ReportMetadata(
report_type=ReportType.DETAILED,
format=ReportFormat.MARKDOWN,
),
sections=[
ReportSection(title="Summary", content=summary_text),
ReportSection(title="Root Cause", content=root_cause_text),
ReportSection(title="Evidence", content=code_output),
],
)
Visualization Construction
The Vis-Agent creates interactive evidence trees using the factory pattern:
from derisk.vis import Vis
vis = Vis.of("code") # creates a VisCode component
vis.sync_display({
"language": "python",
"code": generated_code,
"log": execution_output,
})
The visualization utilities are implemented in derisk-core/src/derisk/vis/vis_converter.py.
Why This Design Enables Effective Root Cause Analysis
Domain Gating – The SRE-Agent's system prompt (prompt_v0.py) rejects non-SRE queries immediately, ensuring the architecture focuses strictly on operational incidents and prevents scope drift.
Skill-First Principle – Before invoking sub-agents, the SRE-Agent enumerates available skills (e.g., "Log-Correlation", "Metric-Anomaly-Detection"). These skills dictate which agent specialization is most appropriate for the task, ensuring domain expertise is applied correctly.
Iterative Refinement – If the Code-Agent's execution fails (syntax error, timeout), the SRE-Agent receives the observation through the ReAct loop and re-plans, potentially rewriting the code or switching languages (e.g., from Python to SQL).
Separation of Concerns – By isolating data handling, code execution, reporting, and visualization into distinct agents, each module becomes independently testable and evolvable. The Data-Agent manages I/O, the Code-Agent manages computation, and the Report-Agent manages presentation.
Unified Evidence Chain – The Vis-Agent stitches together all tool calls and intermediate outputs, providing operators with a traceable audit log that visually matches the final RCA report, crucial for post-incident reviews and compliance.
Summary
- OpenDerisk implements a five-agent architecture (SRE-Agent, Code-Agent, Data-Agent, Report-Agent, Vis-Agent) for autonomous root cause analysis.
- The SRE-Agent (
ReActMasterAgent) orchestrates the workflow fromderisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py, enforcing SRE-domain constraints. - Data-Agent ingests raw logs and metrics, while Code-Agent (
derisk-core/src/derisk/agent/expand/code_agent/agent.py) executes sandboxed analysis code. - Report-Agent (
report_generator.py) synthesizes findings into structured documents, and Vis-Agent renders the evidence chain via GPT-Vis components. - All coordination occurs through deterministic tool calls (
agent_start,read_file, etc.), creating observable and reproducible RCA workflows.
Frequently Asked Questions
How does the SRE-Agent decide which sub-agent to invoke?
The SRE-Agent uses a skill-first planning strategy defined in its system prompt (prompt_v0.py). When processing a query, it first enumerates <available_skills> and <available_agents>, then maps required capabilities to specific agent specializations. For data ingestion tasks, it invokes the Data-Agent; for computational analysis, it delegates to the Code-Agent via the agent_start tool with the appropriate agent_name parameter.
What happens if the Code-Agent's execution fails?
The ReAct loop enables iterative error recovery. If the Code-Agent's sandbox execution returns a non-zero exit code or timeout (handled in derisk-core/src/derisk/agent/expand/code_agent/agent.py), the SRE-Agent receives this observation as negative feedback. It then re-enters the reasoning phase to rewrite the code, adjust parameters (e.g., extend timeout), or switch to a different language interpreter (Bash instead of Python) before retrying.
How is data passed between agents in OpenDerisk?
Agents communicate via the agent file system and structured tool outputs. When the Data-Agent loads a log file, it returns a structured data blob to the SRE-Agent, which stores it with a unique data_key. The SRE-Agent then passes this key to the Code-Agent's agent_start invocation. The Code-Agent retrieves the data, processes it, and writes results back to the file system for the Report-Agent to consume, ensuring loose coupling between components.
Can the Vis-Agent display partial results during long-running analysis?
Yes. The Vis-Agent constructs the evidence chain incrementally using the GPT-Vis protocol implemented in derisk-core/src/derisk/vis/vis_converter.py. As the SRE-Agent completes tool calls (data loading, code execution), the Vis-Agent can push intermediate VisAgentMessages or VisCode components to the front-end before the final Report-Agent completes. This provides real-time visibility into the RCA progress, though the standard workflow described in the source code emphasizes single-turn completion cycles.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →