# OpenDerisk Multi-Agent Architecture for Root Cause Analysis: How Five Specialized Agents Collaborate

> Discover how OpenDerisk's multi-agent architecture uses 5 specialized agents for autonomous root cause analysis. Learn about SRE-Agent, Code-Agent, ReportAgent, Vis-Agent, and Data-Agent collaboration.

- Repository: [derisk-ai/openderisk](https://github.com/derisk-ai/openderisk)
- Tags: architecture
- Published: 2026-02-28

---

**OpenDerisk's multi-agent architecture orchestrates five specialized agents—SRE-Agent, Code-Agent, Data-Agent, Report-Agent, and Vis-Agent—through structured LLM tool calls to execute autonomous root cause analysis workflows.**

OpenDerisk (also referred to as **Open DeRisk**) from the `derisk-ai/openderisk` repository implements a sophisticated **multi-agent architecture for root cause analysis (RCA)** that mimics human SRE workflows. By dividing responsibilities across five specialized agents, the system transforms raw observability data into actionable incident reports through a coordinated chain of analysis, execution, and visualization. This design leverages a master-agent orchestration pattern where the **SRE-Agent** delegates tasks to domain-specific sub-agents via deterministic tool calls defined in the core agent framework.

## The Five Specialized Agents in OpenDerisk

OpenDerisk assigns distinct capabilities to each agent, with implementation details spread across the `derisk-core` and `derisk-ext` packages.

### SRE-Agent (Master Orchestrator)

The **SRE-Agent** acts as the central brain of the RCA operation. Implemented as the `ReActMasterAgent` class in [`derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py), this agent manages the entire ReAct (Reasoning and Acting) loop. Its system prompt, defined in [`derisk-core/src/derisk/agent/expand/tool_agent/prompt_v0.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/agent/expand/tool_agent/prompt_v0.py), enforces an **SRE-only policy** that prevents query drift outside the operations domain. The SRE-Agent plans tool-call graphs, enumerates available skills, and decides when to invoke sub-agents via the `agent_start` tool.

### Code-Agent

The **Code-Agent** (formally `CodeAssistantAgent`) handles all analytical code generation and execution. Located in [`derisk-core/src/derisk/agent/expand/code_agent/agent.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/agent/expand/code_agent/agent.py), this agent generates Python, JavaScript, Bash, or SQL code to perform correlation and aggregation tasks. It executes code safely inside a sandbox environment using `sandbox.shell.exec_command`, storing results in the agent file system for downstream consumption.

### Data-Agent (DataExpert)

The **Data-Agent** specializes in ingesting and structuring raw observability data. While conceptually described as the "DataExpert" in [`docs/OpenDerisk_v0.2.md`](https://github.com/derisk-ai/openderisk/blob/main/docs/OpenDerisk_v0.2.md) (lines 42-44), the concrete implementation utilizes data-ingestion helpers in the `derisk-ext` package. This agent loads log files, metrics, Excel sheets, and trace data, transforming them into structured formats that the Code-Agent can consume for analysis.

### Report-Agent

The **Report-Agent** assembles human-readable RCA documents from intermediate analysis artifacts. Defined in [`derisk-core/src/derisk/agent/expand/react_master_agent/report_generator.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/agent/expand/react_master_agent/report_generator.py) (class starting at line 825), this agent consumes code outputs, logs, and statistics to generate Markdown, HTML, or JSON reports. It structures findings into sections including executive summaries, timelines, and root-cause hypotheses.

### Vis-Agent

The **Vis-Agent** renders the entire reasoning chain as an interactive visual flow. Utilizing [`derisk-core/src/derisk/vis/vis_converter.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/vis/vis_converter.py) and concrete tag implementations under `derisk-ext/src/derisk_ext/vis/gptvis/tags/`, this agent builds **GPT-Vis component trees** (e.g., `VisAgentPlans`, `VisAgentMessages`, `VisCode`). It visualizes skill usage, tool calls, and sub-agent interactions, pushing the evidence chain to the front-end via the `vis` protocol.

## The Root Cause Analysis Workflow: Step-by-Step Collaboration

The OpenDerisk multi-agent root cause analysis follows a deterministic pipeline orchestrated by the SRE-Agent's ReAct loop:

1. **Query Intake and Planning** – The SRE-Agent receives a user query (e.g., "Why did service X experience a latency spike at 10:42 AM?"). It evaluates `<available_skills>`, `<available_knowledges>`, and `<available_agents>` to construct a tool-call execution plan.

2. **Data Ingestion** – The SRE-Agent invokes the **Data-Agent** via the `agent_start` tool with a sub-task to load specific resources (e.g., `svc_x.log`). The Data-Agent uses system tools like `read_file` to ingest and structure the data, returning a structured blob to the SRE-Agent.

3. **Analytical Execution** – The SRE-Agent delegates computational tasks to the **Code-Agent**, passing the structured data and analysis requirements. The Code-Agent generates code (e.g., Python for log correlation), executes it in a sandbox, and returns execution results (stdout, exit codes) saved to the agent file system.

4. **Report Generation** – With raw analysis results available, the SRE-Agent triggers the **Report-Agent** to synthesize findings. The Report-Agent formats the evidence into a comprehensive RCA document with sections for summaries and root-cause hypotheses.

5. **Evidence Visualization** – Finally, the SRE-Agent invokes the **Vis-Agent** to construct the evidence chain. The Vis-Agent creates a visual trace of all steps—data loading, code execution, and report generation—rendering it as interactive GPT-Vis components for the operator.

6. **Unified Output** – The composite result (structured report + visual evidence) returns to the user, completing the RCA loop.

All inter-agent communication occurs through **LLM-structured tool calls** (`agent_start`, `knowledge_search`, `read_file`, etc.) defined in `derisk.agent.core.tools`, making the workflow observable and auditable.

## Code Implementation: How Agents Collaborate

### Initiating the RCA Workflow

The following Python script demonstrates how to initiate a root cause analysis using the SRE-Agent as the entry point:

```python
from derisk.agent import AgentContext, ConversableAgent
from derisk.agent.core.action.agent_action import AgentStart
from derisk.agent.util.llm.llm_client import AIWrapper

async def run_rca(query: str):
    # Initialize the SRE-Agent (ReActMasterAgent)

    ctx = AgentContext(conv_id="rca_demo")
    sre_agent = ConversableAgent.from_name("ReActMasterV2", agent_context=ctx)

    # The ReAct loop automatically invokes Data-Agent, Code-Agent, etc.

    response = await AIWrapper(sre_agent).chat(query)

    # Output contains final report and visualization payload

    print("=== RCA Report ===")
    print(response.message)          # markdown report

    print("\n=== Visualisation ===")
    print(response.vis_payload)      # JSON for front-end

```

The underlying orchestrator logic resides in [`derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py).

### Delegating to the Code-Agent

When the SRE-Agent determines analytical code is required, it generates a structured tool call:

```json
{
  "action": "agent_start",
  "action_input": {
    "agent_name": "CodeAssistant",
    "params": {
      "task": "Analyze latency logs",
      "data_key": "svc_x_log"
    }
  }
}

```

The `agent_start` tool schema is defined in [`derisk/agent/core/tools/agent_start_tool.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk/agent/core/tools/agent_start_tool.py).

### Sandbox Execution in Code-Agent

Inside [`derisk-core/src/derisk/agent/expand/code_agent/agent.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/agent/expand/code_agent/agent.py) (lines 58-66), the Code-Agent executes generated code safely:

```python

# Inside CodeAssistantAgent.execute_code

if language.lower() in ["python", "python3"]:
    result = await sandbox.shell.exec_command(
        command=f"python3 -c {repr(code)}",
        timeout=timeout or self.execution_timeout,
        work_dir=work_dir,
    )

```

This ensures isolated execution of analysis scripts with configurable timeouts.

### Report Generation

The Report-Agent constructs structured documents as shown in [`report_generator.py`](https://github.com/derisk-ai/openderisk/blob/main/report_generator.py) (lines 71-82):

```python
report = Report(
    metadata=ReportMetadata(
        report_type=ReportType.DETAILED,
        format=ReportFormat.MARKDOWN,
    ),
    sections=[
        ReportSection(title="Summary", content=summary_text),
        ReportSection(title="Root Cause", content=root_cause_text),
        ReportSection(title="Evidence", content=code_output),
    ],
)

```

### Visualization Construction

The Vis-Agent creates interactive evidence trees using the factory pattern:

```python
from derisk.vis import Vis

vis = Vis.of("code")          # creates a VisCode component

vis.sync_display({
    "language": "python",
    "code": generated_code,
    "log": execution_output,
})

```

The visualization utilities are implemented in [`derisk-core/src/derisk/vis/vis_converter.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/vis/vis_converter.py).

## Why This Design Enables Effective Root Cause Analysis

**Domain Gating** – The SRE-Agent's system prompt ([`prompt_v0.py`](https://github.com/derisk-ai/openderisk/blob/main/prompt_v0.py)) rejects non-SRE queries immediately, ensuring the architecture focuses strictly on operational incidents and prevents scope drift.

**Skill-First Principle** – Before invoking sub-agents, the SRE-Agent enumerates available skills (e.g., "Log-Correlation", "Metric-Anomaly-Detection"). These skills dictate which agent specialization is most appropriate for the task, ensuring domain expertise is applied correctly.

**Iterative Refinement** – If the Code-Agent's execution fails (syntax error, timeout), the SRE-Agent receives the observation through the ReAct loop and re-plans, potentially rewriting the code or switching languages (e.g., from Python to SQL).

**Separation of Concerns** – By isolating data handling, code execution, reporting, and visualization into distinct agents, each module becomes independently testable and evolvable. The Data-Agent manages I/O, the Code-Agent manages computation, and the Report-Agent manages presentation.

**Unified Evidence Chain** – The Vis-Agent stitches together all tool calls and intermediate outputs, providing operators with a traceable audit log that visually matches the final RCA report, crucial for post-incident reviews and compliance.

## Summary

- OpenDerisk implements a **five-agent architecture** (SRE-Agent, Code-Agent, Data-Agent, Report-Agent, Vis-Agent) for autonomous root cause analysis.
- The **SRE-Agent** (`ReActMasterAgent`) orchestrates the workflow from [`derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py), enforcing SRE-domain constraints.
- **Data-Agent** ingests raw logs and metrics, while **Code-Agent** ([`derisk-core/src/derisk/agent/expand/code_agent/agent.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/agent/expand/code_agent/agent.py)) executes sandboxed analysis code.
- **Report-Agent** ([`report_generator.py`](https://github.com/derisk-ai/openderisk/blob/main/report_generator.py)) synthesizes findings into structured documents, and **Vis-Agent** renders the evidence chain via GPT-Vis components.
- All coordination occurs through **deterministic tool calls** (`agent_start`, `read_file`, etc.), creating observable and reproducible RCA workflows.

## Frequently Asked Questions

### How does the SRE-Agent decide which sub-agent to invoke?

The SRE-Agent uses a **skill-first planning strategy** defined in its system prompt ([`prompt_v0.py`](https://github.com/derisk-ai/openderisk/blob/main/prompt_v0.py)). When processing a query, it first enumerates `<available_skills>` and `<available_agents>`, then maps required capabilities to specific agent specializations. For data ingestion tasks, it invokes the Data-Agent; for computational analysis, it delegates to the Code-Agent via the `agent_start` tool with the appropriate `agent_name` parameter.

### What happens if the Code-Agent's execution fails?

The ReAct loop enables **iterative error recovery**. If the Code-Agent's sandbox execution returns a non-zero exit code or timeout (handled in [`derisk-core/src/derisk/agent/expand/code_agent/agent.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/agent/expand/code_agent/agent.py)), the SRE-Agent receives this observation as negative feedback. It then re-enters the reasoning phase to rewrite the code, adjust parameters (e.g., extend timeout), or switch to a different language interpreter (Bash instead of Python) before retrying.

### How is data passed between agents in OpenDerisk?

Agents communicate via the **agent file system** and structured tool outputs. When the Data-Agent loads a log file, it returns a structured data blob to the SRE-Agent, which stores it with a unique `data_key`. The SRE-Agent then passes this key to the Code-Agent's `agent_start` invocation. The Code-Agent retrieves the data, processes it, and writes results back to the file system for the Report-Agent to consume, ensuring loose coupling between components.

### Can the Vis-Agent display partial results during long-running analysis?

Yes. The Vis-Agent constructs the evidence chain incrementally using the GPT-Vis protocol implemented in [`derisk-core/src/derisk/vis/vis_converter.py`](https://github.com/derisk-ai/openderisk/blob/main/derisk-core/src/derisk/vis/vis_converter.py). As the SRE-Agent completes tool calls (data loading, code execution), the Vis-Agent can push intermediate `VisAgentMessages` or `VisCode` components to the front-end before the final Report-Agent completes. This provides real-time visibility into the RCA progress, though the standard workflow described in the source code emphasizes single-turn completion cycles.