# How the DeepResearch RCA Engine Analyzes Logs, Traces, and Code to Locate Root Causes

> Discover how the DeepResearch RCA engine uses ReAct reasoning to analyze logs, traces, and code for pinpointing system failure causes. Get precise root cause analysis.

- Repository: [derisk-ai/openderisk](https://github.com/derisk-ai/openderisk)
- Tags: deep-dive
- Published: 2026-02-28

---

**The DeepResearch RCA engine is a multi-agent AI system that ingests observability data from the OpenRCA dataset and executes a ReAct reasoning loop—alternating between LLM-driven planning and sub-agent actions—to pinpoint the exact timestamp, component, and cause of system failures.**

The DeepResearch RCA engine transforms raw telemetry into precise diagnoses by orchestrating specialized AI agents against structured failure scenarios. Built within the `derisk-ai/openderisk` repository, this engine leverages the **OpenRCA dataset** (containing logs, metric traces, and code snapshots) to automate root cause analysis without hard-coded file paths or brittle regex patterns.

## Stage 1: Scene & Data Loading with `OpenRcaSceneResource`

The analysis begins when the engine resolves a user-selected **scene**—a logical boundary representing a specific system environment such as *bank*, *telecom*, or *market*. 

In [`packages/derisk-ext/src/derisk_ext/agent/agents/open_rca/resource/open_rca_resource.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-ext/src/derisk_ext/agent/agents/open_rca/resource/open_rca_resource.py), the `OpenRcaSceneResource` class reads the scene enumeration, constructs the local data path under `dataset/openrca/<Scene>`, and exposes resource parameters to downstream agents. This abstraction ensures the master agent never relies on hard-coded directories.

```python
from derisk_ext.agent.agents.open_rca.resource.open_rca_resource import OpenRcaSceneResource

scene_res = OpenRcaSceneResource(scene="bank")

```

The resource automatically locates the scene’s log files, trace data, and optional code snapshots, making the dataset self-describing through the `scene_description` and `data_path` attributes.

## Stage 2: Prompt Generation via `open_rca_scene_prompt_template`

Once the resource resolves the data path, the engine renders a structured scene prompt that contextualizes the failure for the LLM. The `open_rca_scene_prompt_template` (defined in the same resource file) injects the scene name, description, and data directory into an XML-formatted block.

The rendered prompt structure follows this schema:

```xml
<open-rca-scene>
  <scene-info>
    <name>bank</name>
    <description>银行微服务系统场景，包含Tomcat、MySQL、Redis等组件的监控数据</description>
    <data_path>/path/to/dataset/openrca/Bank</data_path>
  </scene-info>
  <scene-background>
    …(optional schema)…
  </scene-background>
</open-rca-scene>

```

This template provides the master agent with a concise view of the dataset schema and investigation scope, grounding the LLM in the specific system topology before reasoning begins.

## Stage 3: Multi-Agent Reasoning Loop via `OpenRcaReActMasterAgent`

The **OpenRcaReActMasterAgent**—a thin wrapper around the generic `ReActMasterAgent` implemented in [`packages/derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-core/src/derisk/agent/expand/react_master_agent/react_master_agent.py)—orchestrates the core reasoning cycle. This agent alternates between *thinking* (generating plans) and *acting* (invoking specialized sub-agents) until it satisfies the diagnosis constraints.

The wrapper class is registered in [`packages/derisk-ext/src/derisk_ext/agent/agents/open_rca/__init__.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-ext/src/derisk_ext/agent/agents/open_rca/__init__.py) and delegates to three specialized agents:

- **SRE-Agent** – Parses raw logs and trace files under `data_path/logs/` and `data_path/trace/`, extracting candidate error events through pattern matching and timestamp parsing.
- **Code-Agent** (`IpythonAssistantAgent`) – Generates and executes Python snippets (often using pandas or NumPy) to perform statistical checks on extracted metrics or analyze code snapshots in a sandboxed interpreter.
- **Report-Agent** – Compiles the final diagnosis into a human-readable markdown report that complies with the OpenRCA skill specification.

The ReAct loop continues iteratively: the master agent proposes an action (e.g., "search logs for ERROR patterns"), the SRE-Agent returns observations, the Code-Agent validates anomalies with statistical analysis, and the master refines its hypothesis.

## Stage 4: Root-Cause Extraction via Skill Specifications

The reasoning loop terminates only when the master agent produces a *final answer* satisfying the **OpenRCA diagnosis skill** constraints. These constraints are encoded in markdown specification files such as [`bank_spec.md`](https://github.com/derisk-ai/openderisk/blob/main/bank_spec.md), [`telecom_spec.md`](https://github.com/derisk-ai/openderisk/blob/main/telecom_spec.md), and [`market_spec.md`](https://github.com/derisk-ai/openderisk/blob/main/market_spec.md) located in `packages/derisk-ext/src/derisk_ext/agent/agents/open_rca/skills/open_rca_diagnosis/specs/`.

A valid root-cause answer must include:

- The exact timestamp of the root-cause event.
- The responsible component or service name.
- A concise causal explanation and remediation hint.

The master agent validates its output against these schema requirements before returning, ensuring the diagnosis contains actionable intelligence rather than vague observations.

## End-to-End Data Flow Implementation

### Executing a Full Diagnosis

To run the complete pipeline from Python, instantiate the client with the scene resource and failure description:

```python
from derisk_client import DeriskClient
from derisk_ext.agent.agents.open_rca.resource.open_rca_resource import OpenRcaSceneResource

client = DeriskClient()
scene_res = OpenRcaSceneResource(scene="bank")

failure_desc = (
    "During the period 2020-05-23 01:30-02:00 a single request failed. "
    "The exact time of the root-cause is unknown."
)

report = client.run_agent(
    agent_name="OpenRcaReActMasterAgent",
    resources=[scene_res],
    user_input=failure_desc,
)

print(report)  # Markdown report with timestamp, component, and remediation

```

The execution chain follows: `DeriskClient.run_agent` → `AgentManager` → `ReActMasterAgent` (generic implementation) → `OpenRcaReActMasterAgent` wrapper.

### Manual Sub-Agent Orchestration

For advanced debugging or custom workflows, you can invoke sub-agents directly without the master loop:

```python
from derisk_ext.agent.agents.open_rca.resource.open_rca_resource import OpenRcaSceneResource
from derisk.agent import Agent

scene_res = OpenRcaSceneResource(scene="bank")
data_path = scene_res.data_path

# 1. Mine logs with SRE-Agent

sre_agent = Agent.get("SRE-Agent")
log_file = f"{data_path}/logs/app.log"
error_lines = sre_agent.run_tool("grep", {"pattern": "ERROR", "file_path": log_file})

# 2. Statistical analysis with Code-Agent

code_agent = Agent.get("IpythonAssistantAgent")
snippet = f"""
import pandas as pd
df = pd.read_csv("{data_path}/metrics/latency.csv")
df['ts'] = pd.to_datetime(df['timestamp'])
error_ts = pd.to_datetime('{error_lines[0].split()[0]}')
window = df[(df['ts'] >= error_ts - pd.Timedelta('5m')) & 
            (df['ts'] <= error_ts + pd.Timedelta('5m'))]
print(window.describe())
"""
analysis = code_agent.execute(snippet)

# 3. Format output with Report-Agent

report_agent = Agent.get("ReportAgent")
final_report = report_agent.format_report(
    timestamp=error_lines[0].split()[0],
    component="payment-service",
    root_cause="database connection pool exhaustion",
    details=analysis,
)

```

### Extending with New Scenes

To add a custom scene (e.g., *ecommerce*), extend the `OpenRcaScene` enum in [`packages/derisk-ext/src/derisk_ext/agent/agents/open_rca/resource/open_rca_base.py`](https://github.com/derisk-ai/openderisk/blob/main/packages/derisk-ext/src/derisk_ext/agent/agents/open_rca/resource/open_rca_base.py) and place the dataset under `dataset/openrca/Ecommerce`. The engine automatically discovers the new scene via `get_open_rca_scenes()` and builds the resource parameters without additional configuration.

## Summary

- The **DeepResearch RCA engine** processes OpenRCA datasets through a four-stage pipeline: scene loading, prompt generation, multi-agent reasoning, and skill-validated extraction.
- **`OpenRcaSceneResource`** abstracts dataset paths, allowing the engine to resolve logs, traces, and code locations dynamically based on the selected scene enum.
- The **ReAct reasoning loop** in [`react_master_agent.py`](https://github.com/derisk-ai/openderisk/blob/main/react_master_agent.py) drives the investigation, delegating concrete actions to specialized SRE, Code, and Report agents.
- **Skill specifications** ([`bank_spec.md`](https://github.com/derisk-ai/openderisk/blob/main/bank_spec.md), etc.) enforce structured output schemas, ensuring every diagnosis includes the root-cause timestamp, responsible component, and remediation guidance.
- Sub-agents execute in isolation: the **SRE-Agent** handles log parsing, the **Code-Agent** (`IpythonAssistantAgent`) performs dynamic Python analysis, and the **Report-Agent** formats the final deliverable.

## Frequently Asked Questions

### What datasets does the DeepResearch RCA engine support by default?

The engine supports the **OpenRCA dataset** scenes: *bank* (financial microservices), *telecom* (communication infrastructure), and *market* (trading systems). Each scene contains approximately 26GB of logs, metric traces, and code snapshots. You can extend support by adding new entries to the `OpenRcaScene` enum in [`open_rca_base.py`](https://github.com/derisk-ai/openderisk/blob/main/open_rca_base.py) and placing corresponding data under `dataset/openrca/<SceneName>`.

### How does the ReAct pattern prevent hallucinations during root cause analysis?

The **ReAct** (Reasoning + Acting) pattern forces the LLM to alternate between *thinking* (planning) and *acting* (executing verified sub-agent tools). Instead of generating conclusions from memory, the master agent in [`react_master_agent.py`](https://github.com/derisk-ai/openderisk/blob/main/react_master_agent.py) requests concrete data—such as specific log lines or statistical calculations—at each step. This grounding in actual file contents and execution results prevents the model from inventing timestamps or components that do not exist in the dataset.

### Can the engine analyze log formats outside the OpenRCA dataset structure?

Yes, though it requires extending the **SRE-Agent** implementation. The current agent expects standard log and trace directories under the scene’s `data_path`, but the modular architecture in `derisk-core` allows you to register custom parsing skills. By modifying the skill registration in the OpenRCA package [`__init__.py`](https://github.com/derisk-ai/openderisk/blob/main/__init__.py), you can inject handlers for JSON, XML, or proprietary binary log formats while retaining the same ReAct master loop.

### What is the role of the skill specification markdown files?

Files like [`bank_spec.md`](https://github.com/derisk-ai/openderisk/blob/main/bank_spec.md) serve as **contracts** between the reasoning engine and the output consumer. They define the required fields for a valid root-cause diagnosis (timestamp, component, cause, remediation) and any validation constraints. The `OpenRcaReActMasterAgent` checks its final answer against these specifications before terminating, ensuring the output is machine-parseable and actionable for downstream automation tools.