# How Dify Implements ReAct-Based Agents: A Deep Dive into the Source Code

> Explore how Dify implements ReAct-based agents. Dive into the source code of CotAgentRunner to understand the stateful streaming loop, tool execution, and observation feedback driving ReAct agent logic.

- Repository: [LangGenius/dify](https://github.com/langgenius/dify)
- Tags: deep-dive
- Published: 2026-02-25

---

**Dify implements ReAct-based agents through a stateful streaming loop in `CotAgentRunner` that parses LLM outputs for "thought" and "action" tokens, executes tools via `ToolEngine`, and feeds observations back into the context until a final answer is reached.**

The `langgenius/dify` repository provides a production-grade implementation of ReAct (Reason + Act) agents that power autonomous AI workflows. Understanding how Dify structures its ReAct-based agents reveals a sophisticated architecture built on top of the generic Chain-of-Thought (CoT) runner, enabling LLMs to reason through complex tasks while dynamically invoking external tools.

## The Architecture of Dify ReAct-Based Agents

Dify’s ReAct implementation treats agent execution as a persistent state machine rather than a single-shot inference. The architecture separates configuration validation, prompt orchestration, streaming parsing, and tool execution into distinct layers.

### Configuration and Strategy Selection

Before execution begins, Dify validates the agent configuration through `AgentConfigManager` in [`api/core/app/app_config/easy_ui_based_app/agent/manager.py`](https://github.com/langgenius/dify/blob/main/api/core/app/app_config/easy_ui_based_app/agent/manager.py). An application must have its `agent_mode` dictionary configured with `"strategy": "react"` (alternatively `"cot"` works identically as they share the same runner).

The configuration manager expands default parameters and validates the tool structure, ensuring that the ReAct-based agent has access to the necessary tool definitions before the runner initializes.

### The CotAgentRunner State Machine

The heart of Dify’s ReAct implementation resides in [`api/core/agent/cot_agent_runner.py`](https://github.com/langgenius/dify/blob/main/api/core/agent/cot_agent_runner.py). The `CotAgentRunner` class maintains state across multiple LLM invocations through three critical components:

- **Scratchpad**: An `AgentScratchpadUnit` instance that accumulates thought → action → observation triples
- **Prompt History**: A running record of system prompts, user queries, and assistant responses
- **Tool Registry**: Runtime `PromptMessageTool` objects derived from the app's configured tools

## The ReAct Execution Loop: Step-by-Step

Dify’s ReAct-based agents execute through an iterative streaming loop that continues until the LLM emits a `Final Answer` action or hits the maximum iteration limit.

### Step 1: Initialization and Scratchpad Creation

When a generation request arrives, `CotAgentRunner` instantiates and immediately calls `_init_react_state(query)` (lines 52-58 in [`cot_agent_runner.py`](https://github.com/langgenius/dify/blob/main/cot_agent_runner.py)). This method initializes a fresh scratchpad and historic prompt history, preparing the state container for the first reasoning iteration.

### Step 2: Prompt Construction with Tool Definitions

Before invoking the LLM, the runner converts configured tools into model-runtime compatible formats through `_init_prompt_tools`. It then builds the complete prompt list via `_organize_prompt_messages`, inserting the current scratchpad content into the context window so the LLM can see previous thoughts and observations.

### Step 3: Streaming LLM Invocation

The runner invokes the model with `stream=True` (lines 124-132). Rather than waiting for the complete response, Dify processes the stream in real-time, feeding each chunk to the output parser. This enables immediate tool execution without buffering the entire LLM output.

### Step 4: Parsing Thoughts and Actions

The `CotAgentOutputParser.handle_react_stream_output` method in [`api/core/agent/output_parser/cot_output_parser.py`](https://github.com/langgenius/dify/blob/main/api/core/agent/output_parser/cot_output_parser.py) walks the streamed text to extract structured components:

- **Thought strings**: Identified by "thought:" markers or implicit reasoning blocks
- **Action blocks**: JSON objects or fenced-code JSON containing `action` and `action_input` keys
- **Final Answer detection**: Termination signal when the LLM indicates completion

The parser yields a mixed stream of plain text (thoughts) and `AgentScratchpadUnit.Action` objects, enabling the runner to distinguish between reasoning and executable commands.

### Step 5: Tool Execution and Observation

When the parser emits an `Action`, the runner invokes `_handle_invoke_action` (lines 81-99). This method:

1. Resolves the tool name to a `Tool` instance
2. Normalizes arguments through the tool's input schema
3. Executes via `ToolEngine.agent_invoke`
4. Captures the tool response as an **observation**

The observation appends to the scratchpad, creating the complete thought → action → observation triple that defines the ReAct pattern.

### Step 6: Iteration Until Final Answer

After each tool call, the runner updates the prompt-tool messages to include the new observation, then repeats the loop (lines 96-145). The cycle continues until:

- The LLM emits a `Final Answer` action (checked at lines 170-176)
- The maximum iteration count configured in `agent_mode` is reached

Upon termination, the final answer streams as a normal `LLMResultChunk`, while all intermediate thoughts, actions, and observations persist as `MessageAgentThought` records accessible through `AgentService.get_agent_logs` in [`api/services/agent_service.py`](https://github.com/langgenius/dify/blob/main/api/services/agent_service.py).

## Key Components and File Structure

Understanding Dify’s ReAct implementation requires familiarity with these specific source files:

| File | Role |
|------|------|
| [`api/core/agent/cot_agent_runner.py`](https://github.com/langgenius/dify/blob/main/api/core/agent/cot_agent_runner.py) | Central ReAct loop, state handling, and tool invocation orchestration |
| [`api/core/agent/output_parser/cot_output_parser.py`](https://github.com/langgenius/dify/blob/main/api/core/agent/output_parser/cot_output_parser.py) | Streaming parser that extracts **thought** and **action** tokens from LLM output |
| [`api/core/rag/retrieval/output_parser/react_output.py`](https://github.com/langgenius/dify/blob/main/api/core/rag/retrieval/output_parser/react_output.py) | Lightweight `ReactAction` data class used by the parser |
| [`api/core/app/app_config/easy_ui_based_app/agent/manager.py`](https://github.com/langgenius/dify/blob/main/api/core/app/app_config/easy_ui_based_app/agent/manager.py) | Configuration validation and default expansion for `agent_mode` |
| [`api/services/agent_service.py`](https://github.com/langgenius/dify/blob/main/api/services/agent_service.py) | API surface for retrieving agent thought logs via `get_agent_logs` |
| [`api/models/model.py`](https://github.com/langgenius/dify/blob/main/api/models/model.py) | Persistence layer for `agent_mode` JSON configuration on the `App` model |

## Practical Implementation Example

To configure a ReAct-based agent in Dify, you must enable the agent mode with the correct strategy and optional tool configurations:

```python
import json
from dify_sdk import DifyClient   # hypothetical wrapper

client = DifyClient(api_key="YOUR_API_KEY")
app_id = "12345"

# Enable ReAct agent mode

config = {
    "agent_mode": {
        "enabled": True,
        "strategy": "react",      # ReAct strategy (alias "cot" also valid)

        "tools": []              # optional tool configurations

    }
}
client.update_app(app_id, {"app_model_config": json.dumps(config)})

```

When generating responses, the streaming interface handles the ReAct loop internally:

```python
from dify_sdk import DifyClient

client = DifyClient(api_key="YOUR_API_KEY")
resp = client.generate(
    app_id="12345",
    query="What is the weather in Paris tomorrow?",
    stream=True               # Enables streaming ReAct processing

)

for chunk in resp:
    print(chunk.delta.message.content, end="")

```

To inspect the agent's reasoning trace after execution:

```python
logs = client.get_agent_logs(
    app_id="12345", 
    conversation_id="c1", 
    message_id="m1"
)

for i, iteration in enumerate(logs["iterations"]):
    print(f"## Iteration {i+1}")

    print("Thought:", iteration["thought"])
    for call in iteration["tool_calls"]:
        print(f"Tool: {call['tool_name']}")
        print(f"Input: {call['tool_input']}")
        print(f"Output: {call['tool_output']}")

```

## Summary

Dify’s ReAct-based agents operate through a sophisticated stateful loop that bridges LLM reasoning with external tool execution:

- **Configuration** requires setting `agent_mode.strategy` to `"react"` via `AgentConfigManager` in [`api/core/app/app_config/easy_ui_based_app/agent/manager.py`](https://github.com/langgenius/dify/blob/main/api/core/app/app_config/easy_ui_based_app/agent/manager.py)
- **Execution** flows through `CotAgentRunner` in [`api/core/agent/cot_agent_runner.py`](https://github.com/langgenius/dify/blob/main/api/core/agent/cot_agent_runner.py), which maintains scratchpad state across iterations
- **Parsing** uses `CotAgentOutputParser.handle_react_stream_output` to extract thoughts and actions from streaming LLM responses
- **Tool invocation** occurs via `_handle_invoke_action`, which normalizes inputs and executes through `ToolEngine.agent_invoke`
- **Persistence** stores all intermediate reasoning as `MessageAgentThought` records accessible through `AgentService.get_agent_logs`

## Frequently Asked Questions

### What is the difference between ReAct and CoT strategies in Dify?

Dify uses the same `CotAgentRunner` for both ReAct and CoT strategies. The `"react"` strategy explicitly enables the tool-calling loop where the LLM generates actions to interact with external tools, while `"cot"` (Chain-of-Thought) may operate without tool invocation depending on configuration. Both strategies share the same parsing logic in [`cot_output_parser.py`](https://github.com/langgenius/dify/blob/main/cot_output_parser.py) and state management through `AgentScratchpadUnit`.

### How does Dify handle malformed tool calls during ReAct execution?

The `CotAgentOutputParser` in [`api/core/agent/output_parser/cot_output_parser.py`](https://github.com/langgenius/dify/blob/main/api/core/agent/output_parser/cot_output_parser.py) recognizes multiple action formats including JSON code fences and raw JSON blocks. If the parser encounters malformed JSON or missing action keys, it continues streaming the text as plain content rather than executing a tool. The `_handle_invoke_action` method in [`cot_agent_runner.py`](https://github.com/langgenius/dify/blob/main/cot_agent_runner.py) includes argument normalization that validates inputs against the tool's schema before calling `ToolEngine.agent_invoke`, preventing execution of invalid tool calls.

### Can I observe the intermediate reasoning steps of a ReAct agent in Dify?

Yes, Dify persists all intermediate thoughts, actions, and observations as `MessageAgentThought` records. You can retrieve these through `AgentService.get_agent_logs` in [`api/services/agent_service.py`](https://github.com/langgenius/dify/blob/main/api/services/agent_service.py), which queries the database for a specific conversation and message ID. The logs reveal the complete scratchpad history including the LLM's reasoning at each iteration, the exact tool inputs generated, and the observations returned from tool execution, enabling full transparency into the agent's decision-making process.

### What triggers the termination of a ReAct agent loop in Dify?

The `CotAgentRunner` terminates the ReAct loop when it detects a `Final Answer` action in the LLM output, checked at lines 170-176 of [`cot_agent_runner.py`](https://github.com/langgenius/dify/blob/main/cot_agent_runner.py). Alternatively, the loop exits when the iteration count exceeds the maximum configured in the `agent_mode` settings. The parser in [`cot_output_parser.py`](https://github.com/langgenius/dify/blob/main/cot_output_parser.py) recognizes various forms of final answer declarations, ensuring the agent stops reasoning and returns the completed response to the user rather than continuing to invoke tools indefinitely.