How Dify Implements ReAct-Based Agents: A Deep Dive into the Source Code
Dify implements ReAct-based agents through a stateful streaming loop in CotAgentRunner that parses LLM outputs for "thought" and "action" tokens, executes tools via ToolEngine, and feeds observations back into the context until a final answer is reached.
The langgenius/dify repository provides a production-grade implementation of ReAct (Reason + Act) agents that power autonomous AI workflows. Understanding how Dify structures its ReAct-based agents reveals a sophisticated architecture built on top of the generic Chain-of-Thought (CoT) runner, enabling LLMs to reason through complex tasks while dynamically invoking external tools.
The Architecture of Dify ReAct-Based Agents
Dify’s ReAct implementation treats agent execution as a persistent state machine rather than a single-shot inference. The architecture separates configuration validation, prompt orchestration, streaming parsing, and tool execution into distinct layers.
Configuration and Strategy Selection
Before execution begins, Dify validates the agent configuration through AgentConfigManager in api/core/app/app_config/easy_ui_based_app/agent/manager.py. An application must have its agent_mode dictionary configured with "strategy": "react" (alternatively "cot" works identically as they share the same runner).
The configuration manager expands default parameters and validates the tool structure, ensuring that the ReAct-based agent has access to the necessary tool definitions before the runner initializes.
The CotAgentRunner State Machine
The heart of Dify’s ReAct implementation resides in api/core/agent/cot_agent_runner.py. The CotAgentRunner class maintains state across multiple LLM invocations through three critical components:
- Scratchpad: An
AgentScratchpadUnitinstance that accumulates thought → action → observation triples - Prompt History: A running record of system prompts, user queries, and assistant responses
- Tool Registry: Runtime
PromptMessageToolobjects derived from the app's configured tools
The ReAct Execution Loop: Step-by-Step
Dify’s ReAct-based agents execute through an iterative streaming loop that continues until the LLM emits a Final Answer action or hits the maximum iteration limit.
Step 1: Initialization and Scratchpad Creation
When a generation request arrives, CotAgentRunner instantiates and immediately calls _init_react_state(query) (lines 52-58 in cot_agent_runner.py). This method initializes a fresh scratchpad and historic prompt history, preparing the state container for the first reasoning iteration.
Step 2: Prompt Construction with Tool Definitions
Before invoking the LLM, the runner converts configured tools into model-runtime compatible formats through _init_prompt_tools. It then builds the complete prompt list via _organize_prompt_messages, inserting the current scratchpad content into the context window so the LLM can see previous thoughts and observations.
Step 3: Streaming LLM Invocation
The runner invokes the model with stream=True (lines 124-132). Rather than waiting for the complete response, Dify processes the stream in real-time, feeding each chunk to the output parser. This enables immediate tool execution without buffering the entire LLM output.
Step 4: Parsing Thoughts and Actions
The CotAgentOutputParser.handle_react_stream_output method in api/core/agent/output_parser/cot_output_parser.py walks the streamed text to extract structured components:
- Thought strings: Identified by "thought:" markers or implicit reasoning blocks
- Action blocks: JSON objects or fenced-code JSON containing
actionandaction_inputkeys - Final Answer detection: Termination signal when the LLM indicates completion
The parser yields a mixed stream of plain text (thoughts) and AgentScratchpadUnit.Action objects, enabling the runner to distinguish between reasoning and executable commands.
Step 5: Tool Execution and Observation
When the parser emits an Action, the runner invokes _handle_invoke_action (lines 81-99). This method:
- Resolves the tool name to a
Toolinstance - Normalizes arguments through the tool's input schema
- Executes via
ToolEngine.agent_invoke - Captures the tool response as an observation
The observation appends to the scratchpad, creating the complete thought → action → observation triple that defines the ReAct pattern.
Step 6: Iteration Until Final Answer
After each tool call, the runner updates the prompt-tool messages to include the new observation, then repeats the loop (lines 96-145). The cycle continues until:
- The LLM emits a
Final Answeraction (checked at lines 170-176) - The maximum iteration count configured in
agent_modeis reached
Upon termination, the final answer streams as a normal LLMResultChunk, while all intermediate thoughts, actions, and observations persist as MessageAgentThought records accessible through AgentService.get_agent_logs in api/services/agent_service.py.
Key Components and File Structure
Understanding Dify’s ReAct implementation requires familiarity with these specific source files:
| File | Role |
|---|---|
api/core/agent/cot_agent_runner.py |
Central ReAct loop, state handling, and tool invocation orchestration |
api/core/agent/output_parser/cot_output_parser.py |
Streaming parser that extracts thought and action tokens from LLM output |
api/core/rag/retrieval/output_parser/react_output.py |
Lightweight ReactAction data class used by the parser |
api/core/app/app_config/easy_ui_based_app/agent/manager.py |
Configuration validation and default expansion for agent_mode |
api/services/agent_service.py |
API surface for retrieving agent thought logs via get_agent_logs |
api/models/model.py |
Persistence layer for agent_mode JSON configuration on the App model |
Practical Implementation Example
To configure a ReAct-based agent in Dify, you must enable the agent mode with the correct strategy and optional tool configurations:
import json
from dify_sdk import DifyClient # hypothetical wrapper
client = DifyClient(api_key="YOUR_API_KEY")
app_id = "12345"
# Enable ReAct agent mode
config = {
"agent_mode": {
"enabled": True,
"strategy": "react", # ReAct strategy (alias "cot" also valid)
"tools": [] # optional tool configurations
}
}
client.update_app(app_id, {"app_model_config": json.dumps(config)})
When generating responses, the streaming interface handles the ReAct loop internally:
from dify_sdk import DifyClient
client = DifyClient(api_key="YOUR_API_KEY")
resp = client.generate(
app_id="12345",
query="What is the weather in Paris tomorrow?",
stream=True # Enables streaming ReAct processing
)
for chunk in resp:
print(chunk.delta.message.content, end="")
To inspect the agent's reasoning trace after execution:
logs = client.get_agent_logs(
app_id="12345",
conversation_id="c1",
message_id="m1"
)
for i, iteration in enumerate(logs["iterations"]):
print(f"## Iteration {i+1}")
print("Thought:", iteration["thought"])
for call in iteration["tool_calls"]:
print(f"Tool: {call['tool_name']}")
print(f"Input: {call['tool_input']}")
print(f"Output: {call['tool_output']}")
Summary
Dify’s ReAct-based agents operate through a sophisticated stateful loop that bridges LLM reasoning with external tool execution:
- Configuration requires setting
agent_mode.strategyto"react"viaAgentConfigManagerinapi/core/app/app_config/easy_ui_based_app/agent/manager.py - Execution flows through
CotAgentRunnerinapi/core/agent/cot_agent_runner.py, which maintains scratchpad state across iterations - Parsing uses
CotAgentOutputParser.handle_react_stream_outputto extract thoughts and actions from streaming LLM responses - Tool invocation occurs via
_handle_invoke_action, which normalizes inputs and executes throughToolEngine.agent_invoke - Persistence stores all intermediate reasoning as
MessageAgentThoughtrecords accessible throughAgentService.get_agent_logs
Frequently Asked Questions
What is the difference between ReAct and CoT strategies in Dify?
Dify uses the same CotAgentRunner for both ReAct and CoT strategies. The "react" strategy explicitly enables the tool-calling loop where the LLM generates actions to interact with external tools, while "cot" (Chain-of-Thought) may operate without tool invocation depending on configuration. Both strategies share the same parsing logic in cot_output_parser.py and state management through AgentScratchpadUnit.
How does Dify handle malformed tool calls during ReAct execution?
The CotAgentOutputParser in api/core/agent/output_parser/cot_output_parser.py recognizes multiple action formats including JSON code fences and raw JSON blocks. If the parser encounters malformed JSON or missing action keys, it continues streaming the text as plain content rather than executing a tool. The _handle_invoke_action method in cot_agent_runner.py includes argument normalization that validates inputs against the tool's schema before calling ToolEngine.agent_invoke, preventing execution of invalid tool calls.
Can I observe the intermediate reasoning steps of a ReAct agent in Dify?
Yes, Dify persists all intermediate thoughts, actions, and observations as MessageAgentThought records. You can retrieve these through AgentService.get_agent_logs in api/services/agent_service.py, which queries the database for a specific conversation and message ID. The logs reveal the complete scratchpad history including the LLM's reasoning at each iteration, the exact tool inputs generated, and the observations returned from tool execution, enabling full transparency into the agent's decision-making process.
What triggers the termination of a ReAct agent loop in Dify?
The CotAgentRunner terminates the ReAct loop when it detects a Final Answer action in the LLM output, checked at lines 170-176 of cot_agent_runner.py. Alternatively, the loop exits when the iteration count exceeds the maximum configured in the agent_mode settings. The parser in cot_output_parser.py recognizes various forms of final answer declarations, ensuring the agent stops reasoning and returns the completed response to the user rather than continuing to invoke tools indefinitely.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →