Core Components of Dify's Agent Service: Architecture and Implementation Guide

Dify's agent service is built as a layered, plugin-friendly stack comprising API services, workflow nodes, agent runners, domain entities, and infrastructure abstractions that enable reactive Chain-of-Thought (CoT) reasoning with tool integration.

The agent service in the langgenius/dify repository provides the infrastructure for building reactive AI agents that can reason, use tools, and maintain state across conversations. Understanding the core components of Dify's agent service is essential for developers extending the platform or debugging agent behavior. The architecture separates concerns between workflow orchestration, prompt handling, and execution logic, allowing modular integration of new LLM providers and custom tools.

API Entry Point: AgentService

The AgentService class in api/services/agent_service.py serves as the central entry point for UI and external callers. This service exposes endpoints for retrieving agent logs, listing available agents, and fetching provider-specific configurations.

The get_agent_logs method aggregates execution metadata from MessageAgentThought records, returning structured data containing iteration history, tool calls, and file attachments. This method handles the log structure assembly for the front-end consumption, spanning lines 19–84 of the implementation.

Workflow Integration: AgentNode

The AgentNode class in api/core/workflow/nodes/agent/agent_node.py bridges the workflow engine and agent execution logic. As a Node implementation, it translates Dify workflow steps into agent runs by handling parameter resolution, credential fetching, and memory injection.

Key responsibilities include:

  • Parameter generation: The _generate_agent_parameters method resolves credentials and model configurations before execution
  • Message transformation: The _transform_message method converts tool invocation streams into workflow-compatible output events
  • Execution orchestration: The _run method (lines 80–98) coordinates the streaming of events including StreamChunkEvent, AgentLogEvent, and final completion signals

The Agent Runner Hierarchy

Dify implements a tiered runner architecture that separates generic agent logic from specific prompt implementations.

BaseAgentRunner and the ReAct Loop

The abstract BaseAgentRunner in api/core/agent/base_agent_runner.py implements the generic thought-action-observation loop (ReAct pattern) used by all concrete runners. It manages the scratch-pad, prompt history, and tool callbacks while providing methods like create_agent_thought and save_agent_thought for persisting intermediate reasoning states.

The runner initializes tool instances through _init_prompt_tools (lines 24–44), constructing the mapping between tool definitions and runtime callable objects.

CoT Agent Runners

The CotAgentRunner in api/core/agent/cot_agent_runner.py extends the base runner with Chain-of-Thought iteration logic. It handles:

  • Streaming LLM output processing
  • Action parsing from model responses
  • Sequential tool invocation management

Two thin specializations adapt this for different prompt styles:

Domain Entities and State Persistence

The AgentEntity, AgentToolEntity, AgentPromptEntity, and AgentScratchpadUnit classes in api/core/agent/entities.py define the Pydantic schemas that serialize configuration and runtime state. These models establish the contract between the UI, workflow engine, and runner layers, ensuring type-safe data flow for agent configurations, tool definitions, and scratch-pad states.

Tool Integration and Memory Management

Tool Management

The ToolManager in api/core/tools/tool_manager.py acts as a factory that transforms tool definitions into runtime Tool instances. For dataset retrieval operations, the DatasetRetrieverTool in api/core/tools/utils/dataset_retriever_tool.py provides specialized integration with Dify's knowledge bases.

Memory Abstraction

The TokenBufferMemory class in api/core/memory/token_buffer_memory.py maintains conversational continuity by storing recent messages in a token-aware buffer. This component queries historic context for prompt injection, ensuring agents maintain awareness across multi-turn interactions while respecting token limits.

Model Abstraction and Plugin Architecture

The ModelManager and ModelInstance classes in api/core/model_manager.py, along with the LargeLanguageModel base class, provide provider-agnostic access to LLMs from OpenAI, Anthropic, and other supported services. This abstraction supplies runners with streaming-compatible model interfaces.

For extensibility, the PluginAgentClient and PluginAgentStrategy in api/core/plugin/impl/agent.py enable third-party agent strategies to integrate without core code modifications. This plugin bridge allows custom reasoning algorithms to replace or augment the standard CoT implementation.

Observability and Debugging

The DifyAgentCallbackHandler in api/core/callback_handler/agent_tool_callback_handler.py provides execution visibility by intercepting tool start/end events. When DEBUG mode is active, it logs detailed execution traces; in production, it forwards structured traces to the Ops queue for monitoring and audit trails.

Practical Implementation Examples

Fetching Agent Logs via Service Layer

from api.services.agent_service import AgentService
from models.model import App

def get_logs(app: App, conv_id: str, msg_id: str):
    # Returns a dict containing meta, iterations, and attached files

    return AgentService.get_agent_logs(app, conv_id, msg_id)

Source: AgentService.get_agent_logs assembles log structures from MessageAgentThought rows (lines 19–84 of api/services/agent_service.py).

Executing an Agent Within a Workflow

from api.core.workflow.nodes.agent.agent_node import AgentNode
from api.core.workflow.graph import WorkflowGraph

# Assume a pre-constructed workflow graph `graph`

node = AgentNode(node_id="agent_1", tenant_id="t1", app_id="app1", node_data=agent_node_data)
graph.add_node(node)

# When the workflow executor reaches this node:

for event in node.run():
    # `event` can be a StreamChunkEvent, AgentLogEvent, or the final StreamCompletedEvent

    handle(event)

Source: Execution initiates at AgentNode._run (lines 80–98) and streams via _transform_message in api/core/workflow/nodes/agent/agent_node.py.

Extending BaseAgentRunner with Custom Tools

from api.core.agent.base_agent_runner import BaseAgentRunner
from api.core.tools.__base.tool import Tool

class MyRunner(BaseAgentRunner):
    def _init_prompt_tools(self):
        # Reuse BaseAgentRunner logic, then inject a custom tool

        instances, prompts = super()._init_prompt_tools()
        my_tool = MyCustomTool()               # must inherit from core.tools.__base.tool.Tool

        instances["my_tool"] = my_tool
        prompts.append(
            PromptMessageTool(
                name="my_tool",
                description="My special utility",
                parameters={"type": "object", "properties": {}, "required": []},
            )
        )
        return instances, prompts

Source: Base runner constructs tool instances in _init_prompt_tools (lines 24–44 of api/core/agent/base_agent_runner.py).

Summary

  • AgentService in api/services/agent_service.py provides the external API for log retrieval and agent management
  • AgentNode in api/core/workflow/nodes/agent/agent_node.py embeds agent execution within Dify's workflow engine
  • BaseAgentRunner and CotAgentRunner in api/core/agent/ implement the ReAct reasoning loop and CoT iteration logic
  • Domain entities in api/core/agent/entities.py define the schemas for configuration, tools, and scratch-pad state
  • ToolManager and TokenBufferMemory handle tool resolution and conversational memory respectively
  • Plugin architecture via api/core/plugin/impl/agent.py allows third-party agent strategies without core modifications
  • Observability is handled by DifyAgentCallbackHandler for execution tracing and debugging

Frequently Asked Questions

What is the difference between CotAgentRunner and BaseAgentRunner?

BaseAgentRunner provides the abstract ReAct (thought-action-observation) loop and state persistence mechanisms, while CotAgentRunner implements the specific Chain-of-Thought iteration logic including streaming LLM output processing and action parsing. The base class handles the generic orchestration infrastructure, whereas the CoT variant adds the concrete reasoning strategy used by Dify's standard agents.

How does Dify handle tool integration in agents?

Dify uses the ToolManager in api/core/tools/tool_manager.py to resolve tool definitions into runtime Tool instances that the runner can invoke. When the agent decides to use a tool, the runner calls the tool's execution method and transforms the output into observation messages that feed back into the LLM context, creating the iterative tool-use loop characteristic of ReAct agents.

Where is agent state persisted during execution?

Agent state persists through the AgentScratchpadUnit entities defined in api/core/agent/entities.py, with methods like create_agent_thought and save_agent_thought in BaseAgentRunner managing the storage of intermediate reasoning steps. Additionally, TokenBufferMemory maintains conversation history across turns, while workflow state is managed by the AgentNode integration.

How can I customize the agent's reasoning strategy?

You can extend BaseAgentRunner or CotAgentRunner to implement custom reasoning logic, or use the PluginAgentClient and PluginAgentStrategy in api/core/plugin/impl/agent.py to register external agent implementations. This plugin architecture allows swapping in entirely new reasoning strategies without modifying Dify's core codebase, enabling integration of custom CoT variations or alternative agent architectures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →