How the Qwen-Agent Reasoning Process Works: From Prompt to Tool Execution
The Qwen-Agent reasoning process is an iterative orchestration loop that drives a large language model through cycles of prompt generation, LLM inference, tool-call detection, and context updates until the model produces a final answer without requesting additional tools.
The Qwen-Agent framework implements this reasoning pipeline through a thin abstraction layer that manages conversation state and tool execution. According to the QwenLM/Qwen-Agent source code, the architecture relies on three primary classes—Agent, FnCallAgent, and ReActChat—to handle everything from standard function calling to structured ReAct-style reasoning chains.
Core Architecture of the Reasoning Pipeline
The reasoning process is distributed across three key classes that define distinct stages of the agent lifecycle:
| Class | Primary File | Responsibility |
|---|---|---|
Agent |
qwen_agent/agent.py |
Base class that initializes the LLM, registers tools in self.function_map, and provides the generic _detect_tool method for parsing function calls. |
FnCallAgent |
qwen_agent/agents/fncall_agent.py |
Implements the core LLM-tool loop, managing the iteration cap (MAX_LLM_CALL_PER_RUN) and delegating to tool execution handlers. |
ReActChat |
qwen_agent/agents/react_chat.py |
Extends FnCallAgent with a ReAct prompt template and custom parsing for Thought/Action/Observation chains. |
During initialization, the base Agent.__init__ instantiates the LLM via get_chat_model ([lines 56–59 in agent.py](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py#L56-L59)) and populates the tool registry from the function_list parameter ([lines 62–66 in agent.py](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py#L62-L66)). Subclasses like FnCallAgent additionally initialize a Memory object for file handling ([lines 55–71 in fncall_agent.py](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agents/fncall_agent.py#L55-L71)).
The Standard Reasoning Loop in FnCallAgent
The heart of the Qwen-Agent reasoning process resides in FnCallAgent._run, which implements a defensive loop bounded by MAX_LLM_CALL_PER_RUN. This method operates on a deep-copied message list to preserve conversation history immutability across iterations.
Step 1: LLM Invocation with Tool Schemas
Each iteration begins with a call to self._call_llm, which forwards the current messages list to the underlying model along with the JSON schema of all registered tools:
output_stream = self._call_llm(
messages=messages,
functions=[func.function for func in self.function_map.values()],
extra_generate_cfg=extra_generate_cfg
)
The LLM receives both the conversation context and function definitions, enabling it to decide whether to generate text or request a tool invocation.
Step 2: Tool Detection and Parsing
After the LLM returns output, Agent._detect_tool ([lines 48–59 in agent.py](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py#L48-L59)) analyzes the response for an OpenAI-style function_call field. This method returns a boolean flag (use_tool), the tool name, and parsed arguments, effectively bridging the LLM's output format with the agent's execution logic.
Step 3: Tool Execution and Context Updates
When use_tool evaluates to true, FnCallAgent triggers _call_tool ([lines 78–110 in agent.py](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py#L78-L110)), which executes the Python implementation of the requested tool. The result is wrapped in a Message object with role FUNCTION and appended to both the internal messages list (for the next LLM turn) and the cumulative response:
fn_msg = Message(
role=FUNCTION,
name=tool_name,
content=tool_result,
extra={'function_id': out.extra.get('function_id', '1')}
)
messages.append(fn_msg)
The loop repeats until the LLM ceases requesting tools or the iteration limit is exhausted.
ReAct-Style Reasoning Implementation
ReActChat customizes the generic loop to enforce structured reasoning through a prompt-engineering approach defined in PROMPT_REACT ([lines 29–46 in react_chat.py](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agents/react_chat.py#L29-L46)).
The ReAct Prompt Template
Before the first LLM call, _prepend_react_prompt ([lines 109–132 in react_chat.py](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agents/react_chat.py#L109-L132)) injects a template that forces the model to output explicit reasoning steps:
- Thought: Internal reasoning about the task
- Action: The tool name to invoke
- Action Input: The arguments for the tool
- Observation: The result returned from tool execution
Parsing Thought-Action-Observation Chains
Unlike the JSON-based function_call detection in the base class, ReActChat._detect_tool ([lines 34–51 in react_chat.py](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agents/react_chat.py#L34-L51)) scans the raw text output for specific markers:
\nAction:\nAction Input:\nObservation:
When detected, the agent extracts the action and input, executes the tool, and injects the result as an Observation back into the prompt history. This yields a visible reasoning trace that can be streamed to users via the yield statements at lines 85–99, providing transparency into the agent's intermediate thought process.
Practical Example: Running a ReAct Agent
The following demonstrates the complete reasoningcycle using the ReActChat implementation:
from qwen_agent.agents.react_chat import ReActChat
# Initialize agent with code interpreter tool
agent = ReActChat(
function_list=['code_interpreter'],
llm={'model': 'qwen-turbo', 'model_type': 'qwen_dashscope'}
)
# Define user query
messages = [
{"role": "user", "content": "Calculate the sum of the first 100 prime numbers."}
]
# Execute non-streaming run
final_response = list(agent.run_nonstream(messages))[0]
print(final_response.content)
Expected reasoning trace:
Thought: I should compute the primes first.
Action: code_interpreter
Action Input: ```python
import sympy as sp
primes = list(sp.primerange(1, 1000))[:100]
print(sum(primes))
Observation: 24133 Thought: I now know the final answer Final Answer: The sum of the first 100 prime numbers is 24,133.
This trace illustrates the完整 reasoning loop: the LLM thinks, acts, observes the tool result, reflects, and terminates with a final answer.
## Summary
The Qwen-Agent reasoning process operates through a tightly controlled iteration mechanism:
- **Three-class architecture**: `Agent` provides base functionality, `FnCallAgent` implements the core loop, and `ReActChat` adds structured prompting.
- **Mutable context management**: Each tool result is appended as a `FUNCTION` message to the conversation history, enabling multi-turn reasoning.
- **Dual parsing strategies**: Standard function calls use JSON detection (`_detect_tool`), while ReAct mode parses custom text markers for Thought/Action/Observation chains.
- **Safety limits**: `MAX_LLM_CALL_PER_RUN` prevents infinite loops by bounding the maximum number of LLM invocations per agent run.
- **File integration**: The `Memory` system in `FnCallAgent` handles document uploads when tools require external data access.
## Frequently Asked Questions
### What terminates the reasoning loop in Qwen-Agent?
The loop exits when the LLM response contains no tool requests (detected by `_detect_tool` returning `use_tool=False`) or when the iteration counter reaches `MAX_LLM_CALL_PER_RUN`, a constant defined in [`qwen_agent/settings.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/settings.py). This prevents runaway execution while allowing complex multi-step tasks to complete.
### How does ReActChat differ from standard function calling?
While `FnCallAgent` relies on the LLM outputting a `function_call` JSON structure, `ReActChat` in [`qwen_agent/agents/react_chat.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agents/react_chat.py) preprocesses prompts with a ReAct template and overrides `_detect_tool` to parse plain-text markers (`\nAction:`, `\nAction Input:`). This produces a human-readable reasoning trace but requires more tokens for the structured output format.
### Where is the maximum number of reasoning steps configured?
The global constant `MAX_LLM_CALL_PER_RUN` controls the upper bound of LLM invocations per agent execution. This setting lives in [`qwen_agent/settings.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/settings.py) and is referenced in `FnCallAgent._run` to decrement the available call counter during each iteration of the reasoning loop.
### Can developers add custom tools to the reasoning process?
Yes. Any class inheriting from `BaseTool` can be registered via the `function_list` parameter during agent initialization. The base `Agent` class automatically adds these to `self.function_map` ([lines 62–66 in [`agent.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/agent.py)](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py#L62-L66)), making their schemas available to `_call_llm` and their implementations accessible to `_call_tool` without modifying the core reasoning logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →