How to Fine-Tune Qwen-Agent for Specific Tasks: A Complete Technical Guide
Fine-tune Qwen-Agent by serving your customized Qwen model via an OpenAI-compatible endpoint, configuring the model_server parameter in get_chat_model, and specializing agent behavior through custom system prompts and domain-specific tools.
Qwen-Agent is a modular framework developed by QwenLM that orchestrates large language models, external tools, and memory systems to build task-specific AI agents. While the core library does not ship with a built-in training loop, you can adapt Qwen-Agent to specialized domains by integrating a fine-tuned Qwen checkpoint and leveraging the framework's pluggable architecture for prompts and tools.
Understanding the Fine-Tuning Architecture
The framework separates concerns into four extensible components that enable customization without modifying core library code. Each component exposes specific entry points for injecting your fine-tuned model and domain logic.
LLM Backend
The LLM backend generates responses and determines when to invoke tools. In qwen_agent/llm/__init__.py, the get_chat_model factory function (lines 31-55) instantiates concrete LLM classes based on the model and model_type configuration fields. By pointing the model parameter to your fine-tuned checkpoint and setting model_server to your custom endpoint, the agent automatically routes all inference requests to your specialized model.
Agent Core
Located in qwen_agent/agent.py, the Agent class (specifically Agent.__init__ at lines 38-45) orchestrates message handling, tool detection, and response streaming. This class accepts a system_message parameter at instantiation, allowing you to inject task-specific instructions that guide the fine-tuned model's behavior without further weight updates.
Tool Registration
The _init_tool method (lines 62-70 in qwen_agent/agent.py) dynamically registers capabilities from the function_list. You can inject domain-specific tools—such as medical database queries or proprietary API calls—by passing custom tool instances or MCP configurations during agent construction.
Memory System
Implemented in qwen_agent/memory/memory.py, the memory component persists long-term context through document storage and embeddings. Fine-tuned agents can integrate retrieval-augmented generation (RAG) pipelines by swapping the default in-memory store for vector database backends, keeping domain knowledge accessible to your specialized model.
Step 1: Fine-Tune the Underlying Qwen Model
Qwen-Agent delegates model training to external frameworks but consumes any OpenAI-compatible or DashScope-compatible endpoint. Choose the fine-tuning approach that matches your computational resources and domain requirements.
LoRA / QLoRA (Parameter-Efficient)
Train a LoRA adapter on your domain data using the official Qwen-2 or Qwen-3 training codebase. Export the adapter weights and serve them via vLLM with the --lora-modules-path flag. Configure Qwen-Agent to point at this endpoint.
Full-Parameter Fine-Tuning
Use the HuggingFace Trainer or similar frameworks to fine-tune Qwen-3-VL on your entire dataset. Serve the resulting checkpoint through an OpenAI-compatible inference server (such as vLLM) and note the endpoint URL.
In-Prompt Fine-Tuning (Prefix Tuning)
For rapid prototyping without GPU training, craft a detailed task-specific prefix in the system prompt. This approach requires no model weight changes and relies entirely on prompt engineering within the agent configuration.
Step 2: Configure Qwen-Agent to Use Your Fine-Tuned Model
Once your model is accessible via a REST endpoint, configure the agent to route requests to it. The get_chat_model factory (lines 59-81 in qwen_agent/llm/__init__.py) reads the model_type field and instantiates the appropriate client class, such as TextChatAtOAI for OpenAI-compatible servers.
llm_cfg = {
"model": "my-finetuned-qwen3-vl", # Your checkpoint name
"model_type": "oai", # OpenAI-compatible API
"model_server": "http://localhost:8000/v1", # vLLM or custom endpoint
"api_key": "EMPTY", # Required for compatibility
"generate_cfg": {
"top_p": 0.9,
"max_new_tokens": 1024
},
}
Pass this configuration dictionary to any agent constructor. The framework handles authentication, request formatting, and response streaming automatically.
Step 3: Customize the System Prompt for Task Specialization
Fine-tuned models often require explicit task framing to activate specialized knowledge. Pass a custom system_message during Agent instantiation to encode domain constraints, output formats, and stylistic requirements.
system_prompt = """
You are a legal-assistant specialized in contract analysis.
- Always cite the relevant clause number.
- If the user asks for a summary, produce a concise bullet-point list.
- When unsure, ask clarifying questions before answering.
"""
bot = Assistant(
llm=llm_cfg,
system_message=system_prompt,
function_list=["keyword_search", "code_interpreter"],
)
The system prompt is concatenated with user messages during the _run cycle, ensuring your fine-tuned model receives consistent task instructions across all interactions.
Step 4: Add Domain-Specific Tools
Extend your agent's capabilities beyond the base toolset by implementing custom functionality. Create a subclass of BaseTool from qwen_agent/tools/base.py, decorate it with @register_tool, and implement the call method.
# my_tools.py
from qwen_agent.tools.base import BaseTool, register_tool
import json5
import requests
@register_tool('medical_lookup')
class MedicalLookupTool(BaseTool):
description = "Lookup medical terminology and drug interactions."
parameters = [
{
"name": "term",
"type": "string",
"description": "Medical term to look up",
"required": True
}
]
def call(self, params: str, **kwargs) -> str:
arguments = json5.loads(params)
term = arguments["term"]
response = requests.get(f"https://medical.api/lookup?term={term}")
return response.text
Instantiate the agent with your custom tool:
from my_tools import MedicalLookupTool
bot = Assistant(
llm=llm_cfg,
system_message=system_prompt,
function_list=[MedicalLookupTool()], # Pass tool instance directly
)
The _init_tool method (lines 62-70 in qwen_agent/agent.py) automatically registers your tool, making it available for the LLM to invoke via function calling.
End-to-End Implementation Example
The following script demonstrates a complete pipeline integrating a fine-tuned model, custom system prompt, and domain-specific tool.
# fine_tune_demo.py
import json5
from qwen_agent.agents import Assistant
from qwen_agent.llm import get_chat_model
from my_tools import MedicalLookupTool
# 1. Configure the fine-tuned LLM endpoint
llm_cfg = {
"model": "my-finetuned-qwen3-vl",
"model_type": "oai",
"model_server": "http://localhost:8000/v1",
"api_key": "EMPTY",
"generate_cfg": {"temperature": 0.2, "top_p": 0.95},
}
llm = get_chat_model(llm_cfg)
# 2. Define task-specific behavior
system_prompt = """
You are a medical-assistant. Answer questions with citations from trusted sources.
If a question is ambiguous, ask for clarification first.
"""
# 3. Initialize agent with custom tool
bot = Assistant(
llm=llm,
system_message=system_prompt,
function_list=[MedicalLookupTool()],
)
# 4. Interactive execution loop
messages = []
while True:
user_input = input("\nUser: ")
if not user_input:
break
messages.append({"role": "user", "content": user_input})
for reply in bot.run(messages=messages):
for msg in reply:
print(msg.content, end="", flush=True)
print()
This implementation routes all LLM calls to your fine-tuned checkpoint, enforces the medical-assistant persona through the system prompt, and allows the model to invoke medical_lookup for factual retrieval.
Key Source Files for Customization
Understanding these core files enables effective debugging and extension of your fine-tuned agent:
qwen_agent/agent.py: Contains theAgentclass with__init__(lines 38-45) for system prompt handling and_init_tool(lines 62-70) for tool registration.qwen_agent/llm/__init__.py: Housesget_chat_model(lines 31-55, 59-81) which maps configuration dictionaries to concrete LLM implementations.qwen_agent/tools/base.py: DefinesBaseTooland theregister_tooldecorator for creating custom capabilities.examples/assistant_qwen3.py: Reference implementation showing standard agent configuration patterns.
Summary
- Fine-tune separately: Train your Qwen model using LoRA, QLoRA, or full-parameter methods, then serve via an OpenAI-compatible endpoint.
- Point and configure: Use
get_chat_modelinqwen_agent/llm/__init__.pyto connect to your fine-tuned checkpoint through themodel_serverparameter. - Specialize via prompts: Inject domain expertise through the
system_messageparameter inAgent.__init__without retraining. - Extend capabilities: Register custom tools by subclassing
BaseTooland passing instances to thefunction_listparameter. - Leverage memory: Integrate vector stores through the memory component for retrieval-augmented generation with domain knowledge bases.
Frequently Asked Questions
Does Qwen-Agent include a built-in fine-tuning pipeline?
No. According to the QwenLM/Qwen-Agent source code, the framework is designed as an orchestration layer that consumes fine-tuned models via API endpoints. You must train your Qwen model using external frameworks like HuggingFace Transformers or the official Qwen training scripts, then serve the checkpoint through vLLM or similar OpenAI-compatible servers.
How do I connect a LoRA adapter to my Qwen-Agent instance?
Serve your LoRA adapter through an inference engine that supports dynamic adapter loading, such as vLLM with the --lora-modules-path flag. Then configure Qwen-Agent's llm_cfg with the model_server pointing to that endpoint and model_type set to "oai". The get_chat_model factory will route requests to your adapted weights automatically.
Can I use fine-tuned models with custom tools simultaneously?
Yes. The Agent class in qwen_agent/agent.py treats the LLM backend and tool registry as independent components. You can pass a fine-tuned model configuration to the llm parameter and a list of custom BaseTool subclasses to function_list in the same constructor call. The _init_tool method registers your tools while the LLM client handles inference from your specialized checkpoint.
What is the difference between fine-tuning the model and fine-tuning the agent?
Fine-tuning the model involves updating the neural network weights (or adapter layers) of the underlying Qwen checkpoint through gradient descent on domain-specific data. Fine-tuning the agent involves configuring the orchestration layer—modifying system prompts, adding custom tools in _init_tool, or adjusting memory settings—without changing model weights. For optimal performance, combine both approaches: train the model on domain data and specialize the agent's behavior through prompts and tools.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →