# How to Fine-Tune Qwen-Agent for Specific Tasks: A Complete Technical Guide

> Learn how to fine-tune Qwen-Agent for your specific tasks. This guide details serving custom models via OpenAI endpoints, configuring parameters, and using custom prompts & tools for specialized agent behavior.

- Repository: [Qwen/Qwen-Agent](https://github.com/qwenlm/Qwen-Agent)
- Tags: how-to-guide
- Published: 2026-03-09

---

**Fine-tune Qwen-Agent by serving your customized Qwen model via an OpenAI-compatible endpoint, configuring the `model_server` parameter in `get_chat_model`, and specializing agent behavior through custom system prompts and domain-specific tools.**

Qwen-Agent is a modular framework developed by QwenLM that orchestrates large language models, external tools, and memory systems to build task-specific AI agents. While the core library does not ship with a built-in training loop, you can adapt Qwen-Agent to specialized domains by integrating a fine-tuned Qwen checkpoint and leveraging the framework's pluggable architecture for prompts and tools.

## Understanding the Fine-Tuning Architecture

The framework separates concerns into four extensible components that enable customization without modifying core library code. Each component exposes specific entry points for injecting your fine-tuned model and domain logic.

### LLM Backend

The LLM backend generates responses and determines when to invoke tools. In [`qwen_agent/llm/__init__.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/__init__.py), the `get_chat_model` factory function (lines 31-55) instantiates concrete LLM classes based on the `model` and `model_type` configuration fields. By pointing the `model` parameter to your fine-tuned checkpoint and setting `model_server` to your custom endpoint, the agent automatically routes all inference requests to your specialized model.

### Agent Core

Located in [`qwen_agent/agent.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py), the `Agent` class (specifically `Agent.__init__` at lines 38-45) orchestrates message handling, tool detection, and response streaming. This class accepts a `system_message` parameter at instantiation, allowing you to inject task-specific instructions that guide the fine-tuned model's behavior without further weight updates.

### Tool Registration

The `_init_tool` method (lines 62-70 in [`qwen_agent/agent.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py)) dynamically registers capabilities from the `function_list`. You can inject domain-specific tools—such as medical database queries or proprietary API calls—by passing custom tool instances or MCP configurations during agent construction.

### Memory System

Implemented in [`qwen_agent/memory/memory.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/memory/memory.py), the memory component persists long-term context through document storage and embeddings. Fine-tuned agents can integrate retrieval-augmented generation (RAG) pipelines by swapping the default in-memory store for vector database backends, keeping domain knowledge accessible to your specialized model.

## Step 1: Fine-Tune the Underlying Qwen Model

Qwen-Agent delegates model training to external frameworks but consumes any OpenAI-compatible or DashScope-compatible endpoint. Choose the fine-tuning approach that matches your computational resources and domain requirements.

**LoRA / QLoRA (Parameter-Efficient)**

Train a LoRA adapter on your domain data using the official Qwen-2 or Qwen-3 training codebase. Export the adapter weights and serve them via vLLM with the `--lora-modules-path` flag. Configure Qwen-Agent to point at this endpoint.

**Full-Parameter Fine-Tuning**

Use the HuggingFace `Trainer` or similar frameworks to fine-tune Qwen-3-VL on your entire dataset. Serve the resulting checkpoint through an OpenAI-compatible inference server (such as vLLM) and note the endpoint URL.

**In-Prompt Fine-Tuning (Prefix Tuning)**

For rapid prototyping without GPU training, craft a detailed task-specific prefix in the system prompt. This approach requires no model weight changes and relies entirely on prompt engineering within the agent configuration.

## Step 2: Configure Qwen-Agent to Use Your Fine-Tuned Model

Once your model is accessible via a REST endpoint, configure the agent to route requests to it. The `get_chat_model` factory (lines 59-81 in [`qwen_agent/llm/__init__.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/__init__.py)) reads the `model_type` field and instantiates the appropriate client class, such as `TextChatAtOAI` for OpenAI-compatible servers.

```python
llm_cfg = {
    "model": "my-finetuned-qwen3-vl",          # Your checkpoint name

    "model_type": "oai",                       # OpenAI-compatible API

    "model_server": "http://localhost:8000/v1", # vLLM or custom endpoint

    "api_key": "EMPTY",                        # Required for compatibility

    "generate_cfg": {
        "top_p": 0.9,
        "max_new_tokens": 1024
    },
}

```

Pass this configuration dictionary to any agent constructor. The framework handles authentication, request formatting, and response streaming automatically.

## Step 3: Customize the System Prompt for Task Specialization

Fine-tuned models often require explicit task framing to activate specialized knowledge. Pass a custom `system_message` during `Agent` instantiation to encode domain constraints, output formats, and stylistic requirements.

```python
system_prompt = """
You are a legal-assistant specialized in contract analysis.
- Always cite the relevant clause number.
- If the user asks for a summary, produce a concise bullet-point list.
- When unsure, ask clarifying questions before answering.
"""

bot = Assistant(
    llm=llm_cfg,
    system_message=system_prompt,
    function_list=["keyword_search", "code_interpreter"],
)

```

The system prompt is concatenated with user messages during the `_run` cycle, ensuring your fine-tuned model receives consistent task instructions across all interactions.

## Step 4: Add Domain-Specific Tools

Extend your agent's capabilities beyond the base toolset by implementing custom functionality. Create a subclass of `BaseTool` from [`qwen_agent/tools/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/tools/base.py), decorate it with `@register_tool`, and implement the `call` method.

```python

# my_tools.py

from qwen_agent.tools.base import BaseTool, register_tool
import json5
import requests

@register_tool('medical_lookup')
class MedicalLookupTool(BaseTool):
    description = "Lookup medical terminology and drug interactions."
    parameters = [
        {
            "name": "term",
            "type": "string",
            "description": "Medical term to look up",
            "required": True
        }
    ]

    def call(self, params: str, **kwargs) -> str:
        arguments = json5.loads(params)
        term = arguments["term"]
        response = requests.get(f"https://medical.api/lookup?term={term}")
        return response.text

```

Instantiate the agent with your custom tool:

```python
from my_tools import MedicalLookupTool

bot = Assistant(
    llm=llm_cfg,
    system_message=system_prompt,
    function_list=[MedicalLookupTool()],  # Pass tool instance directly

)

```

The `_init_tool` method (lines 62-70 in [`qwen_agent/agent.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py)) automatically registers your tool, making it available for the LLM to invoke via function calling.

## End-to-End Implementation Example

The following script demonstrates a complete pipeline integrating a fine-tuned model, custom system prompt, and domain-specific tool.

```python

# fine_tune_demo.py

import json5
from qwen_agent.agents import Assistant
from qwen_agent.llm import get_chat_model
from my_tools import MedicalLookupTool

# 1. Configure the fine-tuned LLM endpoint

llm_cfg = {
    "model": "my-finetuned-qwen3-vl",
    "model_type": "oai",
    "model_server": "http://localhost:8000/v1",
    "api_key": "EMPTY",
    "generate_cfg": {"temperature": 0.2, "top_p": 0.95},
}
llm = get_chat_model(llm_cfg)

# 2. Define task-specific behavior

system_prompt = """
You are a medical-assistant. Answer questions with citations from trusted sources.
If a question is ambiguous, ask for clarification first.
"""

# 3. Initialize agent with custom tool

bot = Assistant(
    llm=llm,
    system_message=system_prompt,
    function_list=[MedicalLookupTool()],
)

# 4. Interactive execution loop

messages = []
while True:
    user_input = input("\nUser: ")
    if not user_input:
        break
    messages.append({"role": "user", "content": user_input})
    
    for reply in bot.run(messages=messages):
        for msg in reply:
            print(msg.content, end="", flush=True)
    print()

```

This implementation routes all LLM calls to your fine-tuned checkpoint, enforces the medical-assistant persona through the system prompt, and allows the model to invoke `medical_lookup` for factual retrieval.

## Key Source Files for Customization

Understanding these core files enables effective debugging and extension of your fine-tuned agent:

- **[`qwen_agent/agent.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py)**: Contains the `Agent` class with `__init__` (lines 38-45) for system prompt handling and `_init_tool` (lines 62-70) for tool registration.
- **[`qwen_agent/llm/__init__.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/__init__.py)**: Houses `get_chat_model` (lines 31-55, 59-81) which maps configuration dictionaries to concrete LLM implementations.
- **[`qwen_agent/tools/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/tools/base.py)**: Defines `BaseTool` and the `register_tool` decorator for creating custom capabilities.
- **[`examples/assistant_qwen3.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/examples/assistant_qwen3.py)**: Reference implementation showing standard agent configuration patterns.

## Summary

- **Fine-tune separately**: Train your Qwen model using LoRA, QLoRA, or full-parameter methods, then serve via an OpenAI-compatible endpoint.
- **Point and configure**: Use `get_chat_model` in [`qwen_agent/llm/__init__.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/__init__.py) to connect to your fine-tuned checkpoint through the `model_server` parameter.
- **Specialize via prompts**: Inject domain expertise through the `system_message` parameter in `Agent.__init__` without retraining.
- **Extend capabilities**: Register custom tools by subclassing `BaseTool` and passing instances to the `function_list` parameter.
- **Leverage memory**: Integrate vector stores through the memory component for retrieval-augmented generation with domain knowledge bases.

## Frequently Asked Questions

### Does Qwen-Agent include a built-in fine-tuning pipeline?

No. According to the QwenLM/Qwen-Agent source code, the framework is designed as an orchestration layer that consumes fine-tuned models via API endpoints. You must train your Qwen model using external frameworks like HuggingFace Transformers or the official Qwen training scripts, then serve the checkpoint through vLLM or similar OpenAI-compatible servers.

### How do I connect a LoRA adapter to my Qwen-Agent instance?

Serve your LoRA adapter through an inference engine that supports dynamic adapter loading, such as vLLM with the `--lora-modules-path` flag. Then configure Qwen-Agent's `llm_cfg` with the `model_server` pointing to that endpoint and `model_type` set to `"oai"`. The `get_chat_model` factory will route requests to your adapted weights automatically.

### Can I use fine-tuned models with custom tools simultaneously?

Yes. The `Agent` class in [`qwen_agent/agent.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py) treats the LLM backend and tool registry as independent components. You can pass a fine-tuned model configuration to the `llm` parameter and a list of custom `BaseTool` subclasses to `function_list` in the same constructor call. The `_init_tool` method registers your tools while the LLM client handles inference from your specialized checkpoint.

### What is the difference between fine-tuning the model and fine-tuning the agent?

Fine-tuning the model involves updating the neural network weights (or adapter layers) of the underlying Qwen checkpoint through gradient descent on domain-specific data. Fine-tuning the agent involves configuring the orchestration layer—modifying system prompts, adding custom tools in `_init_tool`, or adjusting memory settings—without changing model weights. For optimal performance, combine both approaches: train the model on domain data and specialize the agent's behavior through prompts and tools.