How to Integrate Qwen-Agent with Other LLMs: A Complete Plug-in Guide
You can integrate any LLM into Qwen-Agent by subclassing BaseFnCallModel from qwen_agent/llm/base.py, implementing the _chat_stream and _chat_no_stream methods, and registering your class with the @register_llm decorator to make it discoverable via the model_type configuration key.
Qwen-Agent provides a flexible plug-in architecture that enables connections to proprietary APIs, local models, or custom endpoints without modifying the core framework. The system centers on a central registry pattern that maps configuration strings to concrete model implementations, allowing you to integrate everything from Anthropic Claude to private Hugging Face deployments using a consistent subclass-and-register workflow.
Understanding the LLM Plug-in Architecture
The integration mechanism relies on three core components in qwen_agent/llm/base.py:
BaseFnCallModel– The abstract base class that all LLM implementations must extend. It handles common utilities like retry logic, caching, token-budget management, and multimodal preprocessing.LLM_REGISTRY– A global dictionary that maps string identifiers (e.g.,"anthropic","my_custom_llm") to concrete classes.@register_llm– A decorator that adds your class toLLM_REGISTRYat import time.
When you instantiate an Assistant (or any other agent), you pass a dictionary containing a model_type key. The agent queries LLM_REGISTRY[model_type], instantiates the corresponding class with your configuration, and calls its streaming or non-streaming methods to generate responses.
Step-by-Step Integration Process
Step 1: Create a Subclass of BaseFnCallModel
Create a new Python file in qwen_agent/llm/ (e.g., qwen_agent/llm/my_llm.py) and define a class that inherits from BaseFnCallModel. Your __init__ method should accept a configuration dictionary and extract parameters like API keys, model identifiers, and generation settings.
from typing import Dict, Optional
from qwen_agent.llm.base import BaseFnCallModel, register_llm
@register_llm('my_llm')
class MyCustomLLM(BaseFnCallModel):
def __init__(self, cfg: Optional[Dict] = None):
super().__init__(cfg)
self.api_key = cfg.get('api_key') or os.getenv('MY_LLM_API_KEY')
self.model = cfg.get('model', 'default-model-id')
Step 2: Implement Required Chat Methods
You must implement two abstract methods that handle the actual API communication:
_chat_stream– Returns anIterator[List[Message]]for streaming responses. Process the API's streaming chunks and yield partial messages._chat_no_stream– Returns aList[Message]for single-shot responses. Wait for the complete response and return it as a list containing oneMessage.
Both methods receive a List[Message] (the conversation history) and a generate_cfg dictionary containing parameters like temperature and max_new_tokens.
Step 3: Register Your Implementation
Apply the @register_llm decorator with a unique string identifier. This identifier becomes the value you use in the model_type field of your agent configuration. The registration happens automatically when Python imports the module.
from qwen_agent.llm.base import register_llm
@register_llm('anthropic')
class AnthropicLLM(BaseFnCallModel):
# Implementation details...
pass
Step 4: Configure Your Agent
Reference your registered model_type when creating an agent instance. Pass any model-specific parameters (API keys, model IDs) in the same configuration dictionary.
from qwen_agent.agent import Assistant
cfg = {
'model_type': 'my_llm',
'model': 'my-model-id',
'api_key': 'YOUR_API_KEY',
'generate_cfg': {'temperature': 0.7, 'max_new_tokens': 1024}
}
bot = Assistant(llm=cfg)
response = bot.chat([{'role': 'user', 'content': 'Hello!'}])
Complete Working Example: Anthropic Claude Integration
Below is a full implementation that wraps Anthropic's Claude API, demonstrating how to convert between Qwen-Agent's Message format and Anthropic's request/response format.
# File: qwen_agent/llm/anthropic.py
import json
import os
from typing import Dict, Iterator, List, Optional
import requests
from qwen_agent.llm.base import register_llm
from qwen_agent.llm.function_calling import BaseFnCallModel
from qwen_agent.llm.schema import ASSISTANT, Message
@register_llm('anthropic')
class AnthropicLLM(BaseFnCallModel):
"""Thin wrapper around Anthropic's /v1/messages endpoint."""
def __init__(self, cfg: Optional[Dict] = None):
super().__init__(cfg)
self.api_key = cfg.get('api_key') or os.getenv('ANTHROPIC_API_KEY')
if not self.api_key:
raise ValueError('Anthropic API key not found.')
self.model = cfg.get('model', 'claude-3-opus-20240229')
def _call_api(self, messages: List[Message], stream: bool, generate_cfg: dict):
# Convert Qwen-Agent messages to Anthropic format
anthropic_msgs = [
{'role': m.role, 'content': [{'type': 'text', 'text': m.content}]}
for m in messages
]
payload = {
'model': self.model,
'messages': anthropic_msgs,
'max_tokens': generate_cfg.get('max_new_tokens', 1024),
'temperature': generate_cfg.get('temperature', 0.7),
'stream': stream,
}
headers = {'x-api-key': self.api_key, 'Content-Type': 'application/json'}
return requests.post(
'https://api.anthropic.com/v1/messages',
json=payload,
headers=headers,
stream=stream
)
def _chat_stream(
self,
messages: List[Message],
delta_stream: bool,
generate_cfg: dict,
) -> Iterator[List[Message]]:
# Anthropic only supports full-response streaming, so we ignore delta_stream
resp = self._call_api(messages, stream=True, generate_cfg=generate_cfg)
resp.raise_for_status()
for line in resp.iter_lines():
if not line:
continue
data = json.loads(line.decode())
if 'delta' in data and 'text' in data['delta']:
text = data['delta']['text']
yield [Message(ASSISTANT, text)]
def _chat_no_stream(
self,
messages: List[Message],
generate_cfg: dict,
) -> List[Message]:
resp = self._call_api(messages, stream=False, generate_cfg=generate_cfg)
resp.raise_for_status()
content = resp.json()['content'][0]['text']
return [Message(ASSISTANT, content)]
Using the wrapper:
from qwen_agent.agent import Assistant
llm_cfg = {
'model_type': 'anthropic',
'model': 'claude-3-sonnet-20240229',
'api_key': 'sk-ant-...', # Can also use ANTHROPIC_API_KEY env var
'generate_cfg': {'temperature': 0.5, 'max_new_tokens': 512},
}
assistant = Assistant(llm=llm_cfg)
reply = assistant.chat([{'role': 'user', 'content': 'Explain quantum tunneling.'}])
print(reply[0].content)
Supporting Advanced Features
Function Calling Support
If your target model supports OpenAI-style function calling, override the _chat_with_functions method. This method receives a functions list and must parse tool_calls from the response. Reference the implementation in qwen_agent/llm/oai.py for the standard request format and response parsing pattern.
def _chat_with_functions(
self,
messages: List[Message],
functions: List[Dict],
stream: bool,
delta_stream: bool,
generate_cfg: dict,
lang: str = 'en'
):
# Build request with 'tools' parameter
# Parse response for 'tool_calls' field
pass
Multimodal Capabilities
To indicate support for images, audio, or video, override the boolean properties in your subclass:
support_multimodal_input– Set toTrueif the model accepts images or audio.support_multimodal_output– Set toTrueif the model generates images or non-text content.support_audio_input– Set toTrueif the model accepts audio files.
You must also adapt the _get_inputs method to pack media files into the request payload, similar to the pattern used in qwen_agent/llm/transformers_llm.py.
Key Source Files for Reference
qwen_agent/llm/base.py– DefinesBaseFnCallModel, the@register_llmdecorator, and the centralLLM_REGISTRY.qwen_agent/llm/oai.py– Reference implementation for OpenAI-compatible models, including function calling logic.qwen_agent/llm/qwen_dashscope.py– Example of vendor-specific SDK integration (DashScope) with full streaming and tool usage.qwen_agent/llm/transformers_llm.py– Demonstrates wrapping Hugging Face Transformers models, including multimodal processor handling.qwen_agent/agent.py– Contains the high-levelAssistantandAgentclasses that consume the LLM registry.
Summary
- Subclass
BaseFnCallModelfromqwen_agent/llm/base.pyto create a new LLM wrapper. - Implement
_chat_streamand_chat_no_streamto handle streaming and single-shot generation. - Register with
@register_llm('identifier')to add your class to the globalLLM_REGISTRY. - Use
model_typein your agent configuration dictionary to select the registered implementation. - Optionally implement
_chat_with_functionsand multimodal properties for extended capabilities.
Frequently Asked Questions
Do I need to modify Qwen-Agent's core code to add a new LLM?
No. You can place your new subclass in any file within the qwen_agent/llm/ directory or import it in your application code. The @register_llm decorator handles registration automatically via import side-effects, making your implementation discoverable without touching the framework's internals.
Can I integrate local models that don't use an API?
Yes. The qwen_agent/llm/transformers_llm.py file provides a complete example of loading and running Hugging Face Transformers models locally. You simply load the model and tokenizer in __init__, then call them within _chat_stream and _chat_no_stream.
How do I handle models that don't support streaming?
Implement _chat_no_stream to return the complete response as a list containing one Message. For _chat_stream, you can either yield the full response once or raise a NotImplementedError. The base class provides fallback logic, but yielding the complete content ensures compatibility with agents that expect streaming interfaces.
What is the role of the delta_stream parameter in _chat_stream?
The delta_stream boolean indicates whether the caller wants incremental token deltas (like OpenAI's streaming format) or full messages at each step. If your backend only supports full message streaming (like Anthropic), you can ignore this parameter and yield complete Message objects at each iteration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →