How to Integrate Qwen-Agent with Other LLMs: A Complete Plug-in Guide

You can integrate any LLM into Qwen-Agent by subclassing BaseFnCallModel from qwen_agent/llm/base.py, implementing the _chat_stream and _chat_no_stream methods, and registering your class with the @register_llm decorator to make it discoverable via the model_type configuration key.

Qwen-Agent provides a flexible plug-in architecture that enables connections to proprietary APIs, local models, or custom endpoints without modifying the core framework. The system centers on a central registry pattern that maps configuration strings to concrete model implementations, allowing you to integrate everything from Anthropic Claude to private Hugging Face deployments using a consistent subclass-and-register workflow.

Understanding the LLM Plug-in Architecture

The integration mechanism relies on three core components in qwen_agent/llm/base.py:

  • BaseFnCallModel – The abstract base class that all LLM implementations must extend. It handles common utilities like retry logic, caching, token-budget management, and multimodal preprocessing.
  • LLM_REGISTRY – A global dictionary that maps string identifiers (e.g., "anthropic", "my_custom_llm") to concrete classes.
  • @register_llm – A decorator that adds your class to LLM_REGISTRY at import time.

When you instantiate an Assistant (or any other agent), you pass a dictionary containing a model_type key. The agent queries LLM_REGISTRY[model_type], instantiates the corresponding class with your configuration, and calls its streaming or non-streaming methods to generate responses.

Step-by-Step Integration Process

Step 1: Create a Subclass of BaseFnCallModel

Create a new Python file in qwen_agent/llm/ (e.g., qwen_agent/llm/my_llm.py) and define a class that inherits from BaseFnCallModel. Your __init__ method should accept a configuration dictionary and extract parameters like API keys, model identifiers, and generation settings.

from typing import Dict, Optional
from qwen_agent.llm.base import BaseFnCallModel, register_llm

@register_llm('my_llm')
class MyCustomLLM(BaseFnCallModel):
    def __init__(self, cfg: Optional[Dict] = None):
        super().__init__(cfg)
        self.api_key = cfg.get('api_key') or os.getenv('MY_LLM_API_KEY')
        self.model = cfg.get('model', 'default-model-id')

Step 2: Implement Required Chat Methods

You must implement two abstract methods that handle the actual API communication:

  • _chat_stream – Returns an Iterator[List[Message]] for streaming responses. Process the API's streaming chunks and yield partial messages.
  • _chat_no_stream – Returns a List[Message] for single-shot responses. Wait for the complete response and return it as a list containing one Message.

Both methods receive a List[Message] (the conversation history) and a generate_cfg dictionary containing parameters like temperature and max_new_tokens.

Step 3: Register Your Implementation

Apply the @register_llm decorator with a unique string identifier. This identifier becomes the value you use in the model_type field of your agent configuration. The registration happens automatically when Python imports the module.

from qwen_agent.llm.base import register_llm

@register_llm('anthropic')
class AnthropicLLM(BaseFnCallModel):
    # Implementation details...

    pass

Step 4: Configure Your Agent

Reference your registered model_type when creating an agent instance. Pass any model-specific parameters (API keys, model IDs) in the same configuration dictionary.

from qwen_agent.agent import Assistant

cfg = {
    'model_type': 'my_llm',
    'model': 'my-model-id',
    'api_key': 'YOUR_API_KEY',
    'generate_cfg': {'temperature': 0.7, 'max_new_tokens': 1024}
}

bot = Assistant(llm=cfg)
response = bot.chat([{'role': 'user', 'content': 'Hello!'}])

Complete Working Example: Anthropic Claude Integration

Below is a full implementation that wraps Anthropic's Claude API, demonstrating how to convert between Qwen-Agent's Message format and Anthropic's request/response format.


# File: qwen_agent/llm/anthropic.py

import json
import os
from typing import Dict, Iterator, List, Optional

import requests
from qwen_agent.llm.base import register_llm
from qwen_agent.llm.function_calling import BaseFnCallModel
from qwen_agent.llm.schema import ASSISTANT, Message


@register_llm('anthropic')
class AnthropicLLM(BaseFnCallModel):
    """Thin wrapper around Anthropic's /v1/messages endpoint."""

    def __init__(self, cfg: Optional[Dict] = None):
        super().__init__(cfg)
        self.api_key = cfg.get('api_key') or os.getenv('ANTHROPIC_API_KEY')
        if not self.api_key:
            raise ValueError('Anthropic API key not found.')
        self.model = cfg.get('model', 'claude-3-opus-20240229')

    def _call_api(self, messages: List[Message], stream: bool, generate_cfg: dict):
        # Convert Qwen-Agent messages to Anthropic format

        anthropic_msgs = [
            {'role': m.role, 'content': [{'type': 'text', 'text': m.content}]}
            for m in messages
        ]
        payload = {
            'model': self.model,
            'messages': anthropic_msgs,
            'max_tokens': generate_cfg.get('max_new_tokens', 1024),
            'temperature': generate_cfg.get('temperature', 0.7),
            'stream': stream,
        }
        headers = {'x-api-key': self.api_key, 'Content-Type': 'application/json'}
        return requests.post(
            'https://api.anthropic.com/v1/messages',
            json=payload,
            headers=headers,
            stream=stream
        )

    def _chat_stream(
        self,
        messages: List[Message],
        delta_stream: bool,
        generate_cfg: dict,
    ) -> Iterator[List[Message]]:
        # Anthropic only supports full-response streaming, so we ignore delta_stream

        resp = self._call_api(messages, stream=True, generate_cfg=generate_cfg)
        resp.raise_for_status()
        for line in resp.iter_lines():
            if not line:
                continue
            data = json.loads(line.decode())
            if 'delta' in data and 'text' in data['delta']:
                text = data['delta']['text']
                yield [Message(ASSISTANT, text)]

    def _chat_no_stream(
        self,
        messages: List[Message],
        generate_cfg: dict,
    ) -> List[Message]:
        resp = self._call_api(messages, stream=False, generate_cfg=generate_cfg)
        resp.raise_for_status()
        content = resp.json()['content'][0]['text']
        return [Message(ASSISTANT, content)]

Using the wrapper:

from qwen_agent.agent import Assistant

llm_cfg = {
    'model_type': 'anthropic',
    'model': 'claude-3-sonnet-20240229',
    'api_key': 'sk-ant-...',  # Can also use ANTHROPIC_API_KEY env var

    'generate_cfg': {'temperature': 0.5, 'max_new_tokens': 512},
}

assistant = Assistant(llm=llm_cfg)
reply = assistant.chat([{'role': 'user', 'content': 'Explain quantum tunneling.'}])
print(reply[0].content)

Supporting Advanced Features

Function Calling Support

If your target model supports OpenAI-style function calling, override the _chat_with_functions method. This method receives a functions list and must parse tool_calls from the response. Reference the implementation in qwen_agent/llm/oai.py for the standard request format and response parsing pattern.

def _chat_with_functions(
    self,
    messages: List[Message],
    functions: List[Dict],
    stream: bool,
    delta_stream: bool,
    generate_cfg: dict,
    lang: str = 'en'
):
    # Build request with 'tools' parameter

    # Parse response for 'tool_calls' field

    pass

Multimodal Capabilities

To indicate support for images, audio, or video, override the boolean properties in your subclass:

  • support_multimodal_input – Set to True if the model accepts images or audio.
  • support_multimodal_output – Set to True if the model generates images or non-text content.
  • support_audio_input – Set to True if the model accepts audio files.

You must also adapt the _get_inputs method to pack media files into the request payload, similar to the pattern used in qwen_agent/llm/transformers_llm.py.

Key Source Files for Reference

Summary

  • Subclass BaseFnCallModel from qwen_agent/llm/base.py to create a new LLM wrapper.
  • Implement _chat_stream and _chat_no_stream to handle streaming and single-shot generation.
  • Register with @register_llm('identifier') to add your class to the global LLM_REGISTRY.
  • Use model_type in your agent configuration dictionary to select the registered implementation.
  • Optionally implement _chat_with_functions and multimodal properties for extended capabilities.

Frequently Asked Questions

Do I need to modify Qwen-Agent's core code to add a new LLM?

No. You can place your new subclass in any file within the qwen_agent/llm/ directory or import it in your application code. The @register_llm decorator handles registration automatically via import side-effects, making your implementation discoverable without touching the framework's internals.

Can I integrate local models that don't use an API?

Yes. The qwen_agent/llm/transformers_llm.py file provides a complete example of loading and running Hugging Face Transformers models locally. You simply load the model and tokenizer in __init__, then call them within _chat_stream and _chat_no_stream.

How do I handle models that don't support streaming?

Implement _chat_no_stream to return the complete response as a list containing one Message. For _chat_stream, you can either yield the full response once or raise a NotImplementedError. The base class provides fallback logic, but yielding the complete content ensures compatibility with agents that expect streaming interfaces.

What is the role of the delta_stream parameter in _chat_stream?

The delta_stream boolean indicates whether the caller wants incremental token deltas (like OpenAI's streaming format) or full messages at each step. If your backend only supports full message streaming (like Anthropic), you can ignore this parameter and yield complete Message objects at each iteration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →