# How to Integrate Qwen-Agent with Other LLMs: A Complete Plug-in Guide

> Learn how to integrate Qwen-Agent with other LLMs. Follow this plug-in guide to subclass BaseFnCallModel implement essential methods and register your LLM for seamless integration.

- Repository: [Qwen/Qwen-Agent](https://github.com/qwenlm/Qwen-Agent)
- Tags: how-to-guide
- Published: 2026-03-09

---

**You can integrate any LLM into Qwen-Agent by subclassing `BaseFnCallModel` from [`qwen_agent/llm/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/base.py), implementing the `_chat_stream` and `_chat_no_stream` methods, and registering your class with the `@register_llm` decorator to make it discoverable via the `model_type` configuration key.**

Qwen-Agent provides a flexible plug-in architecture that enables connections to proprietary APIs, local models, or custom endpoints without modifying the core framework. The system centers on a central registry pattern that maps configuration strings to concrete model implementations, allowing you to integrate everything from Anthropic Claude to private Hugging Face deployments using a consistent subclass-and-register workflow.

## Understanding the LLM Plug-in Architecture

The integration mechanism relies on three core components in [`qwen_agent/llm/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/base.py):

- **`BaseFnCallModel`** – The abstract base class that all LLM implementations must extend. It handles common utilities like retry logic, caching, token-budget management, and multimodal preprocessing.
- **`LLM_REGISTRY`** – A global dictionary that maps string identifiers (e.g., `"anthropic"`, `"my_custom_llm"`) to concrete classes.
- **`@register_llm`** – A decorator that adds your class to `LLM_REGISTRY` at import time.

When you instantiate an `Assistant` (or any other agent), you pass a dictionary containing a `model_type` key. The agent queries `LLM_REGISTRY[model_type]`, instantiates the corresponding class with your configuration, and calls its streaming or non-streaming methods to generate responses.

## Step-by-Step Integration Process

### Step 1: Create a Subclass of BaseFnCallModel

Create a new Python file in `qwen_agent/llm/` (e.g., [`qwen_agent/llm/my_llm.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/my_llm.py)) and define a class that inherits from `BaseFnCallModel`. Your `__init__` method should accept a configuration dictionary and extract parameters like API keys, model identifiers, and generation settings.

```python
from typing import Dict, Optional
from qwen_agent.llm.base import BaseFnCallModel, register_llm

@register_llm('my_llm')
class MyCustomLLM(BaseFnCallModel):
    def __init__(self, cfg: Optional[Dict] = None):
        super().__init__(cfg)
        self.api_key = cfg.get('api_key') or os.getenv('MY_LLM_API_KEY')
        self.model = cfg.get('model', 'default-model-id')

```

### Step 2: Implement Required Chat Methods

You must implement two abstract methods that handle the actual API communication:

- **`_chat_stream`** – Returns an `Iterator[List[Message]]` for streaming responses. Process the API's streaming chunks and yield partial messages.
- **`_chat_no_stream`** – Returns a `List[Message]` for single-shot responses. Wait for the complete response and return it as a list containing one `Message`.

Both methods receive a `List[Message]` (the conversation history) and a `generate_cfg` dictionary containing parameters like `temperature` and `max_new_tokens`.

### Step 3: Register Your Implementation

Apply the `@register_llm` decorator with a unique string identifier. This identifier becomes the value you use in the `model_type` field of your agent configuration. The registration happens automatically when Python imports the module.

```python
from qwen_agent.llm.base import register_llm

@register_llm('anthropic')
class AnthropicLLM(BaseFnCallModel):
    # Implementation details...

    pass

```

### Step 4: Configure Your Agent

Reference your registered `model_type` when creating an agent instance. Pass any model-specific parameters (API keys, model IDs) in the same configuration dictionary.

```python
from qwen_agent.agent import Assistant

cfg = {
    'model_type': 'my_llm',
    'model': 'my-model-id',
    'api_key': 'YOUR_API_KEY',
    'generate_cfg': {'temperature': 0.7, 'max_new_tokens': 1024}
}

bot = Assistant(llm=cfg)
response = bot.chat([{'role': 'user', 'content': 'Hello!'}])

```

## Complete Working Example: Anthropic Claude Integration

Below is a full implementation that wraps Anthropic's Claude API, demonstrating how to convert between Qwen-Agent's `Message` format and Anthropic's request/response format.

```python

# File: qwen_agent/llm/anthropic.py

import json
import os
from typing import Dict, Iterator, List, Optional

import requests
from qwen_agent.llm.base import register_llm
from qwen_agent.llm.function_calling import BaseFnCallModel
from qwen_agent.llm.schema import ASSISTANT, Message


@register_llm('anthropic')
class AnthropicLLM(BaseFnCallModel):
    """Thin wrapper around Anthropic's /v1/messages endpoint."""

    def __init__(self, cfg: Optional[Dict] = None):
        super().__init__(cfg)
        self.api_key = cfg.get('api_key') or os.getenv('ANTHROPIC_API_KEY')
        if not self.api_key:
            raise ValueError('Anthropic API key not found.')
        self.model = cfg.get('model', 'claude-3-opus-20240229')

    def _call_api(self, messages: List[Message], stream: bool, generate_cfg: dict):
        # Convert Qwen-Agent messages to Anthropic format

        anthropic_msgs = [
            {'role': m.role, 'content': [{'type': 'text', 'text': m.content}]}
            for m in messages
        ]
        payload = {
            'model': self.model,
            'messages': anthropic_msgs,
            'max_tokens': generate_cfg.get('max_new_tokens', 1024),
            'temperature': generate_cfg.get('temperature', 0.7),
            'stream': stream,
        }
        headers = {'x-api-key': self.api_key, 'Content-Type': 'application/json'}
        return requests.post(
            'https://api.anthropic.com/v1/messages',
            json=payload,
            headers=headers,
            stream=stream
        )

    def _chat_stream(
        self,
        messages: List[Message],
        delta_stream: bool,
        generate_cfg: dict,
    ) -> Iterator[List[Message]]:
        # Anthropic only supports full-response streaming, so we ignore delta_stream

        resp = self._call_api(messages, stream=True, generate_cfg=generate_cfg)
        resp.raise_for_status()
        for line in resp.iter_lines():
            if not line:
                continue
            data = json.loads(line.decode())
            if 'delta' in data and 'text' in data['delta']:
                text = data['delta']['text']
                yield [Message(ASSISTANT, text)]

    def _chat_no_stream(
        self,
        messages: List[Message],
        generate_cfg: dict,
    ) -> List[Message]:
        resp = self._call_api(messages, stream=False, generate_cfg=generate_cfg)
        resp.raise_for_status()
        content = resp.json()['content'][0]['text']
        return [Message(ASSISTANT, content)]

```

**Using the wrapper:**

```python
from qwen_agent.agent import Assistant

llm_cfg = {
    'model_type': 'anthropic',
    'model': 'claude-3-sonnet-20240229',
    'api_key': 'sk-ant-...',  # Can also use ANTHROPIC_API_KEY env var

    'generate_cfg': {'temperature': 0.5, 'max_new_tokens': 512},
}

assistant = Assistant(llm=llm_cfg)
reply = assistant.chat([{'role': 'user', 'content': 'Explain quantum tunneling.'}])
print(reply[0].content)

```

## Supporting Advanced Features

### Function Calling Support

If your target model supports OpenAI-style function calling, override the `_chat_with_functions` method. This method receives a `functions` list and must parse `tool_calls` from the response. Reference the implementation in [`qwen_agent/llm/oai.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/oai.py) for the standard request format and response parsing pattern.

```python
def _chat_with_functions(
    self,
    messages: List[Message],
    functions: List[Dict],
    stream: bool,
    delta_stream: bool,
    generate_cfg: dict,
    lang: str = 'en'
):
    # Build request with 'tools' parameter

    # Parse response for 'tool_calls' field

    pass

```

### Multimodal Capabilities

To indicate support for images, audio, or video, override the boolean properties in your subclass:

- **`support_multimodal_input`** – Set to `True` if the model accepts images or audio.
- **`support_multimodal_output`** – Set to `True` if the model generates images or non-text content.
- **`support_audio_input`** – Set to `True` if the model accepts audio files.

You must also adapt the `_get_inputs` method to pack media files into the request payload, similar to the pattern used in [`qwen_agent/llm/transformers_llm.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/transformers_llm.py).

## Key Source Files for Reference

- **[`qwen_agent/llm/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/base.py)** – Defines `BaseFnCallModel`, the `@register_llm` decorator, and the central `LLM_REGISTRY`.
- **[`qwen_agent/llm/oai.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/oai.py)** – Reference implementation for OpenAI-compatible models, including function calling logic.
- **[`qwen_agent/llm/qwen_dashscope.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/qwen_dashscope.py)** – Example of vendor-specific SDK integration (DashScope) with full streaming and tool usage.
- **[`qwen_agent/llm/transformers_llm.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/transformers_llm.py)** – Demonstrates wrapping Hugging Face Transformers models, including multimodal processor handling.
- **[`qwen_agent/agent.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/agent.py)** – Contains the high-level `Assistant` and `Agent` classes that consume the LLM registry.

## Summary

- **Subclass `BaseFnCallModel`** from [`qwen_agent/llm/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/base.py) to create a new LLM wrapper.
- **Implement `_chat_stream` and `_chat_no_stream`** to handle streaming and single-shot generation.
- **Register with `@register_llm('identifier')`** to add your class to the global `LLM_REGISTRY`.
- **Use `model_type`** in your agent configuration dictionary to select the registered implementation.
- **Optionally implement** `_chat_with_functions` and multimodal properties for extended capabilities.

## Frequently Asked Questions

### Do I need to modify Qwen-Agent's core code to add a new LLM?

No. You can place your new subclass in any file within the `qwen_agent/llm/` directory or import it in your application code. The `@register_llm` decorator handles registration automatically via import side-effects, making your implementation discoverable without touching the framework's internals.

### Can I integrate local models that don't use an API?

Yes. The [`qwen_agent/llm/transformers_llm.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/transformers_llm.py) file provides a complete example of loading and running Hugging Face Transformers models locally. You simply load the model and tokenizer in `__init__`, then call them within `_chat_stream` and `_chat_no_stream`.

### How do I handle models that don't support streaming?

Implement `_chat_no_stream` to return the complete response as a list containing one `Message`. For `_chat_stream`, you can either yield the full response once or raise a `NotImplementedError`. The base class provides fallback logic, but yielding the complete content ensures compatibility with agents that expect streaming interfaces.

### What is the role of the `delta_stream` parameter in `_chat_stream`?

The `delta_stream` boolean indicates whether the caller wants incremental token deltas (like OpenAI's streaming format) or full messages at each step. If your backend only supports full message streaming (like Anthropic), you can ignore this parameter and yield complete `Message` objects at each iteration.