# How to Stream Outputs from Qwen-Agent: A Complete Implementation Guide

> Learn how to stream outputs from Qwen-Agent with this implementation guide. Discover how to use stream=True for partial LLM responses and enhance your applications.

- Repository: [Qwen/Qwen-Agent](https://github.com/qwenlm/Qwen-Agent)
- Tags: how-to-guide
- Published: 2026-03-09

---

**Pass `stream=True` to `BaseChatModel.chat()` and iterate over the returned generator to receive partial LLM responses as they are generated by the underlying service.**

Qwen-Agent is a framework for building LLM applications with advanced tool-use and multimodal capabilities. Streaming support is built directly into the `BaseChatModel` interface, enabling real-time token delivery without waiting for complete responses. By setting the `stream` parameter to `True`, developers can consume partial outputs through a unified iterator protocol that works across OpenAI, DashScope, and Qwen-VL backends.

## Understanding the Streaming Architecture

Streaming in Qwen-Agent follows a provider-agnostic iterator pattern defined in the base model class. The architecture separates the public API from backend-specific implementations, ensuring consistent behavior whether you use text-only or multimodal models.

### The Entry Point in base.py

The `BaseChatModel.chat` method in [`qwen_agent/llm/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/base.py) serves as the unified entry point for all LLM interactions. When you invoke `chat` with `stream=True`, the method checks this flag at lines 18-27 and forwards the request to the model-specific `_chat_stream` implementation rather than the standard `_chat_no_stream` method.

The method signature accepts two critical boolean parameters: `stream` to enable the iterator behavior, and `delta_stream` to control chunk format. According to the source code, when streaming is requested, the method returns a generator that yields lists of `Message` objects representing the conversation state at each token generation step.

### Provider-Specific Streaming Implementations

Each LLM backend implements its own `_chat_stream` generator that handles the actual API communication with `stream=True`:

- **OpenAI-compatible models**: The `TextChatAtOAI` class in [`qwen_agent/llm/oai.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/oai.py) implements `_chat_stream` at lines 98-108, invoking the provider's streaming endpoint and converting chunks into `Message` objects.
- **Qwen-VL multimodal models**: The implementation in [`qwen_agent/llm/qwenvl_dashscope.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/qwenvl_dashscope.py) (lines 64-78) handles mixed image-text content, yielding accumulated responses suitable for real-time display.
- **DashScope text models**: Implemented in [`qwen_agent/llm/qwen_dashscope.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/qwen_dashscope.py), providing streaming for Qwen-3 and other text-only models.

All implementations yield `List[Message]` objects, maintaining type consistency across different providers.

## How to Implement Streaming in Your Code

### Basic Text Streaming with OpenAI-Compatible Models

For OpenAI-compatible endpoints, initialize `TextChatAtOAI` and consume the iterator in a `for` loop. Each iteration yields the full accumulated response up to the current token:

```python
from qwen_agent.llm.oai import TextChatAtOAI
from qwen_agent.llm.schema import Message, USER

# Initialize with your API configuration

llm = TextChatAtOAI(cfg={"api_key": "YOUR_KEY"})

messages = [Message(role=USER, content="Tell me a short story about a cat.")]

# stream=True returns an iterator; delta_stream defaults to False

for partial in llm.chat(messages=messages, stream=True):
    # partial is a list containing a single Message with accumulated text

    print(partial[0].content, end="", flush=True)

```

The internal call reaches `_chat_stream` in [`oai.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/oai.py), which manages the HTTP connection and yields incremental updates as the provider sends them.

### Multimodal Streaming with Qwen-VL

Vision-language models use the identical streaming interface, though you must handle potential image references in the yielded content:

```python
from qwen_agent.llm.qwenvl_dashscope import QwenVL
from qwen_agent.llm.schema import Message, USER

llm = QwenVL(cfg={"model": "qwen-vl-plus", "model_server": "dashscope"})

messages = [
    Message(role=USER, content="Describe the image and answer the question."),
    Message(
        role=USER,
        content={"image": "https://example.com/scene.jpg", "text": "What is happening?"},
    ),
]

for chunk in llm.chat(messages=messages, stream=True):
    for msg in chunk:
        if isinstance(msg.content, str):
            print(msg.content, end="", flush=True)
        else:
            # Handle multimodal content (images, structured data)

            pass

```

The `_chat_stream` implementation in [`qwenvl_dashscope.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwenvl_dashscope.py) calls the DashScope endpoint with streaming enabled, yielding accumulated `Message` objects that include both text and image references.

### Using the Unified Model Registry

For backend-agnostic code, use the `LLM_REGISTRY` to instantiate models while keeping the streaming logic identical:

```python
from qwen_agent.llm.base import LLM_REGISTRY

# Load any registered backend: "oai", "qwen_dashscope", "qwenvl_dashscope"

model_cls = LLM_REGISTRY["oai"]
llm = model_cls(cfg={"model": "gpt-4o-mini"})

messages = [{"role": "user", "content": "Explain quantum entanglement."}]

for resp in llm.chat(messages, stream=True):
    # resp is a List[Message] or List[dict] depending on input format

    content = resp[0]["content"] if isinstance(resp[0], dict) else resp[0].content
    print(content, end="")

```

This pattern allows switching between OpenAI, DashScope, and other providers without modifying your streaming consumption code.

## Working with Streamed Output in Agents

When building custom agents, encapsulate the iterator consumption within your agent's methods to process tokens as they arrive:

```python
class StreamingAgent:
    def __init__(self, llm):
        self.llm = llm
    
    def ask(self, query):
        messages = [{"role": "user", "content": query}]
        full_text = ""
        
        for part in self.llm.chat(messages, stream=True):
            # Each part is a list containing Message objects

            token = part[0].content
            full_text += token
            # Optional: Update UI, log progress, or trigger side effects

        
        return full_text

# Usage with any LLM backend

agent = StreamingAgent(TextChatAtOAI(cfg={}))
result = agent.ask("Write a poem about artificial intelligence.")

```

## Configuration Best Practices

### Why delta_stream=False is Recommended

The Qwen-Agent framework explicitly recommends keeping `delta_stream=False` (the default) when streaming. According to the deprecation warning in [`qwen_agent/llm/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/base.py) (lines 69-74), setting `delta_stream=True` is deprecated because it complicates downstream retry logic and state management.

With `delta_stream=False`, each yielded chunk contains the **full accumulated response** up to that point, rather than just the new tokens. This design simplifies:

- **Retry mechanisms**: The full context is available at every step if the connection drops.
- **UI rendering**: You can display the complete current state without maintaining separate accumulation logic.
- **Debugging**: Each chunk represents a valid intermediate state of the conversation.

If you require true delta behavior (only new tokens), calculate the difference between consecutive chunks manually while retaining the reliability of accumulated state.

## Summary

- **Enable streaming** by passing `stream=True` to `BaseChatModel.chat()`, which returns an iterator instead of a complete response.
- **Consume the iterator** using a standard `for` loop; each iteration yields a `List[Message]` containing the accumulated response.
- **Use `delta_stream=False`** (the default) to receive full accumulated responses at each step, simplifying error handling and UI updates.
- **Implementation files**: [`qwen_agent/llm/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/base.py) defines the interface, while [`oai.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/oai.py), [`qwen_dashscope.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_dashscope.py), and [`qwenvl_dashscope.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwenvl_dashscope.py) provide provider-specific streaming logic.

## Frequently Asked Questions

### What return type does stream=True produce?

When `stream=True`, the `chat` method returns a Python **generator** (iterator), not a list. Each iteration yields a `List[Message]` containing the complete accumulated response up to the current token. As implemented in [`qwen_agent/llm/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/base.py), the type signature remains consistent with non-streaming calls, but the consumption pattern changes from direct assignment to iteration.

### Can I use streaming with function calling?

Yes. The streaming interface fully supports function calling agents. The iterator yields partial content that may include function call definitions as they are generated by the model. You can inspect `message.function_call` attributes in each chunk to detect when the model begins emitting tool calls. The [`examples/function_calling.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/examples/function_calling.py) file in the Qwen-Agent repository demonstrates this pattern.

### Is delta_stream=True still supported?

While technically functional, `delta_stream=True` is **deprecated** according to lines 69-74 in [`qwen_agent/llm/base.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/qwen_agent/llm/base.py). The framework maintainers recommend always using `delta_stream=False` (the default) because accumulated chunks provide better reliability for retry logic and are easier to render in user interfaces. Future versions may remove support for delta-only streaming entirely.

### How do I handle multimodal streaming responses?

When streaming with multimodal models like Qwen-VL, each yielded `Message` may contain either text strings or structured content dictionaries with image references. Check `isinstance(msg.content, str)` to distinguish between text and multimodal content. The [`examples/llm_vl_mix_text.py`](https://github.com/QwenLM/Qwen-Agent/blob/main/examples/llm_vl_mix_text.py) example file demonstrates how to process mixed text-image streams in real-time.