How to Stream Outputs from Qwen-Agent: A Complete Implementation Guide
Pass stream=True to BaseChatModel.chat() and iterate over the returned generator to receive partial LLM responses as they are generated by the underlying service.
Qwen-Agent is a framework for building LLM applications with advanced tool-use and multimodal capabilities. Streaming support is built directly into the BaseChatModel interface, enabling real-time token delivery without waiting for complete responses. By setting the stream parameter to True, developers can consume partial outputs through a unified iterator protocol that works across OpenAI, DashScope, and Qwen-VL backends.
Understanding the Streaming Architecture
Streaming in Qwen-Agent follows a provider-agnostic iterator pattern defined in the base model class. The architecture separates the public API from backend-specific implementations, ensuring consistent behavior whether you use text-only or multimodal models.
The Entry Point in base.py
The BaseChatModel.chat method in qwen_agent/llm/base.py serves as the unified entry point for all LLM interactions. When you invoke chat with stream=True, the method checks this flag at lines 18-27 and forwards the request to the model-specific _chat_stream implementation rather than the standard _chat_no_stream method.
The method signature accepts two critical boolean parameters: stream to enable the iterator behavior, and delta_stream to control chunk format. According to the source code, when streaming is requested, the method returns a generator that yields lists of Message objects representing the conversation state at each token generation step.
Provider-Specific Streaming Implementations
Each LLM backend implements its own _chat_stream generator that handles the actual API communication with stream=True:
- OpenAI-compatible models: The
TextChatAtOAIclass inqwen_agent/llm/oai.pyimplements_chat_streamat lines 98-108, invoking the provider's streaming endpoint and converting chunks intoMessageobjects. - Qwen-VL multimodal models: The implementation in
qwen_agent/llm/qwenvl_dashscope.py(lines 64-78) handles mixed image-text content, yielding accumulated responses suitable for real-time display. - DashScope text models: Implemented in
qwen_agent/llm/qwen_dashscope.py, providing streaming for Qwen-3 and other text-only models.
All implementations yield List[Message] objects, maintaining type consistency across different providers.
How to Implement Streaming in Your Code
Basic Text Streaming with OpenAI-Compatible Models
For OpenAI-compatible endpoints, initialize TextChatAtOAI and consume the iterator in a for loop. Each iteration yields the full accumulated response up to the current token:
from qwen_agent.llm.oai import TextChatAtOAI
from qwen_agent.llm.schema import Message, USER
# Initialize with your API configuration
llm = TextChatAtOAI(cfg={"api_key": "YOUR_KEY"})
messages = [Message(role=USER, content="Tell me a short story about a cat.")]
# stream=True returns an iterator; delta_stream defaults to False
for partial in llm.chat(messages=messages, stream=True):
# partial is a list containing a single Message with accumulated text
print(partial[0].content, end="", flush=True)
The internal call reaches _chat_stream in oai.py, which manages the HTTP connection and yields incremental updates as the provider sends them.
Multimodal Streaming with Qwen-VL
Vision-language models use the identical streaming interface, though you must handle potential image references in the yielded content:
from qwen_agent.llm.qwenvl_dashscope import QwenVL
from qwen_agent.llm.schema import Message, USER
llm = QwenVL(cfg={"model": "qwen-vl-plus", "model_server": "dashscope"})
messages = [
Message(role=USER, content="Describe the image and answer the question."),
Message(
role=USER,
content={"image": "https://example.com/scene.jpg", "text": "What is happening?"},
),
]
for chunk in llm.chat(messages=messages, stream=True):
for msg in chunk:
if isinstance(msg.content, str):
print(msg.content, end="", flush=True)
else:
# Handle multimodal content (images, structured data)
pass
The _chat_stream implementation in qwenvl_dashscope.py calls the DashScope endpoint with streaming enabled, yielding accumulated Message objects that include both text and image references.
Using the Unified Model Registry
For backend-agnostic code, use the LLM_REGISTRY to instantiate models while keeping the streaming logic identical:
from qwen_agent.llm.base import LLM_REGISTRY
# Load any registered backend: "oai", "qwen_dashscope", "qwenvl_dashscope"
model_cls = LLM_REGISTRY["oai"]
llm = model_cls(cfg={"model": "gpt-4o-mini"})
messages = [{"role": "user", "content": "Explain quantum entanglement."}]
for resp in llm.chat(messages, stream=True):
# resp is a List[Message] or List[dict] depending on input format
content = resp[0]["content"] if isinstance(resp[0], dict) else resp[0].content
print(content, end="")
This pattern allows switching between OpenAI, DashScope, and other providers without modifying your streaming consumption code.
Working with Streamed Output in Agents
When building custom agents, encapsulate the iterator consumption within your agent's methods to process tokens as they arrive:
class StreamingAgent:
def __init__(self, llm):
self.llm = llm
def ask(self, query):
messages = [{"role": "user", "content": query}]
full_text = ""
for part in self.llm.chat(messages, stream=True):
# Each part is a list containing Message objects
token = part[0].content
full_text += token
# Optional: Update UI, log progress, or trigger side effects
return full_text
# Usage with any LLM backend
agent = StreamingAgent(TextChatAtOAI(cfg={}))
result = agent.ask("Write a poem about artificial intelligence.")
Configuration Best Practices
Why delta_stream=False is Recommended
The Qwen-Agent framework explicitly recommends keeping delta_stream=False (the default) when streaming. According to the deprecation warning in qwen_agent/llm/base.py (lines 69-74), setting delta_stream=True is deprecated because it complicates downstream retry logic and state management.
With delta_stream=False, each yielded chunk contains the full accumulated response up to that point, rather than just the new tokens. This design simplifies:
- Retry mechanisms: The full context is available at every step if the connection drops.
- UI rendering: You can display the complete current state without maintaining separate accumulation logic.
- Debugging: Each chunk represents a valid intermediate state of the conversation.
If you require true delta behavior (only new tokens), calculate the difference between consecutive chunks manually while retaining the reliability of accumulated state.
Summary
- Enable streaming by passing
stream=TruetoBaseChatModel.chat(), which returns an iterator instead of a complete response. - Consume the iterator using a standard
forloop; each iteration yields aList[Message]containing the accumulated response. - Use
delta_stream=False(the default) to receive full accumulated responses at each step, simplifying error handling and UI updates. - Implementation files:
qwen_agent/llm/base.pydefines the interface, whileoai.py,qwen_dashscope.py, andqwenvl_dashscope.pyprovide provider-specific streaming logic.
Frequently Asked Questions
What return type does stream=True produce?
When stream=True, the chat method returns a Python generator (iterator), not a list. Each iteration yields a List[Message] containing the complete accumulated response up to the current token. As implemented in qwen_agent/llm/base.py, the type signature remains consistent with non-streaming calls, but the consumption pattern changes from direct assignment to iteration.
Can I use streaming with function calling?
Yes. The streaming interface fully supports function calling agents. The iterator yields partial content that may include function call definitions as they are generated by the model. You can inspect message.function_call attributes in each chunk to detect when the model begins emitting tool calls. The examples/function_calling.py file in the Qwen-Agent repository demonstrates this pattern.
Is delta_stream=True still supported?
While technically functional, delta_stream=True is deprecated according to lines 69-74 in qwen_agent/llm/base.py. The framework maintainers recommend always using delta_stream=False (the default) because accumulated chunks provide better reliability for retry logic and are easier to render in user interfaces. Future versions may remove support for delta-only streaming entirely.
How do I handle multimodal streaming responses?
When streaming with multimodal models like Qwen-VL, each yielded Message may contain either text strings or structured content dictionaries with image references. Check isinstance(msg.content, str) to distinguish between text and multimodal content. The examples/llm_vl_mix_text.py example file demonstrates how to process mixed text-image streams in real-time.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →