#LlmResponseStream vs AggLlmResponse in NVIDIA Switchyard: Understanding Streaming Behavior
#LlmResponseStream vs AggLlmResponse in NVIDIA Switchyard: Understanding Streaming Behavior
LlmResponseStream delivers an asynchronous iterator of real-time tokens and events for low-latency consumption, while AggLlmResponse provides a single complete dictionary representing the full LLM output only after generation finishes.
The NVIDIA Switchyard framework distinguishes between incremental and batched LLM outputs through two distinct response types defined in the Python bindings. Understanding when algorithms return LlmResponse.Stream versus LlmResponse.Agg is critical for optimizing latency-sensitive chat interfaces versus batch processing pipelines.
Core Response Types in Switchyard
LlmResponse.Stream: Real-Time Event Streaming
The LlmResponse.Stream variant wraps an asynchronous iterator (AsyncIterator[dict]) that yields individual events as soon as the upstream LLM produces them. Defined in switchyard_rust/libsy.py (lines 44-60), this type provides lazy iteration—the algorithm starts sending the first chunk immediately without waiting for the complete response.
Each iteration yields discrete events including:
- Generated tokens
- Tool-call deltas
- Usage statistics updates
- Protocol-level metadata
This design enables true real-time streaming where clients can display tokens as they arrive rather than waiting for the full response.
LlmResponse.Agg: Complete Response Aggregation
The LlmResponse.Agg variant holds a single aggregated dictionary representing the complete response once the LLM has finished generating. Also defined in switchyard_rust/libsy.py (lines 44-52), this type exposes no streaming events; instead, the caller receives the entire payload at once.
Use this mode when downstream processing requires the full context (e.g., structured JSON parsing, batch evaluation, or logging complete interactions) and latency is less critical than having atomic access to the entire output.
Algorithm Integration and Streaming Behavior
Switchyard algorithms expose both response types through the same interface. The core entry point Algorithm.run_stream in crates/libsy/src/core.rs returns an AsyncIterator[Step.CallModel | Step.Done].
During execution:
- Step.CallModel contains a
ModelCallthat canrespondwith eitherLlmResponse.StreamorLlmResponse.Agg - Step.Done delivers the final outcome, which may contain either response type depending on the algorithm's routing decisions
The translation layer (crate switchyard-translation) handles conversion between raw protocol streams and these high-level Python types. According to crates/switchyard-translation/src/helpers.rs:
decode_aggregated_response(lines 46-55) builds anAggLlmResponsefrom a completed JSON bodydecode_stream_event(lines 292-301) producesLlmResponseStreamEventitems for the streaming case
Additionally, crates/switchyard-translation/src/codecs/responses/stream.rs implements the streaming codec that transforms raw HTTP responses into sequences of LlmResponseChunk objects.
Source Code Architecture
The following files define the streaming behavior across the Rust and Python boundary:
switchyard_rust/libsy.py– Python bindings definingLlmResponse.AggandLlmResponse.Streamvariantscrates/protocol/src/llm.rs– Low-level protocol definitions forLlmResponseStreamand event structurescrates/switchyard-translation/src/helpers.rs– Translation functionsdecode_aggregated_responseanddecode_stream_eventcrates/switchyard-translation/src/codecs/responses/stream.rs– HTTP-to-stream decoding logiccrates/libsy/src/core.rs–Algorithm::run_streamorchestration logic
Practical Python Usage
The following pattern demonstrates handling both response types within a Switchyard algorithm:
from switchyard.libsy import LlmResponse, Step, algorithms
async def process_responses():
# Run a routing algorithm that may stream
async for step in algorithms.random().run_stream(
request={},
models={"gpt": ["gpt-4"]}
):
match step:
case Step.CallModel(call):
# The model can choose streaming or aggregate response
call.respond(LlmResponse.Stream(events())) # streaming
# or
call.respond(LlmResponse.Agg({"content": "full answer"})) # aggregate
case Step.Done(outcome):
# At the end you get either an Agg or a Stream (if never switched)
resp = outcome.response
match resp:
case LlmResponse.Agg(data):
print("Aggregated:", data)
case LlmResponse.Stream(stream):
async for ev in stream:
print("Stream event:", ev)
Key implementation details:
- Streaming is ideal for UI-driven chat where displaying tokens as they arrive improves perceived responsiveness
- Aggregate simplifies downstream processing when only the final answer matters (e.g., batch evaluation or structured output parsing)
Summary
- LlmResponseStream provides an asynchronous iterator over incremental events, enabling real-time token-by-token consumption with minimal latency
- AggLlmResponse delivers a single complete dictionary after generation finishes, optimizing for atomic access and batch processing
- Switchyard algorithms expose both types through the unified
run_streaminterface, allowing dynamic selection based on model configuration or routing logic - The translation layer in
switchyard-translationhandles protocol decoding viadecode_stream_eventanddecode_aggregated_responsehelpers
Frequently Asked Questions
When should I use LlmResponse.Stream versus LlmResponse.Agg?
Use LlmResponse.Stream for interactive applications requiring real-time feedback, such as chat interfaces where users see tokens appear as they generate. Use LlmResponse.Agg for batch processing, automated evaluation pipelines, or any scenario where downstream logic requires the complete response before proceeding.
How do I detect which response type an algorithm returns?
Use pattern matching (Python 3.10+ match statements) on the LlmResponse variant within Step.CallModel or Step.Done outcomes. Check for LlmResponse.Agg to handle complete payloads synchronously, or LlmResponse.Stream to initiate asynchronous iteration over events.
Can a single algorithm switch between streaming and aggregate modes mid-execution?
Yes. Switchyard algorithms can dynamically choose between response types for each model call within a single run_stream execution. The algorithm might stream from one model for latency-sensitive display, then aggregate responses from another model for final processing, depending on routing decisions and model configurations.
Where is the streaming protocol defined in the Switchyard source code?
The low-level protocol definition resides in crates/protocol/src/llm.rs, which defines LlmResponseStream and associated event types. The Python bindings exposing these to user code are implemented in switchyard_rust/libsy.py, while the translation logic connecting HTTP responses to these types lives in crates/switchyard-translation/src/helpers.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →