#LlmResponseStream vs AggLlmResponse in NVIDIA Switchyard: Understanding Streaming Behavior

#LlmResponseStream vs AggLlmResponse in NVIDIA Switchyard: Understanding Streaming Behavior

LlmResponseStream delivers an asynchronous iterator of real-time tokens and events for low-latency consumption, while AggLlmResponse provides a single complete dictionary representing the full LLM output only after generation finishes.

The NVIDIA Switchyard framework distinguishes between incremental and batched LLM outputs through two distinct response types defined in the Python bindings. Understanding when algorithms return LlmResponse.Stream versus LlmResponse.Agg is critical for optimizing latency-sensitive chat interfaces versus batch processing pipelines.

Core Response Types in Switchyard

LlmResponse.Stream: Real-Time Event Streaming

The LlmResponse.Stream variant wraps an asynchronous iterator (AsyncIterator[dict]) that yields individual events as soon as the upstream LLM produces them. Defined in switchyard_rust/libsy.py (lines 44-60), this type provides lazy iteration—the algorithm starts sending the first chunk immediately without waiting for the complete response.

Each iteration yields discrete events including:

  • Generated tokens
  • Tool-call deltas
  • Usage statistics updates
  • Protocol-level metadata

This design enables true real-time streaming where clients can display tokens as they arrive rather than waiting for the full response.

LlmResponse.Agg: Complete Response Aggregation

The LlmResponse.Agg variant holds a single aggregated dictionary representing the complete response once the LLM has finished generating. Also defined in switchyard_rust/libsy.py (lines 44-52), this type exposes no streaming events; instead, the caller receives the entire payload at once.

Use this mode when downstream processing requires the full context (e.g., structured JSON parsing, batch evaluation, or logging complete interactions) and latency is less critical than having atomic access to the entire output.

Algorithm Integration and Streaming Behavior

Switchyard algorithms expose both response types through the same interface. The core entry point Algorithm.run_stream in crates/libsy/src/core.rs returns an AsyncIterator[Step.CallModel | Step.Done].

During execution:

  • Step.CallModel contains a ModelCall that can respond with either LlmResponse.Stream or LlmResponse.Agg
  • Step.Done delivers the final outcome, which may contain either response type depending on the algorithm's routing decisions

The translation layer (crate switchyard-translation) handles conversion between raw protocol streams and these high-level Python types. According to crates/switchyard-translation/src/helpers.rs:

  • decode_aggregated_response (lines 46-55) builds an AggLlmResponse from a completed JSON body
  • decode_stream_event (lines 292-301) produces LlmResponseStreamEvent items for the streaming case

Additionally, crates/switchyard-translation/src/codecs/responses/stream.rs implements the streaming codec that transforms raw HTTP responses into sequences of LlmResponseChunk objects.

Source Code Architecture

The following files define the streaming behavior across the Rust and Python boundary:

Practical Python Usage

The following pattern demonstrates handling both response types within a Switchyard algorithm:

from switchyard.libsy import LlmResponse, Step, algorithms

async def process_responses():
    # Run a routing algorithm that may stream

    async for step in algorithms.random().run_stream(
        request={}, 
        models={"gpt": ["gpt-4"]}
    ):
        match step:
            case Step.CallModel(call):
                # The model can choose streaming or aggregate response

                call.respond(LlmResponse.Stream(events()))   # streaming

                # or

                call.respond(LlmResponse.Agg({"content": "full answer"}))  # aggregate

            case Step.Done(outcome):
                # At the end you get either an Agg or a Stream (if never switched)

                resp = outcome.response
                match resp:
                    case LlmResponse.Agg(data):
                        print("Aggregated:", data)
                    case LlmResponse.Stream(stream):
                        async for ev in stream:
                            print("Stream event:", ev)

Key implementation details:

  • Streaming is ideal for UI-driven chat where displaying tokens as they arrive improves perceived responsiveness
  • Aggregate simplifies downstream processing when only the final answer matters (e.g., batch evaluation or structured output parsing)

Summary

  • LlmResponseStream provides an asynchronous iterator over incremental events, enabling real-time token-by-token consumption with minimal latency
  • AggLlmResponse delivers a single complete dictionary after generation finishes, optimizing for atomic access and batch processing
  • Switchyard algorithms expose both types through the unified run_stream interface, allowing dynamic selection based on model configuration or routing logic
  • The translation layer in switchyard-translation handles protocol decoding via decode_stream_event and decode_aggregated_response helpers

Frequently Asked Questions

When should I use LlmResponse.Stream versus LlmResponse.Agg?

Use LlmResponse.Stream for interactive applications requiring real-time feedback, such as chat interfaces where users see tokens appear as they generate. Use LlmResponse.Agg for batch processing, automated evaluation pipelines, or any scenario where downstream logic requires the complete response before proceeding.

How do I detect which response type an algorithm returns?

Use pattern matching (Python 3.10+ match statements) on the LlmResponse variant within Step.CallModel or Step.Done outcomes. Check for LlmResponse.Agg to handle complete payloads synchronously, or LlmResponse.Stream to initiate asynchronous iteration over events.

Can a single algorithm switch between streaming and aggregate modes mid-execution?

Yes. Switchyard algorithms can dynamically choose between response types for each model call within a single run_stream execution. The algorithm might stream from one model for latency-sensitive display, then aggregate responses from another model for final processing, depending on routing decisions and model configurations.

Where is the streaming protocol defined in the Switchyard source code?

The low-level protocol definition resides in crates/protocol/src/llm.rs, which defines LlmResponseStream and associated event types. The Python bindings exposing these to user code are implemented in switchyard_rust/libsy.py, while the translation logic connecting HTTP responses to these types lives in crates/switchyard-translation/src/helpers.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →