# #LlmResponseStream vs AggLlmResponse in NVIDIA Switchyard: Understanding Streaming Behavior

> Understand LlmResponseStream vs AggLlmResponse in NVIDIA Switchyard. Discover how LlmResponseStream offers real-time tokens for low latency while AggLlmResponse provides complete output after generation.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-12

---

#LlmResponseStream vs AggLlmResponse in NVIDIA Switchyard: Understanding Streaming Behavior

**LlmResponseStream delivers an asynchronous iterator of real-time tokens and events for low-latency consumption, while AggLlmResponse provides a single complete dictionary representing the full LLM output only after generation finishes.**

The NVIDIA Switchyard framework distinguishes between incremental and batched LLM outputs through two distinct response types defined in the Python bindings. Understanding when algorithms return `LlmResponse.Stream` versus `LlmResponse.Agg` is critical for optimizing latency-sensitive chat interfaces versus batch processing pipelines.

## Core Response Types in Switchyard

### LlmResponse.Stream: Real-Time Event Streaming

The `LlmResponse.Stream` variant wraps an **asynchronous iterator** (`AsyncIterator[dict]`) that yields individual events as soon as the upstream LLM produces them. Defined in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) (lines 44-60), this type provides **lazy** iteration—the algorithm starts sending the first chunk immediately without waiting for the complete response.

Each iteration yields discrete events including:
- Generated tokens
- Tool-call deltas
- Usage statistics updates
- Protocol-level metadata

This design enables true real-time streaming where clients can display tokens as they arrive rather than waiting for the full response.

### LlmResponse.Agg: Complete Response Aggregation

The `LlmResponse.Agg` variant holds a **single aggregated dictionary** representing the complete response once the LLM has finished generating. Also defined in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) (lines 44-52), this type exposes no streaming events; instead, the caller receives the entire payload at once.

Use this mode when downstream processing requires the full context (e.g., structured JSON parsing, batch evaluation, or logging complete interactions) and latency is less critical than having atomic access to the entire output.

## Algorithm Integration and Streaming Behavior

Switchyard algorithms expose both response types through the same interface. The core entry point `Algorithm.run_stream` in [`crates/libsy/src/core.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core.rs) returns an `AsyncIterator[Step.CallModel | Step.Done]`.

During execution:

- **Step.CallModel** contains a `ModelCall` that can `respond` with either `LlmResponse.Stream` or `LlmResponse.Agg`
- **Step.Done** delivers the final outcome, which may contain either response type depending on the algorithm's routing decisions

The **translation layer** (crate `switchyard-translation`) handles conversion between raw protocol streams and these high-level Python types. According to [`crates/switchyard-translation/src/helpers.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/helpers.rs):

- **`decode_aggregated_response`** (lines 46-55) builds an `AggLlmResponse` from a completed JSON body
- **`decode_stream_event`** (lines 292-301) produces `LlmResponseStreamEvent` items for the streaming case

Additionally, [`crates/switchyard-translation/src/codecs/responses/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/codecs/responses/stream.rs) implements the streaming codec that transforms raw HTTP responses into sequences of `LlmResponseChunk` objects.

## Source Code Architecture

The following files define the streaming behavior across the Rust and Python boundary:

- **[`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py)** – Python bindings defining `LlmResponse.Agg` and `LlmResponse.Stream` variants
- **[`crates/protocol/src/llm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs)** – Low-level protocol definitions for `LlmResponseStream` and event structures
- **[`crates/switchyard-translation/src/helpers.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/helpers.rs)** – Translation functions `decode_aggregated_response` and `decode_stream_event`
- **[`crates/switchyard-translation/src/codecs/responses/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/codecs/responses/stream.rs)** – HTTP-to-stream decoding logic
- **[`crates/libsy/src/core.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core.rs)** – `Algorithm::run_stream` orchestration logic

## Practical Python Usage

The following pattern demonstrates handling both response types within a Switchyard algorithm:

```python
from switchyard.libsy import LlmResponse, Step, algorithms

async def process_responses():
    # Run a routing algorithm that may stream

    async for step in algorithms.random().run_stream(
        request={}, 
        models={"gpt": ["gpt-4"]}
    ):
        match step:
            case Step.CallModel(call):
                # The model can choose streaming or aggregate response

                call.respond(LlmResponse.Stream(events()))   # streaming

                # or

                call.respond(LlmResponse.Agg({"content": "full answer"}))  # aggregate

            case Step.Done(outcome):
                # At the end you get either an Agg or a Stream (if never switched)

                resp = outcome.response
                match resp:
                    case LlmResponse.Agg(data):
                        print("Aggregated:", data)
                    case LlmResponse.Stream(stream):
                        async for ev in stream:
                            print("Stream event:", ev)

```

Key implementation details:
- **Streaming** is ideal for UI-driven chat where displaying tokens as they arrive improves perceived responsiveness
- **Aggregate** simplifies downstream processing when only the final answer matters (e.g., batch evaluation or structured output parsing)

## Summary

- **LlmResponseStream** provides an asynchronous iterator over incremental events, enabling real-time token-by-token consumption with minimal latency
- **AggLlmResponse** delivers a single complete dictionary after generation finishes, optimizing for atomic access and batch processing
- Switchyard algorithms expose both types through the unified `run_stream` interface, allowing dynamic selection based on model configuration or routing logic
- The translation layer in `switchyard-translation` handles protocol decoding via `decode_stream_event` and `decode_aggregated_response` helpers

## Frequently Asked Questions

### When should I use LlmResponse.Stream versus LlmResponse.Agg?

Use **LlmResponse.Stream** for interactive applications requiring real-time feedback, such as chat interfaces where users see tokens appear as they generate. Use **LlmResponse.Agg** for batch processing, automated evaluation pipelines, or any scenario where downstream logic requires the complete response before proceeding.

### How do I detect which response type an algorithm returns?

Use pattern matching (Python 3.10+ `match` statements) on the `LlmResponse` variant within `Step.CallModel` or `Step.Done` outcomes. Check for `LlmResponse.Agg` to handle complete payloads synchronously, or `LlmResponse.Stream` to initiate asynchronous iteration over events.

### Can a single algorithm switch between streaming and aggregate modes mid-execution?

Yes. Switchyard algorithms can dynamically choose between response types for each model call within a single `run_stream` execution. The algorithm might stream from one model for latency-sensitive display, then aggregate responses from another model for final processing, depending on routing decisions and model configurations.

### Where is the streaming protocol defined in the Switchyard source code?

The low-level protocol definition resides in [`crates/protocol/src/llm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs), which defines `LlmResponseStream` and associated event types. The Python bindings exposing these to user code are implemented in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py), while the translation logic connecting HTTP responses to these types lives in [`crates/switchyard-translation/src/helpers.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/helpers.rs).