# How Switchyard Handles Streaming Responses with Protocol Translation

> Learn how Switchyard handles streaming LLM responses with protocol translation. It decodes, translates, and re-encodes events for real-time format conversion.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-17

---

**Switchyard processes streaming LLM responses as a sequence of provider-specific JSON events, decoding each into a canonical internal representation through the `TranslationEngine`, translating between formats, and re-encoding into the target provider's wire format to enable real-time protocol conversion.**

The NVIDIA-NeMo/Switchyard repository provides a high-performance translation layer for LLM inference APIs. When handling streaming responses with protocol translation, Switchyard treats each chunk as a discrete event that flows through a stateful pipeline, allowing seamless conversion between incompatible provider formats like OpenAI Chat and Anthropic Messages without buffering the entire response.

## The Four-Stage Streaming Pipeline

The translation engine in [`crates/switchyard-translation/src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/engine.rs) implements a stateful pipeline that processes each streaming event through four distinct operations.

### 1. Decode and Preserve

The `decode_stream_event` method ingests raw provider-specific JSON while preserving the original payload. This allows the engine to maintain the raw event for potential replay while creating a canonical internal representation.

```rust
decode_stream_event(state, source_format, raw_event) -> LlmResponseStreamEvent

```

### 2. Canonical Translation

Using `translate_event`, the engine converts the canonical representation from the source format to the target format. This operation is one-to-many (or zero), meaning a single source event may generate multiple target events depending on protocol differences.

```rust
translate_event(state, source_format, target_format, raw_event) -> Vec<Value>

```

### 3. Encode and Emit

The `encode_stream_event` method serializes the canonical events back into the target provider's expected JSON structure, preparing them for wire transmission.

```rust
encode_stream_event(state, target_format, canonical_event) -> Vec<Value>

```

### 4. Stream Termination

When the upstream provider closes the connection, `finish_stream` generates any provider-specific termination events required by the target format, ensuring proper stream closure semantics.

```rust
finish_stream(state, target_format) -> Vec<Value>

```

## Registry-Based Codec Architecture

Switchyard maintains two specialized registries for format handling:

- **FormatRegistry**: Maps buffered request/response formats to their corresponding `FormatCodec` implementations. Populated via `FormatRegistry::with_builtins()`.
- **StreamCodecRegistry**: Maps streaming formats to `StreamCodec` instances that handle per-event decode/encode operations. Populated via `StreamCodecRegistry::with_builtins()`.

Both registries include built-in codecs such as `OpenAiChatCodec`, `AnthropicMessagesCodec`, and `OpenAiResponsesCodec`, enabling out-of-the-box support for major LLM providers.

## Stateful Translation Context

The `StreamTranslationState` struct maintains per-stream context across all four pipeline stages. This stateful approach enables the engine to:

- Track whether the source provider emitted a "stop" token
- Carry ordering information between events
- Generate appropriate closing sequences when translating between protocols with different termination semantics

## Implementation Examples

### Low-Level Rust FFI Usage

The Python bindings in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) expose the Rust translation engine directly for fine-grained control over streaming responses with protocol translation:

```python
from switchyard_rust import libsy as sw

engine = sw.TranslationEngine.default()
state = sw.StreamTranslationState.default()

for raw_event in openai_chat_stream:
    # Decode while preserving original JSON

    preserved = engine.decode_stream_event(
        state,
        sw.FormatId.OpenAiChat,
        raw_event
    )
    
    # Translate to Anthropic format (may produce multiple events)

    translated = engine.translate_event(
        state,
        sw.FormatId.OpenAiChat,
        sw.FormatId.AnthropicMessages,
        raw_event
    )
    
    # Encode to target provider's JSON structure

    for canonical in translated:
        out_events = engine.encode_stream_event(
            state,
            sw.FormatId.AnthropicMessages,
            canonical
        )
        for ev in out_events:
            client_send(ev)

# Emit termination events required by target format

final = engine.finish_stream(state, sw.FormatId.AnthropicMessages)
for ev in final:
    client_send(ev)

```

### High-Level Python API

For most use cases, the `switchyard.libsy.algorithms` module abstracts the translation complexity while using the same underlying engine:

```python
import switchyard.libsy.algorithms as alg

async for chunk in alg.algorithms.noop().run_stream(request_body()):
    # Translation happens automatically within the algorithm

    print(chunk.delta["text"])

```

## Key Source Files

The streaming translation logic is distributed across the following source files according to the Switchyard codebase:

- [`crates/switchyard-translation/src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/engine.rs): Core `TranslationEngine` implementing `decode_stream_event`, `translate_event`, `encode_stream_event`, and `finish_stream`
- [`crates/switchyard-translation/src/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/stream.rs): Public façade exposing `StreamCodecRegistry` and `StreamTranslationState`
- [`crates/switchyard-translation/src/codecs/openai_chat.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/codecs/openai_chat.rs): OpenAI Chat-specific streaming codec implementation
- [`crates/switchyard-translation/src/codecs/anthropic.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/codecs/anthropic.rs): Anthropic Messages streaming codec implementation
- [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py): Python FFI bindings exposing the Rust engine
- [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py): High-level Python wrappers that orchestrate stream translation

## Summary

- Switchyard handles streaming responses with protocol translation by treating each chunk as a discrete JSON event in a four-stage pipeline: **decode**, **translate**, **encode**, and **finish**.
- The `TranslationEngine` in [`engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/engine.rs) provides the core methods for stateful stream processing, operating on a persistent `StreamTranslationState` that carries context across events.
- **FormatRegistry** and **StreamCodecRegistry** manage built-in codecs for providers like OpenAI and Anthropic through their respective `with_builtins()` constructors.
- The architecture supports one-to-many event mapping during translation, allowing a single source event to generate multiple target events when protocol semantics differ.
- Both low-level Rust FFI and high-level Python APIs expose the same translation engine, as implemented in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) and [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py).

## Frequently Asked Questions

### What is the difference between FormatRegistry and StreamCodecRegistry?

**FormatRegistry** handles buffered request/response formats using `FormatCodec` implementations, while **StreamCodecRegistry** specifically manages streaming formats through `StreamCodec` instances that process discrete events. Both are initialized with `with_builtins()` to support major providers like OpenAI and Anthropic, but they serve different traffic patterns—buffered versus streaming.

### Can a single source event generate multiple target events during translation?

Yes. The `translate_event` method returns `Vec<Value>`, allowing zero-to-many mapping between protocols. This handles cases where one provider's streaming semantics require multiple events to represent a single event from another provider, such as when translating between different delta encoding schemes.

### How does Switchyard handle stream termination when protocols differ?

The `finish_stream` method examines the `StreamTranslationState` to generate any provider-specific termination events required by the target format. This ensures proper closure semantics even when the source and target protocols use different stop sequences or EOF markers, preventing client-side timeouts or parsing errors.

### Is the streaming translation stateful or stateless?

The translation is **stateful**. Each stream maintains a `StreamTranslationState` instance that persists across the `decode_stream_event`, `translate_event`, `encode_stream_event`, and `finish_stream` calls. This state carries critical context such as stop tokens and ordering information necessary for correct protocol translation throughout the entire response lifecycle.