How Switchyard Handles Streaming Responses with Protocol Translation

Switchyard processes streaming LLM responses as a sequence of provider-specific JSON events, decoding each into a canonical internal representation through the TranslationEngine, translating between formats, and re-encoding into the target provider's wire format to enable real-time protocol conversion.

The NVIDIA-NeMo/Switchyard repository provides a high-performance translation layer for LLM inference APIs. When handling streaming responses with protocol translation, Switchyard treats each chunk as a discrete event that flows through a stateful pipeline, allowing seamless conversion between incompatible provider formats like OpenAI Chat and Anthropic Messages without buffering the entire response.

The Four-Stage Streaming Pipeline

The translation engine in crates/switchyard-translation/src/engine.rs implements a stateful pipeline that processes each streaming event through four distinct operations.

1. Decode and Preserve

The decode_stream_event method ingests raw provider-specific JSON while preserving the original payload. This allows the engine to maintain the raw event for potential replay while creating a canonical internal representation.

decode_stream_event(state, source_format, raw_event) -> LlmResponseStreamEvent

2. Canonical Translation

Using translate_event, the engine converts the canonical representation from the source format to the target format. This operation is one-to-many (or zero), meaning a single source event may generate multiple target events depending on protocol differences.

translate_event(state, source_format, target_format, raw_event) -> Vec<Value>

3. Encode and Emit

The encode_stream_event method serializes the canonical events back into the target provider's expected JSON structure, preparing them for wire transmission.

encode_stream_event(state, target_format, canonical_event) -> Vec<Value>

4. Stream Termination

When the upstream provider closes the connection, finish_stream generates any provider-specific termination events required by the target format, ensuring proper stream closure semantics.

finish_stream(state, target_format) -> Vec<Value>

Registry-Based Codec Architecture

Switchyard maintains two specialized registries for format handling:

  • FormatRegistry: Maps buffered request/response formats to their corresponding FormatCodec implementations. Populated via FormatRegistry::with_builtins().
  • StreamCodecRegistry: Maps streaming formats to StreamCodec instances that handle per-event decode/encode operations. Populated via StreamCodecRegistry::with_builtins().

Both registries include built-in codecs such as OpenAiChatCodec, AnthropicMessagesCodec, and OpenAiResponsesCodec, enabling out-of-the-box support for major LLM providers.

Stateful Translation Context

The StreamTranslationState struct maintains per-stream context across all four pipeline stages. This stateful approach enables the engine to:

  • Track whether the source provider emitted a "stop" token
  • Carry ordering information between events
  • Generate appropriate closing sequences when translating between protocols with different termination semantics

Implementation Examples

Low-Level Rust FFI Usage

The Python bindings in switchyard_rust/libsy.py expose the Rust translation engine directly for fine-grained control over streaming responses with protocol translation:

from switchyard_rust import libsy as sw

engine = sw.TranslationEngine.default()
state = sw.StreamTranslationState.default()

for raw_event in openai_chat_stream:
    # Decode while preserving original JSON

    preserved = engine.decode_stream_event(
        state,
        sw.FormatId.OpenAiChat,
        raw_event
    )
    
    # Translate to Anthropic format (may produce multiple events)

    translated = engine.translate_event(
        state,
        sw.FormatId.OpenAiChat,
        sw.FormatId.AnthropicMessages,
        raw_event
    )
    
    # Encode to target provider's JSON structure

    for canonical in translated:
        out_events = engine.encode_stream_event(
            state,
            sw.FormatId.AnthropicMessages,
            canonical
        )
        for ev in out_events:
            client_send(ev)

# Emit termination events required by target format

final = engine.finish_stream(state, sw.FormatId.AnthropicMessages)
for ev in final:
    client_send(ev)

High-Level Python API

For most use cases, the switchyard.libsy.algorithms module abstracts the translation complexity while using the same underlying engine:

import switchyard.libsy.algorithms as alg

async for chunk in alg.algorithms.noop().run_stream(request_body()):
    # Translation happens automatically within the algorithm

    print(chunk.delta["text"])

Key Source Files

The streaming translation logic is distributed across the following source files according to the Switchyard codebase:

Summary

  • Switchyard handles streaming responses with protocol translation by treating each chunk as a discrete JSON event in a four-stage pipeline: decode, translate, encode, and finish.
  • The TranslationEngine in engine.rs provides the core methods for stateful stream processing, operating on a persistent StreamTranslationState that carries context across events.
  • FormatRegistry and StreamCodecRegistry manage built-in codecs for providers like OpenAI and Anthropic through their respective with_builtins() constructors.
  • The architecture supports one-to-many event mapping during translation, allowing a single source event to generate multiple target events when protocol semantics differ.
  • Both low-level Rust FFI and high-level Python APIs expose the same translation engine, as implemented in switchyard_rust/libsy.py and switchyard/libsy/algorithms.py.

Frequently Asked Questions

What is the difference between FormatRegistry and StreamCodecRegistry?

FormatRegistry handles buffered request/response formats using FormatCodec implementations, while StreamCodecRegistry specifically manages streaming formats through StreamCodec instances that process discrete events. Both are initialized with with_builtins() to support major providers like OpenAI and Anthropic, but they serve different traffic patterns—buffered versus streaming.

Can a single source event generate multiple target events during translation?

Yes. The translate_event method returns Vec<Value>, allowing zero-to-many mapping between protocols. This handles cases where one provider's streaming semantics require multiple events to represent a single event from another provider, such as when translating between different delta encoding schemes.

How does Switchyard handle stream termination when protocols differ?

The finish_stream method examines the StreamTranslationState to generate any provider-specific termination events required by the target format. This ensures proper closure semantics even when the source and target protocols use different stop sequences or EOF markers, preventing client-side timeouts or parsing errors.

Is the streaming translation stateful or stateless?

The translation is stateful. Each stream maintains a StreamTranslationState instance that persists across the decode_stream_event, translate_event, encode_stream_event, and finish_stream calls. This state carries critical context such as stop tokens and ordering information necessary for correct protocol translation throughout the entire response lifecycle.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →