How Switchyard Translates Streaming Responses Between Different LLM APIs

Switchyard normalizes live token streams from any supported provider into a generic intermediate representation (IR) and re-encodes them on-the-fly for any target API, enabling seamless cross-provider streaming without buffering entire responses.

The NVIDIA-NeMo/Switchyard project solves the interoperability challenge of streaming responses between different LLM APIs by treating streaming events as a provider-agnostic stream IR. Instead of buffering complete responses, Switchyard processes tokens incrementally through a stateless pipeline that decodes upstream formats, translates metadata, and encodes downstream formats in real time.

The Stream IR Pipeline: Decode, Translate, Encode

Switchyard routes incoming requests to upstream providers (OpenAI Chat, Anthropic Messages, OpenAI Responses) and processes their streaming HTTP events through a four-stage pipeline. Each stage is implemented in the crates/switchyard-translation crate to ensure type-safe conversion between wire formats.

Decoding Provider Streams into LlmResponseChunk

When an upstream provider returns a streaming response, the provider-specific codec parses raw HTTP payloads into a common enum called LlmResponseChunk. This decoding logic resides in crates/switchyard-translation/src/codecs/<provider>/stream.rs—for example, crates/switchyard-translation/src/codecs/openai_chat/stream.rs for OpenAI Chat and crates/switchyard-translation/src/codecs/anthropic/stream.rs for Anthropic.

The decoder handles provider-specific SSE (Server-Sent Events) formats, extracting chunks, usage deltas, error frames, and stop signals into the normalized IR. The LlmResponseChunk definitions and WireFormat enum live in crates/switchyard-translation/src/stream.rs, providing the type foundation for the translation system.

Translation Engine and State Management

The TranslationEngine defined in crates/switchyard-translation/src/engine.rs receives each LlmResponseChunk alongside a StreamTranslationState struct. The engine performs three critical operations:

  1. Metadata normalization—strips or adds fields like model name, usage details, and reasoning content to match downstream expectations.
  2. Format routing—determines the target WireFormat (e.g., WireFormat::OpenAiChat, WireFormat::AnthropicMessages).
  3. Stream lifecycle management—tracks whether the stream has finished and emits terminal events via engine.finish_stream.

Because translation is stateless per-request, the StreamTranslationState holds only minimal context needed to preserve identity and model information. This design allows Switchyard to pipe a live OpenAI Chat stream directly to an Anthropic-compatible client without accumulating the full response in memory.

Encoding to Target Wire Formats

After translation, the downstream codec re-encodes the LlmResponseChunk into the target provider's streaming JSON format. The encoding logic shares the same stream.rs files as the decoders but operates in reverse, serializing the IR into provider-specific SSE events. For instance, an internal chunk containing a text delta becomes an Anthropic content_block_delta event or an OpenAI delta event depending on the target format.

Stream Termination and Cleanup

When the upstream stream ends, engine.finish_stream emits required terminal events such as final usage summaries, response.incomplete flags, or stop reason declarations. This ensures downstream clients receive well-formed stream terminations even when the upstream provider's final frame differs from the downstream protocol.

Handling Edge Cases in Live Translation

Switchyard's streaming pipeline includes specific logic for preserving fidelity across provider mismatches. The test suite in crates/switchyard-translation/tests/stream_translation.rs validates these behaviors.

Provider Error Frames

When an upstream provider returns an error frame (e.g., OpenAI's {"error":{…}}), the decoder parses it into LlmResponseChunk::StreamError. The translation engine then re-encodes this as a proper error event conforming to the downstream format, ensuring clients receive coherent error signals regardless of the upstream source.

Partial Usage Information

If a provider sends incomplete usage data—such as missing total_tokens—the engine retains cached token counts in the StreamTranslationState and emits corrected usage objects downstream. This prevents clients from receiving partial or inconsistent billing information during cross-provider translation.

Stop Reasons and Extended Metadata

Switchyard preserves stop reasons (e.g., Anthropic's stop_reason or OpenAI's finish_reason) across translations. The engine ensures downstream clients receive the correct stop tokens or response.incomplete flags even when the upstream and downstream formats use different terminologies for stream completion.

Additionally, extended metadata like reasoning content and cache details are copied into the target format when possible. Both OpenAI-Chat and OpenAI-Responses expose detailed usage and reasoning objects; the translation engine maps these fields into the appropriate downstream schema during encoding.

Cross-Provider Streaming Examples

The following examples demonstrate how to leverage Switchyard's streaming translation in both high-level Python clients and low-level Rust implementations.

Python Client: Streaming from OpenAI to Anthropic Format

This Python example requests a streaming completion from an OpenAI upstream and receives the response in Anthropic-compatible format:

import switchyard
import asyncio

async def main():
    # Create a client that talks to the native Switchyard server.

    client = await switchyard.AsyncClient.from_env()

    # Send a request that asks for streaming output.

    request = {
        "model": "gpt-4o-mini",
        "messages": [{"role": "user", "content": "Tell me a joke"}],
        "stream": True,                      # Ask upstream to stream.

    }

    # The client will translate the OpenAI‑Chat stream into the

    # Anthropic‑Messages format because we ask for it explicitly.

    async for chunk in client.chat_completion(
        request,
        target_format="anthropic_messages",   # ← downstream wire format

    ):
        print(chunk)   # Each chunk is a dict conforming to Anthropic’s schema.

asyncio.run(main())

Rust: Low-Level Translation of Streaming Events

For direct control over the translation pipeline, use the Rust TranslationEngine API:

use switchyard_translation::{
    engine::TranslationEngine,
    stream::{StreamTranslationState, WireFormat},
    codecs::openai_chat::stream::decode_stream_event,
};

fn translate_openai_to_anthropic(event: serde_json::Value) {
    let mut state = StreamTranslationState::new("my-stream", WireFormat::OpenAiChat);
    let mut engine = TranslationEngine::new();

    // Decode the OpenAI event into the internal chunk.
    let chunk = decode_stream_event(&mut state, WireFormat::OpenAiChat, &event).unwrap();

    // Re‑encode the same chunk as an Anthropic event.
    let anthro_events = engine
        .encode_stream_event(&mut state, WireFormat::AnthropicMessages, chunk)
        .unwrap();

    // `anthro_events` now contains the JSON representation expected by
    // Anthropic’s streaming endpoint.
    println!("{}", serde_json::to_string_pretty(&anthro_events).unwrap());
}

Summary

Switchyard enables real-time streaming translation between LLM APIs through a provider-agnostic intermediate representation. Key architectural elements include:

  • LlmResponseChunk—the internal enum defined in crates/switchyard-translation/src/stream.rs that normalizes streaming events from any provider.
  • Provider-specific codecs in crates/switchyard-translation/src/codecs/<provider>/stream.rs that handle parsing and serialization.
  • TranslationEngine in crates/switchyard-translation/src/engine.rs that orchestrates metadata preservation and format conversion.
  • Stateless per-request design using StreamTranslationState to avoid buffering complete responses.
  • Robust edge-case handling for error frames, partial usage data, stop reasons, and reasoning content.

Frequently Asked Questions

How does Switchyard handle provider-specific error messages during streaming?

Switchyard parses upstream error frames into LlmResponseChunk::StreamError variants and re-encodes them into the downstream provider's error format. This ensures clients receive coherent error events even when translating between APIs with incompatible error schemas, as implemented in the translation engine's error handling logic.

Can Switchyard translate a stream from Anthropic to OpenAI format without buffering the entire response?

Yes. The translation pipeline is stateless per-request; the StreamTranslationState tracks only minimal context like model identity and accumulated usage. Because the TranslationEngine processes individual LlmResponseChunk objects incrementally, it can forward Anthropic streaming events to OpenAI-compatible clients token-by-token without storing the complete response.

What happens to usage metadata when the upstream provider sends incomplete token counts?

The TranslationEngine caches partial usage information in the StreamTranslationState and emits corrected usage objects downstream. If the upstream stream lacks total_tokens or other fields, Switchyard retains the last known values and ensures the downstream client receives complete billing information in the final stream event.

Does Switchyard preserve extended metadata like reasoning content across translations?

Yes. When translating between formats that support extended metadata—such as OpenAI's reasoning fields or Anthropic's thinking content—the engine copies these details into the target schema. The test suite in crates/switchyard-translation/tests/stream_translation.rs validates that reasoning and cache details survive round-trip translation between supported wire formats.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →