How Switchyard Handles Bidirectional Translation for Streaming and Buffered LLM Responses
Switchyard decodes any supported LLM wire format into a neutral intermediate representation (IR) and re-encodes it into any other supported format, handling both streaming chunks and complete buffered responses through a registry-based engine that maintains per-stream state.
NVIDIA Switchyard acts as a translation layer between client-side LLM SDKs and upstream providers like OpenAI and Anthropic. It enables bidirectional translation for streaming and buffered LLM responses by converting provider-specific JSON into a format-agnostic IR, then serializing it to the target wire format using codec registries and stateful stream tracking.
The Translation Pipeline
Switchyard’s engine operates in three distinct phases for every conversion. The same components handle both streaming (chunk-by-chunk) and buffered (complete payload) modes:
| Direction | Core Component | Function |
|---|---|---|
| Decode (source → IR) | TranslationEngine::decode_* (buffered) or decode_stream_event (streaming) |
Parses provider JSON into neutral types (LlmRequest, AggLlmResponse, LlmResponseChunk). |
| Encode (IR → target) | TranslationEngine::encode_* (buffered) or encode_stream_event (streaming) |
Serializes the neutral IR back into the target wire format. |
| Finish (stream end) | finish_stream |
Emits required terminal chunks when the source stream ends without an explicit stop event. |
This pipeline is orchestrated by the TranslationEngine in [engine.rs](crates/switchyard-translation/src/engine.rs), which holds two registries: FormatRegistry for buffered codecs and StreamCodecRegistry for streaming codecs.
Core Architectural Components
The TranslationEngine Hub
Located in [engine.rs](crates/switchyard-translation/src/engine.rs), the engine provides the primary API for both streaming and buffered operations. It is stateless except for the registries it holds, ensuring thread safety across concurrent translations.
Key public methods include:
pub fn translate_event(
&self,
state: &mut StreamTranslationState,
source: impl Into<FormatId>,
target: impl Into<FormatId>,
event: &Value,
) -> Result<Vec<Value>>; // streaming
pub fn translate_response(
&self,
source: impl Into<FormatId>,
target: impl Into<FormatId>,
body: &Value,
policy: &TranslationPolicy,
) -> Result<TranslationOutput>; // buffered
The engine maps format IDs to specific codec implementations using the registries, invoking the appropriate parser for the source and serializer for the target.
Streaming Codecs and State Management
Each streaming codec implements the StreamCodec trait defined in [codecs/mod.rs](crates/switchyard-translation/src/codecs/mod.rs):
pub trait StreamCodec {
fn format(&self) -> FormatId;
fn decode_event(&self, state: &mut StreamTranslationState, event: &Value)
-> Vec<LlmResponseChunk>;
fn encode_event(&self, state: &mut StreamTranslationState, event: LlmResponseChunk)
-> Vec<Value>;
fn finish(&self, state: &mut StreamTranslationState) -> Vec<Value>;
}
Key implementations include:
- OpenAI Chat streaming: [
openai_chat/stream.rs](crates/switchyard-translation/src/codecs/openai_chat/stream.rs) handleschat.completion.chunk - Anthropic Messages streaming: [
anthropic/stream.rs](crates/switchyard-translation/src/codecs/anthropic/stream.rs) handlesmessage_start,content_block_*, andmessage_delta - OpenAI Responses streaming: [
responses/stream.rs](crates/switchyard-translation/src/codecs/responses/stream.rs) handlesresponse.createdandresponse.output_text.delta
These codecs share per-stream state via StreamTranslationState ([stream.rs](crates/switchyard-translation/src/stream.rs)), which tracks:
sourceandtargetformat IDsmessage_id,model, and usage counters- Flags like
saw_message_start,finished, anderrored
This state enables bidirectional translation: after decoding a source event into a neutral LlmResponseChunk, the engine passes that chunk to the target codec’s encode_event, which may belong to a different format entirely.
Buffered Response Handling
Buffered codecs operate on complete JSON bodies and implement the FormatCodec trait. Registered in FormatRegistry, they use the same TranslationEngine workflow (decode_response → encode_response) as the streaming path.
The buffered path guarantees that the output of a buffered response is identical to re-assembling all chunks from a streaming response of the same data. This is verified by the test responses_buffered_and_streamed_outputs_match in the test suite.
Source Preservation and Cross-Format Encoding
When source and target formats match, decode_stream_event returns an LlmResponseStreamEvent containing the preserved raw JSON (event field). The engine replays this verbatim, avoiding unnecessary re-encoding. When formats differ, the engine discards provider-specific fields (e.g., system_fingerprint) and re-encodes using only the normalized chunks, ensuring semantic fidelity.
Handling Stream Termination
When an upstream provider closes a stream without an explicit stop signal, the engine’s finish_stream method invokes the target codec’s finish implementation. For OpenAI Chat, this synthesizes a terminal chunk (finish_openai_chat_stream) containing the appropriate finish_reason and usage payload, ensuring the downstream client receives a well-formed termination regardless of how the upstream ended.
Implementation Examples
Streaming Translation Example
The following demonstrates translating an OpenAI Chat streaming chunk into Anthropic Messages format:
use switchyard_translation::{
TranslationEngine,
StreamTranslationState,
WireFormat,
};
use serde_json::json;
fn stream_example() -> anyhow::Result<()> {
// Initialize the engine with built-in codecs
let engine = TranslationEngine::default();
// Set up per-stream state
let mut state = StreamTranslationState::new(
WireFormat::OpenAiChat,
WireFormat::AnthropicMessages,
);
// Simulate receiving an OpenAI chunk
let openai_chunk = json!({
"id": "chatcmpl-123",
"object": "chat.completion.chunk",
"model": "gpt-4o",
"choices": [{
"index": 0,
"delta": {"content": "Hello"},
"finish_reason": null
}]
});
// Decode → neutral IR
let preserved = engine.decode_stream_event(
&mut state,
WireFormat::OpenAiChat,
openai_chunk
)?;
// Encode → Anthropic format
let anthro_events = engine.encode_stream_event(
&mut state,
WireFormat::AnthropicMessages,
preserved
)?;
// Flush terminal message when stream ends
let final_events = engine.finish_stream(
&mut state,
WireFormat::AnthropicMessages
)?;
Ok(())
}
Buffered Response Example
For complete responses, use the buffered API:
use switchyard_translation::{
TranslationEngine,
TranslationPolicy,
WireFormat,
};
use serde_json::json;
fn buffered_example() -> anyhow::Result<()> {
let engine = TranslationEngine::default();
let policy = TranslationPolicy::default();
let openai_body = json!({
"id": "resp-1",
"object": "response",
"model": "gpt-4o",
"status": "completed",
"output": [{
"type": "message",
"role": "assistant",
"content": [{"type": "output_text", "text": "Hello"}]
}]
});
// Decode → neutral IR
let decoded = engine.decode_response(
WireFormat::OpenAiResponses,
&openai_body,
&policy,
)?;
// Encode → Anthropic format
let encoded = engine.encode_response(
WireFormat::AnthropicMessages,
&decoded.response,
&policy,
)?;
println!("{}", serde_json::to_string_pretty(&encoded.body)?);
Ok(())
}
Summary
- Stateless Engine:
TranslationEnginecontains only codec registries and contains no per-request state. - Stateful Tracking:
StreamTranslationStatetracks identifiers, usage, and terminal event status across chunks. - Unified IR: All codecs translate to/from
LlmResponseChunkandAggLlmResponse, enabling arbitrary source-to-target combinations. - Format Preservation: When source and target match, original JSON is replayed verbatim; when they differ, normalized chunks ensure semantic accuracy.
- Termination Guarantees: The
finish_streammethod ensures proper terminal chunks are emitted even when upstream closes silently.
Frequently Asked Questions
How does Switchyard maintain state across streaming chunks?
Switchyard uses a StreamTranslationState object that lives on the caller side (Python client, proxy, or server). This state is passed to every decode_stream_event and encode_stream_event call, allowing the engine to remember message_id, model, usage statistics, and flags like saw_message_start across the lifetime of a single client connection.
What happens when source and target formats are identical?
When the source and target formats match, the decode_stream_event method returns the preserved raw JSON alongside the normalized chunks. The engine detects this match and replays the original provider JSON verbatim, avoiding unnecessary encoding overhead and preserving provider-specific fields like system_fingerprint that might not exist in the neutral IR.
How does the engine handle stream termination without explicit stop events?
The engine exposes a finish_stream method that must be called when the upstream connection closes. This invokes the target codec’s finish method (e.g., finish_openai_chat_stream), which synthesizes the required terminal chunk with the appropriate finish_reason and optional usage payload, ensuring downstream clients receive a well-formed stream termination regardless of upstream behavior.
Can Switchyard translate between any supported formats bidirectionally?
Yes. Because all codecs convert to the neutral intermediate representation (LlmResponseChunk for streaming, AggLlmResponse for buffered), the engine can route any supported source format to any supported target format. The IR is format-agnostic, making the translation pipeline reversible and allowing cross-format conversion between OpenAI Chat, Anthropic Messages, and the generic "Responses" schema in either direction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →