How Switchyard's Protocol Translation Layer Converts Between OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages
Switchyard's protocol translation layer normalizes OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages into a vendor-agnostic intermediate representation (IR) using codec-based streaming and buffered converters, then re-encodes responses back to the target provider format.
The NVIDIA-NeMo/Switchyard repository implements a high-performance LLM routing proxy that abstracts disparate vendor APIs. The protocol translation layer, located in crates/switchyard-translation, bridges OpenAI and Anthropic schemas by bidirectionally converting their JSON payloads through a unified intermediate representation.
Codec Architecture and Intermediate Representation
The translation system is built around a common Codec trait defined in crates/switchyard-translation/src/lib.rs. This abstraction allows the Engine (src/engine.rs) to register multiple vendor-specific codecs and select the appropriate implementation based on request headers or configuration policies.
All codecs transform external data into intermediate representation (IR) types:
LlmRequest– Normalized incoming request structureLlmResponseChunk– Incremental response fragments for streamingLlmResponse– Complete response payload for buffered modes
Decoding Vendor Formats into IR
The decoding path converts provider-specific JSON into the IR. Switchyard implements both streaming (Server-Sent Events) and buffered (full payload) decoders.
Streaming Decoding
For real-time SSE streams, codecs process discrete events:
- OpenAI Chat Completions and Responses: The
decode_openai_chat_streamfunction incodecs/openai_chat/stream.rsparses SSE values to extractrole,content, and tool calls, emittingLlmResponseChunkstructs. Both the Chat Completions and Responses formats share this streaming logic. - Anthropic Messages: The
decode_anthropic_streamfunction incodecs/anthropic/stream.rshandlescontent_block_start,content_block_delta, andcontent_block_stopevents, mapping them to the sameLlmResponseChunkshape.
Buffered Decoding
For standard HTTP POST requests, buffered codecs parse complete JSON objects:
codecs/openai_chat/buffered.rsimplements full-body parsing for OpenAI Chat Completions and Responses formats.codecs/anthropic/buffered.rshandles Anthropic'smessagesschema, including tool-use ID normalization viasanitize_anthropic_tool_use_idanddesanitize_anthropic_tool_use_idlocated insrc/util.rs.
Encoding from IR to Vendor Outputs
After routing logic processes the IR, codecs encode the result back to the provider's native format.
Streaming Encoding
Streaming responses generate SSE event sequences:
- OpenAI:
encode_openai_chat_streamconstructs{"role":"assistant","content":"..."}fragments, whilefinish_openai_chat_streamappendsstop_reasonand usage statistics. - Anthropic:
encode_anthropic_streambuilds{"type":"assistant","content":"...","id":...}events, withfinish_anthropic_streamadding termination markers.
Buffered Encoding
Buffered paths assemble single JSON objects matching the provider schema:
codecs/openai_chat/buffered.rsreconstructs Chat Completions responses.codecs/anthropic/buffered.rsrebuilds Anthropic Messages format, preserving metadata through round-trip conversion.
Preserving Extensions and Metadata
Switchyard ensures user-provided extensions survive translation. The helper copy_openai_chat_request_extensions (and analogous Anthropic helpers) copies custom fields into the IR using the PRESERVATION_METADATA_KEY constant defined in src/lib.rs. During encoding, these values are restored to the output JSON, ensuring no loss of caller-supplied metadata.
Policy Enforcement and Diagnostics
Each codec receives a Policy and mutable Diagnostic structure. The Policy can restrict features (e.g., disabling Anthropic tool usage), while Diagnostic collects warnings—such as noting Anthropic's optional done marker in src/sse.rs—allowing Switchyard to enforce safety constraints while maintaining format fidelity.
Implementation Examples
The following Rust snippets demonstrate translation workflows using the public API:
use switchyard_translation::codecs::{openai_chat, anthropic};
// OpenAI Chat Completions → IR (Buffered)
let request_body = std::fs::read_to_string("request.json")?;
let codec = openai_chat::OpenAiChatBufferedCodec::new();
let ir_request = codec.decode(&request_body)?;
// IR → Anthropic Messages (Streaming)
let mut state = StreamTranslationState::new();
for sse_event in incoming_sse_stream {
let chunks = anthropic::decode_anthropic_stream(&mut state, &sse_event)?;
// Process LlmResponseChunk items...
}
let final_events = anthropic::finish_anthropic_stream(&mut state)?;
// Tool ID sanitization for Anthropic compatibility
use switchyard_translation::util::{sanitize_anthropic_tool_use_id, desanitize_anthropic_tool_use_id};
let safe_id = sanitize_anthropic_tool_use_id("tool-123");
let original = desanitize_anthropic_tool_use_id(&safe_id);
Summary
- The protocol translation layer resides in
crates/switchyard-translationand implements a codec pattern for bidirectional conversion. - Streaming codecs (
stream.rsmodules) handle Server-Sent Events for real-time responses, while buffered codecs (buffered.rsmodules) process complete JSON payloads. - The intermediate representation (
LlmRequest,LlmResponse,LlmResponseChunk) decouples routing logic from vendor specifics. - Extension preservation via
PRESERVATION_METADATA_KEYensures custom metadata survives round-trip translation. - Policy and Diagnostic hooks in the Engine (
src/engine.rs) enable feature restrictions and observability without breaking format compatibility.
Frequently Asked Questions
How does Switchyard determine which codec to use for incoming requests?
The Engine (src/engine.rs) inspects request headers—typically the Content-Type or a provider-specific header—and matches them against registered codecs in src/lib.rs. Configuration policies in the Switchyard TOML can also explicitly force a specific codec target, overriding automatic detection.
What is the difference between streaming and buffered translation in Switchyard?
Streaming translation processes Server-Sent Events (SSE) incrementally using functions like decode_openai_chat_stream and encode_anthropic_stream, minimizing latency for real-time applications. Buffered translation parses entire JSON bodies via the buffered.rs modules, suitable for standard request-response cycles where the full payload is available immediately.
How does Switchyard handle tool-use identifiers between OpenAI and Anthropic formats?
When translating Anthropic requests, Switchyard normalizes tool-use IDs using sanitize_anthropic_tool_use_id and reverses the transformation with desanitize_anthropic_tool_use_id in src/util.rs. This ensures compatibility between OpenAI's string-based tool references and Anthropic's ID requirements during buffered encoding and decoding.
Where is the intermediate representation defined in the Switchyard source code?
The core IR types—LlmRequest, LlmResponse, and LlmResponseChunk—are defined in crates/switchyard-translation/src/lib.rs. These structures provide the vendor-neutral data model that all codecs target during translation, enabling the routing algorithm to operate independently of external API schemas.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →