How Switchyard's IR Translation Layer Normalizes OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages
Switchyard converts OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages into a unified internal representation called LlmResponseChunk using provider-specific streaming codecs that implement the StreamCodec trait for bidirectional translation.
The NVIDIA-NeMo/Switchyard project provides a vendor-agnostic routing layer for large language model APIs. Its IR translation layer enables seamless interoperability between different provider formats by normalizing all incoming streams into a single intermediate representation before processing or re-encoding.
The StreamCodec Trait and Translation Interface
The translation layer resides in crates/switchyard-translation/src/codecs/ where each codec implements the StreamCodec trait. This trait defines four methods that enable bidirectional streaming translation:
decode_event– Converts provider-specific JSON chunks into one or moreLlmResponseChunkeventsencode_event– Converts neutralLlmResponseChunkinstances back into provider JSON formatsobserve_replayed_event– Updates internal state when replaying raw events (Anthropic-specific)finish– Emits missing terminal events when source streams terminate
All codecs operate on a shared StreamTranslationState managed by the translation engine in crates/switchyard-translation/src/engine.rs.
Decoding: Normalizing Vendor Streams into LlmResponseChunk
OpenAI Chat Completions
In crates/switchyard-translation/src/codecs/openai_chat/stream.rs, the decoder extracts model, message id, usage, content, tool calls, reasoning, and stop reason from incoming chunks. Errors convert into LlmResponseChunk::StreamError, while usage fields normalize via openai_usage to expose input_tokens, output_tokens, total_tokens, and optional reasoning_tokens (see lines 92‑127).
OpenAI Responses
The decoder in crates/switchyard-translation/src/codecs/responses/stream.rs mirrors the chat codec logic. While the JSON structure differs—with "choices" containing complete messages rather than incremental deltas—the decoder rewrites payloads into identical LlmResponseChunk variants. This ensures unified handling across both OpenAI endpoints despite structural differences in the raw API responses.
Anthropic Messages
In crates/switchyard-translation/src/codecs/anthropic/stream.rs, the decode_anthropic_stream function maps Anthropic's event-driven protocol into neutral chunks:
MessageStart→LlmResponseChunk::MessageStart(captures model and ID)content_block_*events →TextDelta,ReasoningDelta, orToolCallDeltabased on block typemessage_delta→ Captures usage viacapture_anthropic_usageand stop reason; emitsMessageStop- Errors →
StreamError
The capture_anthropic_usage function (lines 35‑55) copies raw input_tokens, output_tokens, and cache fields into the standard Usage struct.
Encoding: Denormalizing to Provider Formats
OpenAI Chat and Responses
The encode_openai_chat_stream function builds JSON chat.completion.chunk objects with fields for id, model, choices, usage, and finish_reason. It sanitizes tool-call IDs and maps stop reasons via openai_finish_reason (lines 80‑144). The Responses codec reuses helper functions (openai_stream_chunk, openai_usage_value) to emit a final usage-only chunk after terminal messages.
Anthropic Messages
The encode_anthropic_stream function creates event sequences including message_start, content_block_*, message_delta, and message_stop. It manages block lifecycles for text, reasoning, and tool blocks, ensuring required termination events via finish_anthropic_stream. Stop reason translation occurs through anthropic_stop_reason (lines 92‑103).
Normalizing Tokens and Stop Reasons
Switchyard unifies provider-specific metadata through dedicated normalization functions:
Token Usage
- OpenAI:
openai_usageconverts cached-token details and reasoning tokens into the unifiedUsagestruct - Anthropic:
capture_anthropic_usagecopies raw input/output tokens and cache fields into the sameUsagestruct
Stop Reasons
openai_finish_reasonmaps internal states like"end_turn","max_tokens", and"tool_use"to OpenAI vocabulariesanthropic_stop_reasonmaps equivalent states to Anthropic vocabularies including"max_tokens","tool_use", and"refusal"
Message IDs
openai_stream_idrewrites upstream IDs into"chatcmpl_…"patternsanthropic_message_idensures IDs start with"msg_"
Why Vendor-Agnostic Translation Matters
By translating every incoming event into LlmResponseChunk and back, Switchyard achieves three critical capabilities:
- Universal Routing – Requests can route to any model regardless of which API format the client uses
- Algorithmic Consistency – Routing algorithms operate on unified streams without vendor-specific handling
- Standardized Telemetry – Consistent usage statistics reach callers even when underlying providers report metrics in different shapes
Summary
- Switchyard uses the
StreamCodectrait to abstract OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages intoLlmResponseChunk - Decoding occurs in provider-specific files:
openai_chat/stream.rs,responses/stream.rs, andanthropic/stream.rs decode_eventnormalizes model IDs, usage statistics, content deltas, tool calls, and stop reasons into vendor-neutral variants- Encoding functions reverse the process, ensuring downstream services receive properly formatted provider-specific JSON
- Usage normalization via
openai_usageandcapture_anthropic_usageexposes consistentinput_tokens,output_tokens, and cache metrics - The neutral IR enables seamless request routing and analytics across heterogeneous LLM providers
Frequently Asked Questions
What is the internal representation used by Switchyard's translation layer?
Switchyard normalizes all provider formats into LlmResponseChunk, an enum defined in crates/switchyard-translation/src/lib.rs. This intermediate representation captures message metadata, text deltas, reasoning content, tool calls, usage statistics, and error states in a vendor-neutral format.
How does Switchyard handle streaming differences between OpenAI and Anthropic?
OpenAI streams use incremental deltas within a single JSON structure per chunk, while Anthropic uses discrete event types (message_start, content_block_delta, etc.). The respective codecs in openai_chat/stream.rs and anthropic/stream.rs map these disparate patterns into identical LlmResponseChunk sequences, abstracting the structural differences.
What happens to token usage metrics when translating between providers?
Switchyard extracts provider-specific usage fields—such as OpenAI's prompt_tokens or Anthropic's input_tokens—and normalizes them into a unified Usage struct containing input_tokens, output_tokens, total_tokens, and optional cache-related fields. This ensures consistent billing and monitoring regardless of the upstream provider.
Why does the Anthropic codec require observe_replayed_event?
The Anthropic Messages API uses complex block-level lifecycle events that require state tracking when replaying or reconstructing streams. The observe_replayed_event method updates the StreamTranslationState to maintain accurate block indices and content boundaries during stream manipulation, whereas OpenAI's simpler delta-based protocol does not require this additional state management.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →