How Switchyard's IR Translation Layer Normalizes OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages

Switchyard converts OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages into a unified internal representation called LlmResponseChunk using provider-specific streaming codecs that implement the StreamCodec trait for bidirectional translation.

The NVIDIA-NeMo/Switchyard project provides a vendor-agnostic routing layer for large language model APIs. Its IR translation layer enables seamless interoperability between different provider formats by normalizing all incoming streams into a single intermediate representation before processing or re-encoding.

The StreamCodec Trait and Translation Interface

The translation layer resides in crates/switchyard-translation/src/codecs/ where each codec implements the StreamCodec trait. This trait defines four methods that enable bidirectional streaming translation:

  • decode_event – Converts provider-specific JSON chunks into one or more LlmResponseChunk events
  • encode_event – Converts neutral LlmResponseChunk instances back into provider JSON formats
  • observe_replayed_event – Updates internal state when replaying raw events (Anthropic-specific)
  • finish – Emits missing terminal events when source streams terminate

All codecs operate on a shared StreamTranslationState managed by the translation engine in crates/switchyard-translation/src/engine.rs.

Decoding: Normalizing Vendor Streams into LlmResponseChunk

OpenAI Chat Completions

In crates/switchyard-translation/src/codecs/openai_chat/stream.rs, the decoder extracts model, message id, usage, content, tool calls, reasoning, and stop reason from incoming chunks. Errors convert into LlmResponseChunk::StreamError, while usage fields normalize via openai_usage to expose input_tokens, output_tokens, total_tokens, and optional reasoning_tokens (see lines 92‑127).

OpenAI Responses

The decoder in crates/switchyard-translation/src/codecs/responses/stream.rs mirrors the chat codec logic. While the JSON structure differs—with "choices" containing complete messages rather than incremental deltas—the decoder rewrites payloads into identical LlmResponseChunk variants. This ensures unified handling across both OpenAI endpoints despite structural differences in the raw API responses.

Anthropic Messages

In crates/switchyard-translation/src/codecs/anthropic/stream.rs, the decode_anthropic_stream function maps Anthropic's event-driven protocol into neutral chunks:

  • MessageStart → LlmResponseChunk::MessageStart (captures model and ID)
  • content_block_* events → TextDelta, ReasoningDelta, or ToolCallDelta based on block type
  • message_delta → Captures usage via capture_anthropic_usage and stop reason; emits MessageStop
  • Errors → StreamError

The capture_anthropic_usage function (lines 35‑55) copies raw input_tokens, output_tokens, and cache fields into the standard Usage struct.

Encoding: Denormalizing to Provider Formats

OpenAI Chat and Responses

The encode_openai_chat_stream function builds JSON chat.completion.chunk objects with fields for id, model, choices, usage, and finish_reason. It sanitizes tool-call IDs and maps stop reasons via openai_finish_reason (lines 80‑144). The Responses codec reuses helper functions (openai_stream_chunk, openai_usage_value) to emit a final usage-only chunk after terminal messages.

Anthropic Messages

The encode_anthropic_stream function creates event sequences including message_start, content_block_*, message_delta, and message_stop. It manages block lifecycles for text, reasoning, and tool blocks, ensuring required termination events via finish_anthropic_stream. Stop reason translation occurs through anthropic_stop_reason (lines 92‑103).

Normalizing Tokens and Stop Reasons

Switchyard unifies provider-specific metadata through dedicated normalization functions:

Token Usage

  • OpenAI: openai_usage converts cached-token details and reasoning tokens into the unified Usage struct
  • Anthropic: capture_anthropic_usage copies raw input/output tokens and cache fields into the same Usage struct

Stop Reasons

  • openai_finish_reason maps internal states like "end_turn", "max_tokens", and "tool_use" to OpenAI vocabularies
  • anthropic_stop_reason maps equivalent states to Anthropic vocabularies including "max_tokens", "tool_use", and "refusal"

Message IDs

  • openai_stream_id rewrites upstream IDs into "chatcmpl_…" patterns
  • anthropic_message_id ensures IDs start with "msg_"

Why Vendor-Agnostic Translation Matters

By translating every incoming event into LlmResponseChunk and back, Switchyard achieves three critical capabilities:

  1. Universal Routing – Requests can route to any model regardless of which API format the client uses
  2. Algorithmic Consistency – Routing algorithms operate on unified streams without vendor-specific handling
  3. Standardized Telemetry – Consistent usage statistics reach callers even when underlying providers report metrics in different shapes

Summary

  • Switchyard uses the StreamCodec trait to abstract OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages into LlmResponseChunk
  • Decoding occurs in provider-specific files: openai_chat/stream.rs, responses/stream.rs, and anthropic/stream.rs
  • decode_event normalizes model IDs, usage statistics, content deltas, tool calls, and stop reasons into vendor-neutral variants
  • Encoding functions reverse the process, ensuring downstream services receive properly formatted provider-specific JSON
  • Usage normalization via openai_usage and capture_anthropic_usage exposes consistent input_tokens, output_tokens, and cache metrics
  • The neutral IR enables seamless request routing and analytics across heterogeneous LLM providers

Frequently Asked Questions

What is the internal representation used by Switchyard's translation layer?

Switchyard normalizes all provider formats into LlmResponseChunk, an enum defined in crates/switchyard-translation/src/lib.rs. This intermediate representation captures message metadata, text deltas, reasoning content, tool calls, usage statistics, and error states in a vendor-neutral format.

How does Switchyard handle streaming differences between OpenAI and Anthropic?

OpenAI streams use incremental deltas within a single JSON structure per chunk, while Anthropic uses discrete event types (message_start, content_block_delta, etc.). The respective codecs in openai_chat/stream.rs and anthropic/stream.rs map these disparate patterns into identical LlmResponseChunk sequences, abstracting the structural differences.

What happens to token usage metrics when translating between providers?

Switchyard extracts provider-specific usage fields—such as OpenAI's prompt_tokens or Anthropic's input_tokens—and normalizes them into a unified Usage struct containing input_tokens, output_tokens, total_tokens, and optional cache-related fields. This ensures consistent billing and monitoring regardless of the upstream provider.

Why does the Anthropic codec require observe_replayed_event?

The Anthropic Messages API uses complex block-level lifecycle events that require state tracking when replaying or reconstructing streams. The observe_replayed_event method updates the StreamTranslationState to maintain accurate block indices and content boundaries during stream manipulation, whereas OpenAI's simpler delta-based protocol does not require this additional state management.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →