# How Switchyard's IR Translation Layer Normalizes OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages

> Learn how Switchyard's IR translation layer normalizes OpenAI Chat Completions and Anthropic Messages into a unified format, enabling seamless integration and efficient processing of LLM outputs.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: internals
- Published: 2026-09-12

---

**Switchyard converts OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages into a unified internal representation called `LlmResponseChunk` using provider-specific streaming codecs that implement the `StreamCodec` trait for bidirectional translation.**

The NVIDIA-NeMo/Switchyard project provides a vendor-agnostic routing layer for large language model APIs. Its IR translation layer enables seamless interoperability between different provider formats by normalizing all incoming streams into a single intermediate representation before processing or re-encoding.

## The `StreamCodec` Trait and Translation Interface

The translation layer resides in `crates/switchyard-translation/src/codecs/` where each codec implements the **`StreamCodec`** trait. This trait defines four methods that enable bidirectional streaming translation:

- **`decode_event`** – Converts provider-specific JSON chunks into one or more `LlmResponseChunk` events
- **`encode_event`** – Converts neutral `LlmResponseChunk` instances back into provider JSON formats
- **`observe_replayed_event`** – Updates internal state when replaying raw events (Anthropic-specific)
- **`finish`** – Emits missing terminal events when source streams terminate

All codecs operate on a shared **`StreamTranslationState`** managed by the translation engine in [`crates/switchyard-translation/src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/engine.rs).

## Decoding: Normalizing Vendor Streams into `LlmResponseChunk`

### OpenAI Chat Completions

In [`crates/switchyard-translation/src/codecs/openai_chat/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/codecs/openai_chat/stream.rs), the decoder extracts **model**, **message id**, **usage**, **content**, **tool calls**, **reasoning**, and **stop reason** from incoming chunks. Errors convert into `LlmResponseChunk::StreamError`, while usage fields normalize via `openai_usage` to expose `input_tokens`, `output_tokens`, `total_tokens`, and optional `reasoning_tokens` (see lines 92‑127).

### OpenAI Responses

The decoder in [`crates/switchyard-translation/src/codecs/responses/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/codecs/responses/stream.rs) mirrors the chat codec logic. While the JSON structure differs—with `"choices"` containing complete messages rather than incremental deltas—the decoder rewrites payloads into identical `LlmResponseChunk` variants. This ensures unified handling across both OpenAI endpoints despite structural differences in the raw API responses.

### Anthropic Messages

In [`crates/switchyard-translation/src/codecs/anthropic/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/codecs/anthropic/stream.rs), the `decode_anthropic_stream` function maps Anthropic's event-driven protocol into neutral chunks:

- `MessageStart` → `LlmResponseChunk::MessageStart` (captures model and ID)
- `content_block_*` events → `TextDelta`, `ReasoningDelta`, or `ToolCallDelta` based on block type
- `message_delta` → Captures usage via `capture_anthropic_usage` and stop reason; emits `MessageStop`
- Errors → `StreamError`

The `capture_anthropic_usage` function (lines 35‑55) copies raw `input_tokens`, `output_tokens`, and cache fields into the standard `Usage` struct.

## Encoding: Denormalizing to Provider Formats

### OpenAI Chat and Responses

The `encode_openai_chat_stream` function builds JSON `chat.completion.chunk` objects with fields for `id`, `model`, `choices`, `usage`, and `finish_reason`. It sanitizes tool-call IDs and maps stop reasons via `openai_finish_reason` (lines 80‑144). The Responses codec reuses helper functions (`openai_stream_chunk`, `openai_usage_value`) to emit a final usage-only chunk after terminal messages.

### Anthropic Messages

The `encode_anthropic_stream` function creates event sequences including `message_start`, `content_block_*`, `message_delta`, and `message_stop`. It manages block lifecycles for text, reasoning, and tool blocks, ensuring required termination events via `finish_anthropic_stream`. Stop reason translation occurs through `anthropic_stop_reason` (lines 92‑103).

## Normalizing Tokens and Stop Reasons

Switchyard unifies provider-specific metadata through dedicated normalization functions:

**Token Usage**
- **OpenAI**: `openai_usage` converts cached-token details and reasoning tokens into the unified `Usage` struct
- **Anthropic**: `capture_anthropic_usage` copies raw input/output tokens and cache fields into the same `Usage` struct

**Stop Reasons**
- **`openai_finish_reason`** maps internal states like `"end_turn"`, `"max_tokens"`, and `"tool_use"` to OpenAI vocabularies
- **`anthropic_stop_reason`** maps equivalent states to Anthropic vocabularies including `"max_tokens"`, `"tool_use"`, and `"refusal"`

**Message IDs**
- `openai_stream_id` rewrites upstream IDs into `"chatcmpl_…"` patterns
- `anthropic_message_id` ensures IDs start with `"msg_"`

## Why Vendor-Agnostic Translation Matters

By translating every incoming event into `LlmResponseChunk` and back, Switchyard achieves three critical capabilities:

1. **Universal Routing** – Requests can route to any model regardless of which API format the client uses
2. **Algorithmic Consistency** – Routing algorithms operate on unified streams without vendor-specific handling
3. **Standardized Telemetry** – Consistent usage statistics reach callers even when underlying providers report metrics in different shapes

## Summary

- Switchyard uses the **`StreamCodec`** trait to abstract OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages into **`LlmResponseChunk`**
- **Decoding** occurs in provider-specific files: [`openai_chat/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/openai_chat/stream.rs), [`responses/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/responses/stream.rs), and [`anthropic/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/anthropic/stream.rs)
- **`decode_event`** normalizes model IDs, usage statistics, content deltas, tool calls, and stop reasons into vendor-neutral variants
- **Encoding** functions reverse the process, ensuring downstream services receive properly formatted provider-specific JSON
- **Usage normalization** via `openai_usage` and `capture_anthropic_usage` exposes consistent `input_tokens`, `output_tokens`, and cache metrics
- The neutral IR enables seamless request routing and analytics across heterogeneous LLM providers

## Frequently Asked Questions

### What is the internal representation used by Switchyard's translation layer?

Switchyard normalizes all provider formats into **`LlmResponseChunk`**, an enum defined in [`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs). This intermediate representation captures message metadata, text deltas, reasoning content, tool calls, usage statistics, and error states in a vendor-neutral format.

### How does Switchyard handle streaming differences between OpenAI and Anthropic?

OpenAI streams use incremental deltas within a single JSON structure per chunk, while Anthropic uses discrete event types (`message_start`, `content_block_delta`, etc.). The respective codecs in [`openai_chat/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/openai_chat/stream.rs) and [`anthropic/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/anthropic/stream.rs) map these disparate patterns into identical `LlmResponseChunk` sequences, abstracting the structural differences.

### What happens to token usage metrics when translating between providers?

Switchyard extracts provider-specific usage fields—such as OpenAI's `prompt_tokens` or Anthropic's `input_tokens`—and normalizes them into a unified **`Usage`** struct containing `input_tokens`, `output_tokens`, `total_tokens`, and optional cache-related fields. This ensures consistent billing and monitoring regardless of the upstream provider.

### Why does the Anthropic codec require `observe_replayed_event`?

The Anthropic Messages API uses complex block-level lifecycle events that require state tracking when replaying or reconstructing streams. The `observe_replayed_event` method updates the `StreamTranslationState` to maintain accurate block indices and content boundaries during stream manipulation, whereas OpenAI's simpler delta-based protocol does not require this additional state management.