# How Switchyard's Protocol Translation Layer Converts Between OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages

> Discover how Switchyard's protocol translation layer converts OpenAI Chat Completions, Responses, and Anthropic Messages to a vendor-agnostic format using codecs and re-encodes them for target providers.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-13

---

**Switchyard's protocol translation layer normalizes OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages into a vendor-agnostic intermediate representation (IR) using codec-based streaming and buffered converters, then re-encodes responses back to the target provider format.**

The NVIDIA-NeMo/Switchyard repository implements a high-performance LLM routing proxy that abstracts disparate vendor APIs. The **protocol translation layer**, located in `crates/switchyard-translation`, bridges OpenAI and Anthropic schemas by bidirectionally converting their JSON payloads through a unified intermediate representation.

## Codec Architecture and Intermediate Representation

The translation system is built around a common **Codec** trait defined in [`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs). This abstraction allows the **Engine** ([`src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/engine.rs)) to register multiple vendor-specific codecs and select the appropriate implementation based on request headers or configuration policies.

All codecs transform external data into **intermediate representation (IR)** types:

- **`LlmRequest`** – Normalized incoming request structure
- **`LlmResponseChunk`** – Incremental response fragments for streaming
- **`LlmResponse`** – Complete response payload for buffered modes

## Decoding Vendor Formats into IR

The decoding path converts provider-specific JSON into the IR. Switchyard implements both streaming (Server-Sent Events) and buffered (full payload) decoders.

### Streaming Decoding

For real-time SSE streams, codecs process discrete events:

- **OpenAI Chat Completions and Responses**: The `decode_openai_chat_stream` function in [`codecs/openai_chat/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/codecs/openai_chat/stream.rs) parses SSE values to extract `role`, `content`, and tool calls, emitting `LlmResponseChunk` structs. Both the Chat Completions and Responses formats share this streaming logic.
- **Anthropic Messages**: The `decode_anthropic_stream` function in [`codecs/anthropic/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/codecs/anthropic/stream.rs) handles `content_block_start`, `content_block_delta`, and `content_block_stop` events, mapping them to the same `LlmResponseChunk` shape.

### Buffered Decoding

For standard HTTP POST requests, buffered codecs parse complete JSON objects:

- [`codecs/openai_chat/buffered.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/codecs/openai_chat/buffered.rs) implements full-body parsing for OpenAI Chat Completions and Responses formats.
- [`codecs/anthropic/buffered.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/codecs/anthropic/buffered.rs) handles Anthropic's `messages` schema, including tool-use ID normalization via `sanitize_anthropic_tool_use_id` and `desanitize_anthropic_tool_use_id` located in [`src/util.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/util.rs).

## Encoding from IR to Vendor Outputs

After routing logic processes the IR, codecs encode the result back to the provider's native format.

### Streaming Encoding

Streaming responses generate SSE event sequences:

- **OpenAI**: `encode_openai_chat_stream` constructs `{"role":"assistant","content":"..."}` fragments, while `finish_openai_chat_stream` appends `stop_reason` and usage statistics.
- **Anthropic**: `encode_anthropic_stream` builds `{"type":"assistant","content":"...","id":...}` events, with `finish_anthropic_stream` adding termination markers.

### Buffered Encoding

Buffered paths assemble single JSON objects matching the provider schema:

- [`codecs/openai_chat/buffered.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/codecs/openai_chat/buffered.rs) reconstructs Chat Completions responses.
- [`codecs/anthropic/buffered.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/codecs/anthropic/buffered.rs) rebuilds Anthropic Messages format, preserving metadata through round-trip conversion.

## Preserving Extensions and Metadata

Switchyard ensures user-provided extensions survive translation. The helper `copy_openai_chat_request_extensions` (and analogous Anthropic helpers) copies custom fields into the IR using the `PRESERVATION_METADATA_KEY` constant defined in [`src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/lib.rs). During encoding, these values are restored to the output JSON, ensuring no loss of caller-supplied metadata.

## Policy Enforcement and Diagnostics

Each codec receives a **Policy** and mutable **Diagnostic** structure. The `Policy` can restrict features (e.g., disabling Anthropic tool usage), while `Diagnostic` collects warnings—such as noting Anthropic's optional `done` marker in [`src/sse.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/sse.rs)—allowing Switchyard to enforce safety constraints while maintaining format fidelity.

## Implementation Examples

The following Rust snippets demonstrate translation workflows using the public API:

```rust
use switchyard_translation::codecs::{openai_chat, anthropic};

// OpenAI Chat Completions → IR (Buffered)
let request_body = std::fs::read_to_string("request.json")?;
let codec = openai_chat::OpenAiChatBufferedCodec::new();
let ir_request = codec.decode(&request_body)?;

```

```rust
// IR → Anthropic Messages (Streaming)
let mut state = StreamTranslationState::new();
for sse_event in incoming_sse_stream {
    let chunks = anthropic::decode_anthropic_stream(&mut state, &sse_event)?;
    // Process LlmResponseChunk items...
}
let final_events = anthropic::finish_anthropic_stream(&mut state)?;

```

```rust
// Tool ID sanitization for Anthropic compatibility
use switchyard_translation::util::{sanitize_anthropic_tool_use_id, desanitize_anthropic_tool_use_id};

let safe_id = sanitize_anthropic_tool_use_id("tool-123");
let original = desanitize_anthropic_tool_use_id(&safe_id);

```

## Summary

- The **protocol translation layer** resides in `crates/switchyard-translation` and implements a codec pattern for bidirectional conversion.
- **Streaming codecs** ([`stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/stream.rs) modules) handle Server-Sent Events for real-time responses, while **buffered codecs** ([`buffered.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/buffered.rs) modules) process complete JSON payloads.
- The **intermediate representation** (`LlmRequest`, `LlmResponse`, `LlmResponseChunk`) decouples routing logic from vendor specifics.
- **Extension preservation** via `PRESERVATION_METADATA_KEY` ensures custom metadata survives round-trip translation.
- **Policy and Diagnostic hooks** in the Engine ([`src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/engine.rs)) enable feature restrictions and observability without breaking format compatibility.

## Frequently Asked Questions

### How does Switchyard determine which codec to use for incoming requests?

The **Engine** ([`src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/engine.rs)) inspects request headers—typically the `Content-Type` or a provider-specific header—and matches them against registered codecs in [`src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/lib.rs). Configuration policies in the Switchyard TOML can also explicitly force a specific codec target, overriding automatic detection.

### What is the difference between streaming and buffered translation in Switchyard?

**Streaming** translation processes Server-Sent Events (SSE) incrementally using functions like `decode_openai_chat_stream` and `encode_anthropic_stream`, minimizing latency for real-time applications. **Buffered** translation parses entire JSON bodies via the [`buffered.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/buffered.rs) modules, suitable for standard request-response cycles where the full payload is available immediately.

### How does Switchyard handle tool-use identifiers between OpenAI and Anthropic formats?

When translating Anthropic requests, Switchyard normalizes tool-use IDs using `sanitize_anthropic_tool_use_id` and reverses the transformation with `desanitize_anthropic_tool_use_id` in [`src/util.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/util.rs). This ensures compatibility between OpenAI's string-based tool references and Anthropic's ID requirements during buffered encoding and decoding.

### Where is the intermediate representation defined in the Switchyard source code?

The core IR types—**`LlmRequest`**, **`LlmResponse`**, and **`LlmResponseChunk`**—are defined in [`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs). These structures provide the vendor-neutral data model that all codecs target during translation, enabling the routing algorithm to operate independently of external API schemas.