# How Switchyard-Translation Enables API Format Conversion for LLMs

> Discover how Switchyard-translation, a Rust crate, converts API formats for LLMs. Seamlessly route requests between OpenAI, Anthropic, and other backends with bidirectional mapping.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-22

---

**Switchyard-translation is a pure-Rust crate that performs bidirectional mapping between provider-specific HTTP formats (OpenAI, Anthropic, etc.) and Switchyard's neutral conversation IR, enabling seamless routing across heterogeneous LLM backends.**

The `switchyard-translation` crate sits at the boundary between external LLM providers and the NVIDIA-NeMo/Switchyard orchestration layer. It translates disparate API schemas into a unified internal representation (`switchyard_protocol::llm`) and back again, allowing the Python orchestration layer to interact with any supported backend through a single, consistent interface.

## Core Architecture of Switchyard-Translation

The translation engine operates through a trait-based codec system that isolates provider-specific serialization logic from the core routing infrastructure.

### The FormatCodec Trait

At the heart of the system lies the `FormatCodec` trait defined in [`crates/switchyard-translation/src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/engine.rs). This trait mandates four essential operations for every supported provider format:

- **`decode_request`** – Transforms incoming provider JSON into an `LlmRequest`
- **`encode_request`** – Converts an `LlmRequest` into the target provider's wire format
- **`decode_response`** – Parses provider responses into `AggLlmResponse`
- **`encode_response`** – Serializes internal responses back to provider-specific JSON

Each provider implementation registers itself against a `FormatId`, allowing the engine to dynamically select the appropriate codec at runtime.

### Provider Implementations

Concrete codecs reside in `crates/switchyard-translation/src/codecs/`. The OpenAI Chat implementation serves as the reference architecture, split into two specialized variants:

- **`OpenAiChatCodec`** in [`codecs/openai_chat/buffered.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/codecs/openai_chat/buffered.rs) handles complete request/response cycles
- **`OpenAiChatStreamCodec`** in [`codecs/openai_chat/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/codecs/openai_chat/stream.rs) manages Server-Sent Events (SSE) for real-time token streaming

Both variants implement the same `FormatCodec` interface, ensuring consistent behavior whether operating in buffered or streaming modes.

## Bidirectional Translation Workflow

The conversion process preserves semantic meaning while adapting structural differences between provider schemas.

### Decoding Provider Requests

When decoding an OpenAI-compatible payload, the `decode_request` method (lines 38-78 in [`buffered.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/buffered.rs)) orchestrates field-level extraction through helper functions:

- **`decode_openai_content`** – Parses message content blocks, handling text, images, and tool results
- **`decode_openai_tool_call`** – Extracts function calling metadata into the internal tool representation
- **`role_from_openai`** – Maps OpenAI role strings (`system`, `user`, `assistant`, `tool`) to the internal `Role` enum

The method constructs an `LlmRequest` containing the model identifier, message history, sampling parameters, and tool definitions, normalizing provider quirks into the canonical IR.

### Encoding Internal Responses

The `encode_request` method (lines 81-124 in [`buffered.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/buffered.rs)) performs the inverse operation, generating provider-specific JSON from an `LlmRequest` or `AggLlmResponse`. During encoding, the codec applies **deterministic ID policies** to ensure traceability across routing hops, copies extension fields, and implements fallback logic—emitting text representations when the target provider lacks support for specific content blocks like images or files.

### The Preservation Layer

Unknown fields that lack direct IR mappings undergo lossless capture via the preservation mechanism implemented in [`crates/switchyard-translation/src/util.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/util.rs). The functions `capture_request_preservation` and `capture_response_preservation` serialize unrecognized JSON subtrees into a sidecar structure. Later, `embed_preservation` re-inserts these fields verbatim into the encoded output. This guarantees **lossless round-tripping** of provider-specific metadata (custom headers, vendor extensions, or experimental features) even when Switchyard lacks native support for those fields.

## Policy and Streaming Support

Beyond basic serialization, the translation layer enforces behavioral policies and handles real-time data flows.

### TranslationPolicy

The `TranslationPolicy` struct in [`crates/switchyard-translation/src/policy.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/policy.rs) drives non-structural transformations:

- **Deterministic ID generation** – Ensures consistent identifiers across retry and cache operations
- **Content-type handling** – Determines when to preserve binary data versus convert to text
- **Lossy diagnostics** – Flags when translation compromises data fidelity, alerting upstream systems to potential information loss

Codecs consult the policy during ambiguous conversion scenarios, such as mapping between providers with incompatible tool schemas.

### Streaming Codecs

The `OpenAiChatStreamCodec` mirrors the buffered implementation but operates on `LlmResponseStream` events rather than complete responses. It translates incremental token deltas into SSE-formatted JSON lines while maintaining the same preservation and policy guarantees as the buffered variant. This enables Switchyard to route streaming requests through intermediate processing (caching, logging, tool execution) without breaking the real-time delivery contract to the client.

## Practical Code Examples

The following examples demonstrate API format conversion through the PyO3 bindings exposed as `switchyard_rust`.

### Decoding an OpenAI Chat Request

```python
from switchyard_rust import libsy
import json

openai_payload = json.loads("""{
    "model": "gpt-4",
    "messages": [
        {"role": "system", "content": "You are helpful."},
        {"role": "user", "content": "Explain quantum tunnelling."}
    ],
    "temperature": 0.7,
    "stream": false
}""")

request = libsy.decode_request(openai_payload, "openai_chat")
print(request)  # LlmRequest with normalized model, messages, sampling

```

### Encoding an Internal Response to OpenAI Format

```python
from switchyard_rust import libsy, llm

response = llm.AggLlmResponse(
    id="chatcmpl-123",
    model="gpt-4",
    outputs=[
        llm.ResponseOutput(
            role=llm.Role.Assistant,
            content=[llm.ContentBlock.Text(text="Quantum tunnelling is ...")],
            stop_reason=llm.StopReason.EndTurn,
        )
    ],
    usage=llm.Usage(input_tokens=10, output_tokens=20)
)

openai_response = libsy.encode_response(response, "openai_chat")
print(json.dumps(openai_response, indent=2))

```

### Streaming Encoding

```python
from switchyard_rust import libsy

# stream_events yields LlmResponseChunk objects

encoded_stream = libsy.encode_stream(stream_events, "openai_chat")

# Returns JSON line strings suitable for HTTP chunked transfer-encoding

```

## Summary

- **switchyard-translation** provides bidirectional mapping between provider-specific HTTP formats (OpenAI, Anthropic) and Switchyard's neutral `LlmRequest`/`AggLlmResponse` IR.
- The **FormatCodec** trait in [`engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/engine.rs) standardizes decode/encode operations across all providers, with implementations in `codecs/<provider>/`.
- **Preservation helpers** in [`util.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/util.rs) capture unknown fields during decoding and re-embed them during encoding, ensuring lossless round-tripping of vendor extensions.
- **TranslationPolicy** enforces deterministic IDs and content handling rules during conversion.
- **Streaming codecs** extend these capabilities to real-time SSE flows without breaking the abstraction layer.
- Being pure Rust with no external SDK dependencies, the crate compiles into both the native `switchyard-server` binary and Python bindings via PyO3.

## Frequently Asked Questions

### How does switchyard-translation handle unknown fields in provider payloads?

When decoding, the `capture_request_preservation` and `capture_response_preservation` functions in [`crates/switchyard-translation/src/util.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/util.rs) extract JSON subtrees that lack IR mappings. These fields travel with the internal request object and are re-inserted via `embed_preservation` during encoding, ensuring provider-specific metadata survives round-trips through Switchyard's routing layer.

### What is the FormatCodec trait and where is it defined?

The `FormatCodec` trait is defined in [`crates/switchyard-translation/src/engine.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/engine.rs) and specifies four required methods: `decode_request`, `encode_request`, `decode_response`, and `encode_response`. Each supported LLM provider implements this trait to handle its specific wire format, allowing the translation engine to treat all providers polymorphically.

### Does switchyard-translation support streaming responses?

Yes. Each provider codec includes a streaming variant (e.g., `OpenAiChatStreamCodec` in [`crates/switchyard-translation/src/codecs/openai_chat/stream.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/codecs/openai_chat/stream.rs)) that implements the same `FormatCodec` interface but operates on `LlmResponseStream` events. This converts incremental token deltas into SSE-formatted JSON lines while maintaining preservation and policy enforcement.

### Why is switchyard-translation implemented in Rust rather than Python?

The crate is implemented in pure Rust to eliminate external SDK dependencies and maximize performance. This design allows the translation engine to compile into the native `switchyard-server` binary for high-throughput scenarios while also exposing functionality to Python via PyO3 bindings (`switchyard_rust`), avoiding the overhead of importing heavyweight provider client libraries into the orchestration environment.