How Switchyard-Translation Enables API Format Conversion for LLMs

Switchyard-translation is a pure-Rust crate that performs bidirectional mapping between provider-specific HTTP formats (OpenAI, Anthropic, etc.) and Switchyard's neutral conversation IR, enabling seamless routing across heterogeneous LLM backends.

The switchyard-translation crate sits at the boundary between external LLM providers and the NVIDIA-NeMo/Switchyard orchestration layer. It translates disparate API schemas into a unified internal representation (switchyard_protocol::llm) and back again, allowing the Python orchestration layer to interact with any supported backend through a single, consistent interface.

Core Architecture of Switchyard-Translation

The translation engine operates through a trait-based codec system that isolates provider-specific serialization logic from the core routing infrastructure.

The FormatCodec Trait

At the heart of the system lies the FormatCodec trait defined in crates/switchyard-translation/src/engine.rs. This trait mandates four essential operations for every supported provider format:

  • decode_request – Transforms incoming provider JSON into an LlmRequest
  • encode_request – Converts an LlmRequest into the target provider's wire format
  • decode_response – Parses provider responses into AggLlmResponse
  • encode_response – Serializes internal responses back to provider-specific JSON

Each provider implementation registers itself against a FormatId, allowing the engine to dynamically select the appropriate codec at runtime.

Provider Implementations

Concrete codecs reside in crates/switchyard-translation/src/codecs/. The OpenAI Chat implementation serves as the reference architecture, split into two specialized variants:

Both variants implement the same FormatCodec interface, ensuring consistent behavior whether operating in buffered or streaming modes.

Bidirectional Translation Workflow

The conversion process preserves semantic meaning while adapting structural differences between provider schemas.

Decoding Provider Requests

When decoding an OpenAI-compatible payload, the decode_request method (lines 38-78 in buffered.rs) orchestrates field-level extraction through helper functions:

  • decode_openai_content – Parses message content blocks, handling text, images, and tool results
  • decode_openai_tool_call – Extracts function calling metadata into the internal tool representation
  • role_from_openai – Maps OpenAI role strings (system, user, assistant, tool) to the internal Role enum

The method constructs an LlmRequest containing the model identifier, message history, sampling parameters, and tool definitions, normalizing provider quirks into the canonical IR.

Encoding Internal Responses

The encode_request method (lines 81-124 in buffered.rs) performs the inverse operation, generating provider-specific JSON from an LlmRequest or AggLlmResponse. During encoding, the codec applies deterministic ID policies to ensure traceability across routing hops, copies extension fields, and implements fallback logic—emitting text representations when the target provider lacks support for specific content blocks like images or files.

The Preservation Layer

Unknown fields that lack direct IR mappings undergo lossless capture via the preservation mechanism implemented in crates/switchyard-translation/src/util.rs. The functions capture_request_preservation and capture_response_preservation serialize unrecognized JSON subtrees into a sidecar structure. Later, embed_preservation re-inserts these fields verbatim into the encoded output. This guarantees lossless round-tripping of provider-specific metadata (custom headers, vendor extensions, or experimental features) even when Switchyard lacks native support for those fields.

Policy and Streaming Support

Beyond basic serialization, the translation layer enforces behavioral policies and handles real-time data flows.

TranslationPolicy

The TranslationPolicy struct in crates/switchyard-translation/src/policy.rs drives non-structural transformations:

  • Deterministic ID generation – Ensures consistent identifiers across retry and cache operations
  • Content-type handling – Determines when to preserve binary data versus convert to text
  • Lossy diagnostics – Flags when translation compromises data fidelity, alerting upstream systems to potential information loss

Codecs consult the policy during ambiguous conversion scenarios, such as mapping between providers with incompatible tool schemas.

Streaming Codecs

The OpenAiChatStreamCodec mirrors the buffered implementation but operates on LlmResponseStream events rather than complete responses. It translates incremental token deltas into SSE-formatted JSON lines while maintaining the same preservation and policy guarantees as the buffered variant. This enables Switchyard to route streaming requests through intermediate processing (caching, logging, tool execution) without breaking the real-time delivery contract to the client.

Practical Code Examples

The following examples demonstrate API format conversion through the PyO3 bindings exposed as switchyard_rust.

Decoding an OpenAI Chat Request

from switchyard_rust import libsy
import json

openai_payload = json.loads("""{
    "model": "gpt-4",
    "messages": [
        {"role": "system", "content": "You are helpful."},
        {"role": "user", "content": "Explain quantum tunnelling."}
    ],
    "temperature": 0.7,
    "stream": false
}""")

request = libsy.decode_request(openai_payload, "openai_chat")
print(request)  # LlmRequest with normalized model, messages, sampling

Encoding an Internal Response to OpenAI Format

from switchyard_rust import libsy, llm

response = llm.AggLlmResponse(
    id="chatcmpl-123",
    model="gpt-4",
    outputs=[
        llm.ResponseOutput(
            role=llm.Role.Assistant,
            content=[llm.ContentBlock.Text(text="Quantum tunnelling is ...")],
            stop_reason=llm.StopReason.EndTurn,
        )
    ],
    usage=llm.Usage(input_tokens=10, output_tokens=20)
)

openai_response = libsy.encode_response(response, "openai_chat")
print(json.dumps(openai_response, indent=2))

Streaming Encoding

from switchyard_rust import libsy

# stream_events yields LlmResponseChunk objects

encoded_stream = libsy.encode_stream(stream_events, "openai_chat")

# Returns JSON line strings suitable for HTTP chunked transfer-encoding

Summary

  • switchyard-translation provides bidirectional mapping between provider-specific HTTP formats (OpenAI, Anthropic) and Switchyard's neutral LlmRequest/AggLlmResponse IR.
  • The FormatCodec trait in engine.rs standardizes decode/encode operations across all providers, with implementations in codecs/<provider>/.
  • Preservation helpers in util.rs capture unknown fields during decoding and re-embed them during encoding, ensuring lossless round-tripping of vendor extensions.
  • TranslationPolicy enforces deterministic IDs and content handling rules during conversion.
  • Streaming codecs extend these capabilities to real-time SSE flows without breaking the abstraction layer.
  • Being pure Rust with no external SDK dependencies, the crate compiles into both the native switchyard-server binary and Python bindings via PyO3.

Frequently Asked Questions

How does switchyard-translation handle unknown fields in provider payloads?

When decoding, the capture_request_preservation and capture_response_preservation functions in crates/switchyard-translation/src/util.rs extract JSON subtrees that lack IR mappings. These fields travel with the internal request object and are re-inserted via embed_preservation during encoding, ensuring provider-specific metadata survives round-trips through Switchyard's routing layer.

What is the FormatCodec trait and where is it defined?

The FormatCodec trait is defined in crates/switchyard-translation/src/engine.rs and specifies four required methods: decode_request, encode_request, decode_response, and encode_response. Each supported LLM provider implements this trait to handle its specific wire format, allowing the translation engine to treat all providers polymorphically.

Does switchyard-translation support streaming responses?

Yes. Each provider codec includes a streaming variant (e.g., OpenAiChatStreamCodec in crates/switchyard-translation/src/codecs/openai_chat/stream.rs) that implements the same FormatCodec interface but operates on LlmResponseStream events. This converts incremental token deltas into SSE-formatted JSON lines while maintaining preservation and policy enforcement.

Why is switchyard-translation implemented in Rust rather than Python?

The crate is implemented in pure Rust to eliminate external SDK dependencies and maximize performance. This design allows the translation engine to compile into the native switchyard-server binary for high-throughput scenarios while also exposing functionality to Python via PyO3 bindings (switchyard_rust), avoiding the overhead of importing heavyweight provider client libraries into the orchestration environment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →