How Switchyard-Translation Enables API Format Conversion for LLMs
Switchyard-translation is a pure-Rust crate that performs bidirectional mapping between provider-specific HTTP formats (OpenAI, Anthropic, etc.) and Switchyard's neutral conversation IR, enabling seamless routing across heterogeneous LLM backends.
The switchyard-translation crate sits at the boundary between external LLM providers and the NVIDIA-NeMo/Switchyard orchestration layer. It translates disparate API schemas into a unified internal representation (switchyard_protocol::llm) and back again, allowing the Python orchestration layer to interact with any supported backend through a single, consistent interface.
Core Architecture of Switchyard-Translation
The translation engine operates through a trait-based codec system that isolates provider-specific serialization logic from the core routing infrastructure.
The FormatCodec Trait
At the heart of the system lies the FormatCodec trait defined in crates/switchyard-translation/src/engine.rs. This trait mandates four essential operations for every supported provider format:
decode_request– Transforms incoming provider JSON into anLlmRequestencode_request– Converts anLlmRequestinto the target provider's wire formatdecode_response– Parses provider responses intoAggLlmResponseencode_response– Serializes internal responses back to provider-specific JSON
Each provider implementation registers itself against a FormatId, allowing the engine to dynamically select the appropriate codec at runtime.
Provider Implementations
Concrete codecs reside in crates/switchyard-translation/src/codecs/. The OpenAI Chat implementation serves as the reference architecture, split into two specialized variants:
OpenAiChatCodecincodecs/openai_chat/buffered.rshandles complete request/response cyclesOpenAiChatStreamCodecincodecs/openai_chat/stream.rsmanages Server-Sent Events (SSE) for real-time token streaming
Both variants implement the same FormatCodec interface, ensuring consistent behavior whether operating in buffered or streaming modes.
Bidirectional Translation Workflow
The conversion process preserves semantic meaning while adapting structural differences between provider schemas.
Decoding Provider Requests
When decoding an OpenAI-compatible payload, the decode_request method (lines 38-78 in buffered.rs) orchestrates field-level extraction through helper functions:
decode_openai_content– Parses message content blocks, handling text, images, and tool resultsdecode_openai_tool_call– Extracts function calling metadata into the internal tool representationrole_from_openai– Maps OpenAI role strings (system,user,assistant,tool) to the internalRoleenum
The method constructs an LlmRequest containing the model identifier, message history, sampling parameters, and tool definitions, normalizing provider quirks into the canonical IR.
Encoding Internal Responses
The encode_request method (lines 81-124 in buffered.rs) performs the inverse operation, generating provider-specific JSON from an LlmRequest or AggLlmResponse. During encoding, the codec applies deterministic ID policies to ensure traceability across routing hops, copies extension fields, and implements fallback logic—emitting text representations when the target provider lacks support for specific content blocks like images or files.
The Preservation Layer
Unknown fields that lack direct IR mappings undergo lossless capture via the preservation mechanism implemented in crates/switchyard-translation/src/util.rs. The functions capture_request_preservation and capture_response_preservation serialize unrecognized JSON subtrees into a sidecar structure. Later, embed_preservation re-inserts these fields verbatim into the encoded output. This guarantees lossless round-tripping of provider-specific metadata (custom headers, vendor extensions, or experimental features) even when Switchyard lacks native support for those fields.
Policy and Streaming Support
Beyond basic serialization, the translation layer enforces behavioral policies and handles real-time data flows.
TranslationPolicy
The TranslationPolicy struct in crates/switchyard-translation/src/policy.rs drives non-structural transformations:
- Deterministic ID generation – Ensures consistent identifiers across retry and cache operations
- Content-type handling – Determines when to preserve binary data versus convert to text
- Lossy diagnostics – Flags when translation compromises data fidelity, alerting upstream systems to potential information loss
Codecs consult the policy during ambiguous conversion scenarios, such as mapping between providers with incompatible tool schemas.
Streaming Codecs
The OpenAiChatStreamCodec mirrors the buffered implementation but operates on LlmResponseStream events rather than complete responses. It translates incremental token deltas into SSE-formatted JSON lines while maintaining the same preservation and policy guarantees as the buffered variant. This enables Switchyard to route streaming requests through intermediate processing (caching, logging, tool execution) without breaking the real-time delivery contract to the client.
Practical Code Examples
The following examples demonstrate API format conversion through the PyO3 bindings exposed as switchyard_rust.
Decoding an OpenAI Chat Request
from switchyard_rust import libsy
import json
openai_payload = json.loads("""{
"model": "gpt-4",
"messages": [
{"role": "system", "content": "You are helpful."},
{"role": "user", "content": "Explain quantum tunnelling."}
],
"temperature": 0.7,
"stream": false
}""")
request = libsy.decode_request(openai_payload, "openai_chat")
print(request) # LlmRequest with normalized model, messages, sampling
Encoding an Internal Response to OpenAI Format
from switchyard_rust import libsy, llm
response = llm.AggLlmResponse(
id="chatcmpl-123",
model="gpt-4",
outputs=[
llm.ResponseOutput(
role=llm.Role.Assistant,
content=[llm.ContentBlock.Text(text="Quantum tunnelling is ...")],
stop_reason=llm.StopReason.EndTurn,
)
],
usage=llm.Usage(input_tokens=10, output_tokens=20)
)
openai_response = libsy.encode_response(response, "openai_chat")
print(json.dumps(openai_response, indent=2))
Streaming Encoding
from switchyard_rust import libsy
# stream_events yields LlmResponseChunk objects
encoded_stream = libsy.encode_stream(stream_events, "openai_chat")
# Returns JSON line strings suitable for HTTP chunked transfer-encoding
Summary
- switchyard-translation provides bidirectional mapping between provider-specific HTTP formats (OpenAI, Anthropic) and Switchyard's neutral
LlmRequest/AggLlmResponseIR. - The FormatCodec trait in
engine.rsstandardizes decode/encode operations across all providers, with implementations incodecs/<provider>/. - Preservation helpers in
util.rscapture unknown fields during decoding and re-embed them during encoding, ensuring lossless round-tripping of vendor extensions. - TranslationPolicy enforces deterministic IDs and content handling rules during conversion.
- Streaming codecs extend these capabilities to real-time SSE flows without breaking the abstraction layer.
- Being pure Rust with no external SDK dependencies, the crate compiles into both the native
switchyard-serverbinary and Python bindings via PyO3.
Frequently Asked Questions
How does switchyard-translation handle unknown fields in provider payloads?
When decoding, the capture_request_preservation and capture_response_preservation functions in crates/switchyard-translation/src/util.rs extract JSON subtrees that lack IR mappings. These fields travel with the internal request object and are re-inserted via embed_preservation during encoding, ensuring provider-specific metadata survives round-trips through Switchyard's routing layer.
What is the FormatCodec trait and where is it defined?
The FormatCodec trait is defined in crates/switchyard-translation/src/engine.rs and specifies four required methods: decode_request, encode_request, decode_response, and encode_response. Each supported LLM provider implements this trait to handle its specific wire format, allowing the translation engine to treat all providers polymorphically.
Does switchyard-translation support streaming responses?
Yes. Each provider codec includes a streaming variant (e.g., OpenAiChatStreamCodec in crates/switchyard-translation/src/codecs/openai_chat/stream.rs) that implements the same FormatCodec interface but operates on LlmResponseStream events. This converts incremental token deltas into SSE-formatted JSON lines while maintaining preservation and policy enforcement.
Why is switchyard-translation implemented in Rust rather than Python?
The crate is implemented in pure Rust to eliminate external SDK dependencies and maximize performance. This design allows the translation engine to compile into the native switchyard-server binary for high-throughput scenarios while also exposing functionality to Python via PyO3 bindings (switchyard_rust), avoiding the overhead of importing heavyweight provider client libraries into the orchestration environment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →