How Switchyard Translates Between OpenAI and Anthropic API Formats
Switchyard provides a bidirectional translation layer that converts OpenAI Chat-Completions API requests into Anthropic Claude format and vice-versa using a neutral intermediate representation and provider-specific codec modules.
Switchyard enables seamless interoperability between competing LLM APIs by acting as a protocol bridge. This open-source project from NVIDIA-NeMo allows developers to write client code against the OpenAI SDK while routing requests to Anthropic's Claude models, or use Anthropic's message format with OpenAI backends. The translation system is implemented in the switchyard-translation crate and maintains full fidelity for streaming responses, tool calls, and usage metrics.
The Neutral Intermediate Representation (IR)
At the core of Switchyard's translation architecture sits a provider-agnostic intermediate representation that decouples client and server protocols. When a request arrives, it is decoded into an LlmRequest IR that captures normalized message roles, content blocks, tool definitions, and stop parameters.
This IR abstraction allows the switchyard-translation::Engine to handle cross-provider mapping without hardcoding bilateral conversions. The engine, defined in crates/switchyard-translation/src/engine.rs, maintains translation state and routes payloads through the appropriate codec pairs.
Request Flow: Decoding Provider-Specific Formats
Incoming client requests are parsed by provider-specific request codecs that extract fields and populate the neutral IR.
OpenAI Request Parsing
For OpenAI Chat-Completions payloads, the engine selects either openai_chat::buffered or openai_chat::stream from crates/switchyard-translation/src/codecs/openai_chat/mod.rs. The codec's decode method processes OpenAI-specific fields like tool_choice and logprobs, mapping them to Anthropic equivalents or normalizing them into the IR structure.
Anthropic Request Parsing
Anthropic Message API requests are handled by the anthropic::messages codec located in crates/switchyard-translation/src/codecs/anthropic/mod.rs. This module translates Anthropic-native fields—such as max_tokens and stop_sequences—into the IR's standardized parameter set, ensuring equivalent behavior when the request reaches an OpenAI backend.
Response Flow: Encoding to Target Formats
After the upstream provider generates a response, the engine selects the matching response codec to serialize the LlmResponse IR back to the client's expected format.
Key mapping operations include:
- Converting Anthropic "content_block" objects into OpenAI Chat "chunks" and vice-versa
- Translating Anthropic "thinking" blocks into OpenAI's
reasoning_contentfield - Normalizing usage statistics (
prompt_tokens,completion_tokens) into either OpenAI'susageobject or Anthropic'susagestructure - Mapping stop reasons via helper functions in
crates/switchyard-translation/src/util.rs, convertingstop_reasonto OpenAI'sfinish_reasonandstop_detailsto thestopfield
Handling Streaming Responses
For real-time streams, the engine maintains a StreamTranslationState that tracks partial messages, accumulating usage statistics and managing stop sequence detection. OpenAI stream events containing delta objects are decoded into IR LlmResponseChunk structures, then re-encoded as Anthropic stream events (or the reverse direction). The state machine ensures that tool-result deltas, reasoning content, and final terminal chunks are emitted in the correct order regardless of provider directionality.
Preservation Metadata for Round-Trip Integrity
Switchyard embeds provenance data using a hidden PRESERVATION_METADATA_KEY that records the original provider and model identifiers. This metadata allows requests to traverse multiple translation layers—OpenAI to Anthropic and back to OpenAI—without losing request provenance or requiring client-side tracking.
Implementation Examples
The following Python example demonstrates how a client using OpenAI-style requests communicates with an Anthropic backend through Switchyard:
import switchyard
import asyncio
async def main():
client = switchyard.SwitchyardClient(
base_url="http://localhost:4000",
provider="anthropic"
)
request = {
"model": "claude-2.0",
"messages": [
{"role": "user", "content": "Explain quantum tunneling in simple terms."}
],
"max_tokens": 256,
"temperature": 0.7,
}
response = await client.chat_completion(request)
print(response["choices"][0]["message"]["content"])
asyncio.run(main())
For server-side translation, the Rust implementation in the switchyard-translation crate demonstrates the codec workflow:
use switchyard_translation::engine::Engine;
use switchyard_translation::codecs::openai_chat::buffered::OpenAiChatCodec;
use switchyard_translation::codecs::anthropic::messages::AnthropicMessagesCodec;
let openai_json = r#"{ "model":"gpt-4", "messages":[{"role":"user","content":"Hello"}] }"#;
let mut engine = Engine::new();
let ir_request = OpenAiChatCodec::decode(&mut engine, openai_json)?;
let openai_response = OpenAiChatCodec::encode(&mut engine, ir_response)?;
Key Source Files
The translation pipeline relies on these specific modules within the NVIDIA-NeMo/Switchyard repository:
crates/switchyard-translation/src/engine.rs– Core translation engine and state managementcrates/switchyard-translation/src/codecs/openai_chat/buffered.rs– Buffered OpenAI Chat codeccrates/switchyard-translation/src/codecs/openai_chat/stream.rs– Streaming OpenAI translation logiccrates/switchyard-translation/src/codecs/anthropic/messages.rs– Anthropic Messages API codeccrates/switchyard-translation/src/util.rs– Stop reason mapping and ID normalization utilitiescrates/switchyard-translation/tests/request_translation.rs– OpenAI to Anthropic request test suitecrates/switchyard-translation/tests/response_translation.rs– Anthropic to OpenAI response validation
Summary
- Switchyard uses a neutral IR (
LlmRequest/LlmResponse) to avoid N-to-N translation complexity, requiring only one codec per provider. - Provider-specific codecs handle schema normalization, dropping unsupported fields like OpenAI's
logprobswhen targeting Anthropic, or mapping Anthropic'sstop_sequencesto OpenAI'sstopparameter. - Streaming translation maintains state via
StreamTranslationStateto properly order deltas, tool results, and reasoning content across real-time connections. - Bidirectional preservation metadata enables transparent proxy chains and round-trip routing without data loss.
- Modular architecture allows new providers to be added by implementing the codec trait, without modifying the core routing logic in
engine.rs.
Frequently Asked Questions
Can Switchyard translate streaming responses between OpenAI and Anthropic in real-time?
Yes. The StreamTranslationState struct in the translation engine tracks partial message deltas, usage accumulation, and stop sequences. It decodes incoming stream events—such as OpenAI's delta objects—into neutral LlmResponseChunk instances, then re-encodes them as Anthropic content blocks or vice versa. This ensures that reasoning content, tool calls, and usage statistics arrive in the correct order on both sides.
How does Switchyard handle fields that exist in one API but not the other?
Provider-specific codecs normalize or drop incompatible fields during the decode phase. For example, OpenAI's tool_choice parameter is either mapped to Anthropic equivalents or omitted, while Anthropic's max_tokens is converted to OpenAI's identically named field. The util.rs module provides helper functions for mapping stop reasons and sanitizing tool identifiers to maintain functional parity.
Is it possible to route OpenAI client code to Anthropic models without modifying the client?
Yes. Switchyard exposes a server endpoint that accepts standard OpenAI Chat-Completions JSON. When configured with provider="anthropic", the engine transparently converts the request to Anthropic's Messages format, forwards it to Claude, then translates the response back to OpenAI format before returning it to the client. The client remains unaware that the backend is Anthropic rather than OpenAI.
Where is the preservation metadata stored during translation?
The engine injects a PRESERVATION_METADATA_KEY into the request context that records the original provider and model identifiers. This metadata persists through the IR transformation and allows subsequent translation layers to maintain provenance information. This is particularly useful for logging, debugging, and scenarios where requests traverse multiple translation hops.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →