What Is NVIDIA Switchyard? A Rust-Based LLM Proxy and Router

NVIDIA Switchyard is an open-source Rust proxy that translates between OpenAI and Anthropic API formats while routing requests across multiple LLM backends using pluggable algorithms.

NVIDIA Switchyard is an open-source Rust-based proxy and library developed under the NVIDIA-NeMo organization. It acts as an intermediary layer between LLM client applications—such as Claude Code, the OpenAI SDK, or the Anthropic SDK—and underlying model backends. By handling protocol translation and intelligent routing, Switchyard enables teams to deploy multi-backend inference architectures without modifying existing client code.

Core Architecture and Responsibilities

According to the repository's README.md, Switchyard operates as a thin, stateless layer that does not perform inference itself. Instead, it receives requests in the client's native format, selects a backend via configurable routing logic, translates the payload into the provider's expected format, and returns the translated response to the client.

Protocol Translation

The crates/switchyard-translation/README.md module handles bidirectional conversion between OpenAI Chat, OpenAI Responses, and Anthropic Messages formats. This allows clients to continue using their native SDKs while Switchyard communicates with any OpenAI-compatible or Anthropic-compatible endpoint. The translation layer ensures that request schemas, authentication headers, and response structures are correctly mapped between providers.

Multi-Backend Routing

Switchyard implements a pluggable routing layer that determines which target model should serve each request. The routing algorithms available in docs/routing_algorithms/overview.md include:

  • Random split – Distributes traffic probabilistically across multiple models
  • LLM-as-classifier – Uses a language model to categorize and route requests
  • Stage-router – Implements multi-stage routing pipelines
  • Escalation router – Attempts cheaper models first, escalating to more expensive ones on failure

Operational Observability

The proxy emits Prometheus-compatible metrics for production monitoring. As implemented in the server crate, Switchyard tracks request counts, error rates, end-to-end latency, token usage statistics, and routing decision overhead. These metrics expose the health and performance characteristics of your multi-backend LLM infrastructure.

Deployment Options: Server vs. Library

Switchyard supports two primary integration patterns: running as a standalone HTTP proxy or embedding the routing logic directly into a Rust application.

Running the Standalone Server (switchyard-server)

The server path installs a compiled binary that acts as an HTTP proxy. According to crates/switchyard-server/README.md, installation and execution follow standard Cargo workflows:


# Install the server binary

cargo install --locked switchyard-server

# Configure environment and validate

export OPENROUTER_API_KEY="your-openrouter-key"
switchyard-server --config routes.toml --dry-run

# Start the proxy

switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

Verify the server health using the built-in endpoint:

curl http://localhost:4000/health

# {"status":"ok"}

Embedding the Library (switchyard-libsy)

For Rust applications requiring direct integration, add the switchyard-libsy and switchyard-protocol crates to your Cargo.toml:

[dependencies]
switchyard-libsy = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
switchyard-protocol = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }

The library exposes the RoutingConfig and Router types from crates/libsy/README.md. Use RoutingConfig::from_file to load TOML configurations and Router::new to initialize the routing engine:

use switchyard_libsy::{RoutingConfig, Router};
use switchyard_protocol::{ChatRequest, ChatResponse};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Initialize from TOML configuration
    let cfg = RoutingConfig::from_file("routes.toml")?;
    let router = Router::new(cfg);

    // Create request in OpenAI Chat format
    let request = ChatRequest::new("gpt-4o-mini", "You are a helpful assistant.", "Hello!");
    
    // Route and receive response
    let response: ChatResponse = router.handle(request)?;
    println!("LLM replied: {}", response.choices[0].message.content);
    Ok(())
}

Configuring Routing Algorithms

Routing behavior is defined in TOML configuration files. As documented in docs/getting_started.md, the following example configures a random A/B split routing 70% of traffic to gpt-4o-mini and 30% to gpt-4o:

[[routes]]
name = "my_random_route"
type = "random"
targets = [
  { model_id = "openai:gpt-4o-mini", weight = 0.7 },
  { model_id = "openai:gpt-4o",     weight = 0.3 }
]

This configuration file is consumed by both the standalone server and the embedded library, ensuring consistent routing logic across deployment modes.

Key Source Files and Implementation Details

The Switchyard codebase is organized into distinct crates under the crates/ directory:

File Path Responsibility
README.md High-level overview, quick-start guide, and feature list
crates/switchyard-server/README.md Build instructions and runtime configuration for the standalone proxy
crates/libsy/README.md Core routing library with Router and RoutingConfig implementations
crates/protocol/README.md Provider-neutral type definitions for ChatRequest and ChatResponse
crates/switchyard-translation/README.md Cross-API translation logic between OpenAI and Anthropic formats
docs/routing_algorithms/overview.md Comprehensive routing strategy documentation and configuration schemas

Summary

  • NVIDIA Switchyard is a Rust-based LLM proxy that separates client SDK compatibility from backend provider selection.
  • It provides protocol translation between OpenAI and Anthropic API formats without requiring client code changes.
  • Pluggable routing algorithms (random, classifier-based, escalation) determine backend selection via TOML configuration.
  • Two deployment modes are supported: a standalone server (switchyard-server) and an embeddable library (switchyard-libsy).
  • The architecture exposes Prometheus metrics for monitoring multi-backend LLM deployments.
  • Core implementations reside in crates/libsy/README.md for routing logic and crates/switchyard-translation/README.md for API format conversion.

Frequently Asked Questions

What is NVIDIA Switchyard used for?

NVIDIA Switchyard enables development teams to route LLM traffic across multiple model providers (OpenAI, Anthropic, or compatible endpoints) from a single entry point. It solves the problem of vendor lock-in by allowing applications to use their native SDKs while the proxy handles backend selection, failover, and protocol translation.

How does Switchyard handle different LLM API formats?

Switchyard uses the translation layer implemented in crates/switchyard-translation/README.md to convert between OpenAI Chat, OpenAI Responses, and Anthropic Messages formats. When a client sends a request in Anthropic's native format, Switchyard translates it to the target backend's expected schema, forwards it, then translates the response back to the client's expected format.

Can I use Switchyard as a library in my Rust application?

Yes. By adding the switchyard-libsy and switchyard-protocol crates to your Cargo.toml, you can embed the routing engine directly into your Rust code. The library exposes RoutingConfig::from_file for configuration loading and Router::new for request handling, allowing you to integrate intelligent routing without running a separate proxy process.

What routing algorithms does NVIDIA Switchyard support?

According to docs/routing_algorithms/overview.md, Switchyard supports multiple routing strategies including random traffic splitting, LLM-as-classifier routing for content-based decisions, stage-routers for multi-step pipelines, and escalation routers that attempt cost-effective models before falling back to premium options.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →