# What Is NVIDIA Switchyard? A Rust-Based LLM Proxy and Router

> Discover NVIDIA Switchyard, a Rust proxy that translates OpenAI and Anthropic API formats. Route requests to multiple LLM backends with pluggable algorithms.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: getting-started
- Published: 2026-08-23

---

**NVIDIA Switchyard is an open-source Rust proxy that translates between OpenAI and Anthropic API formats while routing requests across multiple LLM backends using pluggable algorithms.**

NVIDIA Switchyard is an open-source Rust-based proxy and library developed under the NVIDIA-NeMo organization. It acts as an intermediary layer between LLM client applications—such as Claude Code, the OpenAI SDK, or the Anthropic SDK—and underlying model backends. By handling protocol translation and intelligent routing, Switchyard enables teams to deploy multi-backend inference architectures without modifying existing client code.

## Core Architecture and Responsibilities

According to the repository's [`README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/README.md), Switchyard operates as a thin, stateless layer that does not perform inference itself. Instead, it receives requests in the client's native format, selects a backend via configurable routing logic, translates the payload into the provider's expected format, and returns the translated response to the client.

### Protocol Translation

The [`crates/switchyard-translation/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/README.md) module handles bidirectional conversion between **OpenAI Chat**, **OpenAI Responses**, and **Anthropic Messages** formats. This allows clients to continue using their native SDKs while Switchyard communicates with any OpenAI-compatible or Anthropic-compatible endpoint. The translation layer ensures that request schemas, authentication headers, and response structures are correctly mapped between providers.

### Multi-Backend Routing

Switchyard implements a pluggable routing layer that determines which target model should serve each request. The routing algorithms available in [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md) include:

- **Random split** – Distributes traffic probabilistically across multiple models
- **LLM-as-classifier** – Uses a language model to categorize and route requests
- **Stage-router** – Implements multi-stage routing pipelines
- **Escalation router** – Attempts cheaper models first, escalating to more expensive ones on failure

### Operational Observability

The proxy emits **Prometheus-compatible metrics** for production monitoring. As implemented in the server crate, Switchyard tracks request counts, error rates, end-to-end latency, token usage statistics, and routing decision overhead. These metrics expose the health and performance characteristics of your multi-backend LLM infrastructure.

## Deployment Options: Server vs. Library

Switchyard supports two primary integration patterns: running as a standalone HTTP proxy or embedding the routing logic directly into a Rust application.

### Running the Standalone Server (`switchyard-server`)

The server path installs a compiled binary that acts as an HTTP proxy. According to [`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md), installation and execution follow standard Cargo workflows:

```bash

# Install the server binary

cargo install --locked switchyard-server

# Configure environment and validate

export OPENROUTER_API_KEY="your-openrouter-key"
switchyard-server --config routes.toml --dry-run

# Start the proxy

switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

Verify the server health using the built-in endpoint:

```bash
curl http://localhost:4000/health

# {"status":"ok"}

```

### Embedding the Library (`switchyard-libsy`)

For Rust applications requiring direct integration, add the `switchyard-libsy` and `switchyard-protocol` crates to your [`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml):

```toml
[dependencies]
switchyard-libsy = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
switchyard-protocol = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }

```

The library exposes the `RoutingConfig` and `Router` types from [`crates/libsy/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md). Use `RoutingConfig::from_file` to load TOML configurations and `Router::new` to initialize the routing engine:

```rust
use switchyard_libsy::{RoutingConfig, Router};
use switchyard_protocol::{ChatRequest, ChatResponse};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Initialize from TOML configuration
    let cfg = RoutingConfig::from_file("routes.toml")?;
    let router = Router::new(cfg);

    // Create request in OpenAI Chat format
    let request = ChatRequest::new("gpt-4o-mini", "You are a helpful assistant.", "Hello!");
    
    // Route and receive response
    let response: ChatResponse = router.handle(request)?;
    println!("LLM replied: {}", response.choices[0].message.content);
    Ok(())
}

```

## Configuring Routing Algorithms

Routing behavior is defined in TOML configuration files. As documented in [`docs/getting_started.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/getting_started.md), the following example configures a random A/B split routing 70% of traffic to `gpt-4o-mini` and 30% to `gpt-4o`:

```toml
[[routes]]
name = "my_random_route"
type = "random"
targets = [
  { model_id = "openai:gpt-4o-mini", weight = 0.7 },
  { model_id = "openai:gpt-4o",     weight = 0.3 }
]

```

This configuration file is consumed by both the standalone server and the embedded library, ensuring consistent routing logic across deployment modes.

## Key Source Files and Implementation Details

The Switchyard codebase is organized into distinct crates under the `crates/` directory:

| File Path | Responsibility |
|-----------|----------------|
| [`README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/README.md) | High-level overview, quick-start guide, and feature list |
| [`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md) | Build instructions and runtime configuration for the standalone proxy |
| [`crates/libsy/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md) | Core routing library with `Router` and `RoutingConfig` implementations |
| [`crates/protocol/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/README.md) | Provider-neutral type definitions for `ChatRequest` and `ChatResponse` |
| [`crates/switchyard-translation/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/README.md) | Cross-API translation logic between OpenAI and Anthropic formats |
| [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md) | Comprehensive routing strategy documentation and configuration schemas |

## Summary

- **NVIDIA Switchyard** is a Rust-based LLM proxy that separates client SDK compatibility from backend provider selection.
- It provides **protocol translation** between OpenAI and Anthropic API formats without requiring client code changes.
- **Pluggable routing algorithms** (random, classifier-based, escalation) determine backend selection via TOML configuration.
- Two deployment modes are supported: a **standalone server** (`switchyard-server`) and an **embeddable library** (`switchyard-libsy`).
- The architecture exposes **Prometheus metrics** for monitoring multi-backend LLM deployments.
- Core implementations reside in [`crates/libsy/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md) for routing logic and [`crates/switchyard-translation/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/README.md) for API format conversion.

## Frequently Asked Questions

### What is NVIDIA Switchyard used for?

NVIDIA Switchyard enables development teams to route LLM traffic across multiple model providers (OpenAI, Anthropic, or compatible endpoints) from a single entry point. It solves the problem of vendor lock-in by allowing applications to use their native SDKs while the proxy handles backend selection, failover, and protocol translation.

### How does Switchyard handle different LLM API formats?

Switchyard uses the translation layer implemented in [`crates/switchyard-translation/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/README.md) to convert between OpenAI Chat, OpenAI Responses, and Anthropic Messages formats. When a client sends a request in Anthropic's native format, Switchyard translates it to the target backend's expected schema, forwards it, then translates the response back to the client's expected format.

### Can I use Switchyard as a library in my Rust application?

Yes. By adding the `switchyard-libsy` and `switchyard-protocol` crates to your [`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml), you can embed the routing engine directly into your Rust code. The library exposes `RoutingConfig::from_file` for configuration loading and `Router::new` for request handling, allowing you to integrate intelligent routing without running a separate proxy process.

### What routing algorithms does NVIDIA Switchyard support?

According to [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md), Switchyard supports multiple routing strategies including random traffic splitting, LLM-as-classifier routing for content-based decisions, stage-routers for multi-step pipelines, and escalation routers that attempt cost-effective models before falling back to premium options.