# How to Set Up Switchyard as an LLM Traffic Orchestration Proxy

> Learn how to set up Switchyard, a Rust proxy, to orchestrate LLM traffic. Route requests to optimal backends with configurable algorithms like random split or LLM-as-classifier.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-21

---

**Switchyard is a Rust-based proxy that intercepts OpenAI and Anthropic SDK requests, translates them into a provider-neutral internal representation, and routes them to optimal backend LLMs using configurable algorithms like random split, LLM-as-classifier, or escalation routing.**

Setting up **Switchyard as an LLM traffic orchestration proxy** allows you to sit between client applications and backend providers to perform intelligent load balancing, cost optimization, and failover handling. The NVIDIA-NeMo/Switchyard repository provides both a standalone server binary for HTTP proxying and a library crate for embedded Rust applications.

## What Is Switchyard?

Switchyard functions as a protocol-aware gateway that normalizes disparate LLM API formats into a unified routing layer. According to the architecture documentation in the NVIDIA-NeMo/Switchyard repository, the system performs three core functions: **protocol translation** between OpenAI Chat, OpenAI Responses, and Anthropic Messages formats; **intelligent routing** using pluggable algorithms; and **metrics emission** with automatic fallback handling.

### Core Architecture Components

The proxy operates through a provider-neutral intermediate representation (IR) defined in [`crates/protocol/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/README.md). When a client sends a request, Switchyard:

1. Translates the native format (OpenAI or Anthropic) into the internal IR
2. Applies the configured routing algorithm to select a backend target
3. Converts the request back to the target's native format
4. Returns the response to the client in their original expected format

As implemented in `crates/switchyard-server`, the server automatically exposes health endpoints and Prometheus metrics at `/metrics` for observability into request counts, latencies, and token usage.

### Execution Paths: Server vs. Library

Switchyard supports two distinct deployment patterns:

- **Server path** (`switchyard-server` binary): Use this when you need a standalone HTTP proxy that any OpenAI-compatible client can call without code changes. This is the most common setup for LLM traffic orchestration.
- **Library path** (`switchyard-libsy` crate): Use this when building a custom Rust application that needs embedded routing logic without the HTTP layer overhead.

This guide focuses on the **server path**, which provides immediate integration with existing Python, JavaScript, or CLI-based LLM clients.

## Prerequisites

Before installing Switchyard, ensure your environment meets the following requirements:

- **Rust toolchain** ≥ 1.96 with Cargo (install via [rustup](https://rustup.rs/))
- **API credentials** for an OpenAI-compatible provider (OpenRouter, OpenAI, Anthropic, or Azure OpenAI)
- **Network access** to your chosen LLM backend endpoints

The [`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md) file specifies that the server binds to TCP ports by default and requires environment variables for API key injection rather than hardcoded secrets.

## Installation and Basic Configuration

### Install the Server Binary

Install the Switchyard server using Cargo's locked installation to ensure reproducible builds:

```bash
cargo install --locked switchyard-server

```

This command compiles the release binary and places it in `~/.cargo/bin`. Verify the installation by checking the help output:

```bash
switchyard-server --help

```

### Create the routes.toml Configuration

Switchyard reads a TOML configuration file that declares LLM clients, backend targets, and routing rules. Create a file named [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) in your working directory with the following structure:

```toml
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[routes.smart]
id = "switchyard"
type = "llm_classifier"
mode = "capability"
classifier_target = "weak"
strong_target = "strong"
weak_target = "weak"
base_threshold = 0.5

```

Key configuration parameters:
- `format`: Specifies the upstream protocol (`openai_chat`, `openai_responses`, or `anthropic_messages`)
- `api_key_env`: Names the environment variable containing the provider secret (secrets never appear in the configuration file)
- `type`: Defines the routing algorithm (see routing options below)

### Validate with Dry-Run

Before opening network sockets, validate your configuration syntax and environment variable presence:

```bash
export OPENROUTER_API_KEY="your-api-key-here"
switchyard-server --config routes.toml --dry-run

```

The `--dry-run` flag parses the TOML schema, checks that referenced environment variables exist, validates target references, and reports configuration errors without starting the HTTP server.

## Running the LLM Traffic Orchestration Proxy

### Start the Server

Launch the proxy with explicit host and port bindings:

```bash
switchyard-server --config routes.toml \
  --host 127.0.0.1 \
  --port 4000

```

The server now listens on `http://127.0.0.1:4000` and is ready to accept OpenAI-compatible requests. For production deployments, bind to `0.0.0.0` or a specific interface as documented in [`docs/getting_started.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/getting_started.md).

### Verify the Proxy Endpoints

Test the running server using standard HTTP clients:

```bash

# Health check

curl http://localhost:4000/health

# Expected: {"status":"ok"}

# List available models (includes your route IDs)

curl http://localhost:4000/v1/models

# Chat completion through the routing layer

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard","messages":[{"role":"user","content":"hello"}]}'

```

A successful chat request returns a completion routed through your configured algorithm (in this example, the LLM classifier selects either the weak or strong target based on input complexity).

## Configuring Routing Algorithms

The [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md) file documents multiple strategies for LLM traffic orchestration. Update the `type` field in your [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) to switch between algorithms.

### LLM-as-Classifier Routing

The `llm_classifier` type uses a lightweight model to determine whether to route to a cost-effective "weak" model or a high-capability "strong" model:

```toml
[routes.smart]
id = "switchyard"
type = "llm_classifier"
mode = "capability"
classifier_target = "weak"
strong_target = "strong"
weak_target = "weak"
base_threshold = 0.5

```

### Random A/B Testing

Use `random` routing for fixed traffic splits between model versions or providers:

```toml
[routes.ab_test]
id = "switchyard"
type = "random"
weights = { weak = 0.7, strong = 0.3 }
weak_target = "weak"
strong_target = "strong"

```

This configuration sends 70% of traffic to the weak target and 30% to the strong target, ideal for benchmarking prompt performance across model tiers.

### Stage Router and Escalation Patterns

For complex workflows involving tool use or error recovery, use `stage_router` to route based on intermediate results. For cost-sensitive escalation patterns, use `llm_classifier` with `mode = "escalation"` to run a cheap tier first, then conditionally retry with an expensive model based on a judge's evaluation.

## Monitoring and Observability

### Prometheus Metrics Integration

Switchyard automatically emits Prometheus-compatible metrics at the `/metrics` endpoint. Configure your Prometheus server to scrape `http://localhost:4000/metrics` to collect:

- `switchyard_requests_total`: Counter of requests by route and target
- `switchyard_latency_seconds`: Histogram of end-to-end request latency
- `switchyard_routing_overhead_seconds`: Time spent in routing logic

Refer to [`examples/prometheus/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/prometheus/README.md) for complete scrape configuration examples and Grafana dashboard templates.

## Library Path Integration

For Rust applications requiring embedded routing without the HTTP overhead, import the `switchyard-libsy` crate directly:

```toml
[dependencies]
switchyard-libsy = "0.1"

```

This provides programmatic access to the same routing algorithms available in the server path, allowing you to build custom LLM traffic orchestration into your application binary. The [`crates/libsy/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md) contains the API documentation for the `Router` and `Target` trait implementations.

## Summary

- **Switchyard** provides protocol translation between OpenAI and Anthropic formats using a provider-neutral IR defined in `crates/protocol`.
- Install the server via `cargo install --locked switchyard-server` and configure routing logic in a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file.
- Choose from multiple routing algorithms including random splitting, LLM-based classification, and escalation patterns documented in [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md).
- Validate configurations safely using the `--dry-run` flag before exposing network services.
- Monitor traffic and latency through the built-in Prometheus metrics endpoint at `/metrics`.

## Frequently Asked Questions

### What is the difference between the server path and library path in Switchyard?

The **server path** runs Switchyard as a standalone HTTP proxy that any OpenAI-compatible client can connect to via HTTP requests, making it ideal for polyglot environments with Python, JavaScript, or CLI tools. The **library path** (`switchyard-libsy`) embeds the routing logic directly into your Rust application as a library dependency, eliminating HTTP overhead but requiring Rust as the implementation language.

### How does Switchyard handle API key security?

Switchyard never stores API keys in the [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration file. Instead, you specify an `api_key_env` parameter that names an environment variable; the server reads the secret from the process environment at startup. This approach prevents credential leakage in version control and allows secure injection via Kubernetes secrets or Docker environment files.

### Can Switchyard route between different LLM providers (e.g., OpenAI and Anthropic)?

Yes. Switchyard's protocol translation layer in `crates/protocol` normalizes requests into a provider-neutral format, allowing you to define targets with different `format` values (e.g., `openai_chat` for one target and `anthropic_messages` for another) within the same [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration. The proxy automatically converts requests and responses between the client's expected format and each backend's native protocol.

### What happens if a routed request fails?

Switchyard implements automatic fallback handling as documented in the operations guides. If a selected target returns an error or times out, the proxy can retry the request against a default fallback target defined in your configuration. Additionally, the system emits Prometheus metrics tracking failure rates by target, enabling alerting on backend health degradation.