# How to Deploy Switchyard Models: Server and Library Configuration Guide

> Learn to deploy Switchyard models efficiently. Configure clients, targets, and routes in TOML, then run the server or embed the library in your Rust app.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Deploy Switchyard models by defining LLM clients, targets, and routes in a TOML configuration file, then run the `switchyard-server` binary or embed the `switchyard-libsy` library in your Rust application.**

Switchyard is an LLM traffic proxy maintained by NVIDIA that sits between client applications and model backends. According to the NVIDIA-NeMo/Switchyard repository, deploying Switchyard involves configuring a native TOML deployment file that specifies how to route requests between efficient and capable targets while translating provider formats. The stable client-facing API allows you to switch underlying models without changing application code.

## Architecture Overview

Switchyard deployments consist of three distinct layers defined in [`docs/architecture.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/architecture.md). **LLM clients** define wire formats (`openai_chat`, `openai_responses`, `anthropic_messages`) and authentication for upstream providers like OpenAI or Anthropic. **Targets** map concrete model identifiers (e.g., `openai/gpt-4o`) to specific LLM clients. **Routes** declare the routing algorithm (`auto`, `random`, `llm_classifier`, `stage_router`) and designate which target is *efficient* versus *capable* for intelligent traffic splitting.

## Prerequisites

Before you deploy Switchyard models, install the following dependencies:

- Git
- A C/C++ toolchain
- Rust (via `rustup`)

These tools are required to compile the Rust-based server binary from the `crates/switchyard-server` source.

## Server Path Deployment

The server path is the most common deployment method for production use. Install the standalone `switchyard-server` binary and configure it via TOML.

### Install the Server Binary

Compile and install the server using Cargo:

```bash
cargo install --locked switchyard-server

```

This installs the binary from the `crates/switchyard-server` package as documented in [`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md).

### Create the Deployment TOML

Create a file named [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) that defines your LLM clients, targets, and routing logic:

```toml
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[routes.smart]
id = "switchyard"
type = "auto"
capable_target = "strong"
efficient_target = "weak"

```

This configuration routes traffic to OpenRouter, defining a cost-efficient mini model and a capable full-scale model, then applies the `auto` routing strategy.

### Validate the Configuration

Validate your TOML file and environment variables without starting the server:

```bash
export OPENROUTER_API_KEY="your-openrouter-key"
switchyard-server --config routes.toml --dry-run

```

The `--dry-run` flag parses the configuration and checks for schema errors or missing environment variables without binding to a network port, as specified in [`docs/cli_reference.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/cli_reference.md).

### Start the Proxy Server

Launch the server to expose an OpenAI-compatible API:

```bash
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

Optional flags include `--tls-cert` and `--tls-key` for TLS termination, and `--routing-log` for traffic audit trails.

### Test the Deployment

Verify the deployment using any OpenAI-compatible client:

```bash
curl http://localhost:4000/v1/models

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard","messages":[{"role":"user","content":"Hello"}]}'

```

## Library Path Deployment

For Rust applications requiring embedded routing logic, use the library path instead of the standalone server. Add the following dependencies to your [`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml):

```toml
[dependencies]
async-trait = "0.1"
futures = "0.3"
switchyard-libsy = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
switchyard-protocol = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
tokio = { version = "1", features = ["macros", "rt"] }

```

The `switchyard-libsy` crate provides the core routing algorithm implemented in `crates/switchyard-libsy/src/`, allowing you to drive traffic management directly from your application code. See [`docs/getting_started.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/getting_started.md) for integration examples.

## Configuration Reference

The complete TOML schema specification resides in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md). This document defines valid fields for:

- **LLM clients**: Provider endpoints, API key environment variables, and request formats
- **Targets**: Model IDs and client associations
- **Routes**: Algorithm types, target weighting, and fallback chains

For a practical integration example with LiteLLM, consult [`examples/litellm/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/README.md).

## Summary

- Switchyard deployments require a TOML file with three layers: **LLM clients**, **targets**, and **routes**
- Use the **server path** (`switchyard-server` binary) for standalone proxy deployments
- Use the **library path** (`switchyard-libsy` crate) to embed routing logic in Rust applications
- Always validate configurations with `--dry-run` before exposing network ports
- The proxy exposes an **OpenAI-compatible API** at configurable host/port endpoints

## Frequently Asked Questions

### What is the difference between `efficient_target` and `capable_target` in Switchyard routes?

The `efficient_target` designates the model for cost-effective, high-throughput inference, while the `capable_target` specifies the model for complex requests requiring higher accuracy. When using the `auto` routing algorithm, Switchyard intelligently selects between these targets based on request characteristics defined in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md).

### How do I validate a Switchyard deployment without exposing network ports?

Run `switchyard-server --config routes.toml --dry-run` to perform schema validation and environment variable checks without binding a socket or accepting traffic. This command validates the TOML structure and ensures all referenced environment variables like `OPENROUTER_API_KEY` are present.

### Can I deploy Switchyard without installing the standalone server binary?

Yes. Import the `switchyard-libsy` crate directly into your Rust application to embed the routing engine. This library path, documented in [`docs/getting_started.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/getting_started.md), allows programmatic control over routing decisions without running a separate proxy process.

### Where is the complete TOML schema for Switchyard deployments documented?

The authoritative specification resides in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md) within the NVIDIA-NeMo/Switchyard repository. This file documents all valid configuration sections, field types, and routing algorithm parameters for production deployments.