How to Deploy Switchyard Models: Server and Library Configuration Guide

Deploy Switchyard models by defining LLM clients, targets, and routes in a TOML configuration file, then run the switchyard-server binary or embed the switchyard-libsy library in your Rust application.

Switchyard is an LLM traffic proxy maintained by NVIDIA that sits between client applications and model backends. According to the NVIDIA-NeMo/Switchyard repository, deploying Switchyard involves configuring a native TOML deployment file that specifies how to route requests between efficient and capable targets while translating provider formats. The stable client-facing API allows you to switch underlying models without changing application code.

Architecture Overview

Switchyard deployments consist of three distinct layers defined in docs/architecture.md. LLM clients define wire formats (openai_chat, openai_responses, anthropic_messages) and authentication for upstream providers like OpenAI or Anthropic. Targets map concrete model identifiers (e.g., openai/gpt-4o) to specific LLM clients. Routes declare the routing algorithm (auto, random, llm_classifier, stage_router) and designate which target is efficient versus capable for intelligent traffic splitting.

Prerequisites

Before you deploy Switchyard models, install the following dependencies:

  • Git
  • A C/C++ toolchain
  • Rust (via rustup)

These tools are required to compile the Rust-based server binary from the crates/switchyard-server source.

Server Path Deployment

The server path is the most common deployment method for production use. Install the standalone switchyard-server binary and configure it via TOML.

Install the Server Binary

Compile and install the server using Cargo:

cargo install --locked switchyard-server

This installs the binary from the crates/switchyard-server package as documented in crates/switchyard-server/README.md.

Create the Deployment TOML

Create a file named routes.toml that defines your LLM clients, targets, and routing logic:

schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[routes.smart]
id = "switchyard"
type = "auto"
capable_target = "strong"
efficient_target = "weak"

This configuration routes traffic to OpenRouter, defining a cost-efficient mini model and a capable full-scale model, then applies the auto routing strategy.

Validate the Configuration

Validate your TOML file and environment variables without starting the server:

export OPENROUTER_API_KEY="your-openrouter-key"
switchyard-server --config routes.toml --dry-run

The --dry-run flag parses the configuration and checks for schema errors or missing environment variables without binding to a network port, as specified in docs/cli_reference.md.

Start the Proxy Server

Launch the server to expose an OpenAI-compatible API:

switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

Optional flags include --tls-cert and --tls-key for TLS termination, and --routing-log for traffic audit trails.

Test the Deployment

Verify the deployment using any OpenAI-compatible client:

curl http://localhost:4000/v1/models

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard","messages":[{"role":"user","content":"Hello"}]}'

Library Path Deployment

For Rust applications requiring embedded routing logic, use the library path instead of the standalone server. Add the following dependencies to your Cargo.toml:

[dependencies]
async-trait = "0.1"
futures = "0.3"
switchyard-libsy = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
switchyard-protocol = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
tokio = { version = "1", features = ["macros", "rt"] }

The switchyard-libsy crate provides the core routing algorithm implemented in crates/switchyard-libsy/src/, allowing you to drive traffic management directly from your application code. See docs/getting_started.md for integration examples.

Configuration Reference

The complete TOML schema specification resides in docs/reference/toml_schema.md. This document defines valid fields for:

  • LLM clients: Provider endpoints, API key environment variables, and request formats
  • Targets: Model IDs and client associations
  • Routes: Algorithm types, target weighting, and fallback chains

For a practical integration example with LiteLLM, consult examples/litellm/README.md.

Summary

  • Switchyard deployments require a TOML file with three layers: LLM clients, targets, and routes
  • Use the server path (switchyard-server binary) for standalone proxy deployments
  • Use the library path (switchyard-libsy crate) to embed routing logic in Rust applications
  • Always validate configurations with --dry-run before exposing network ports
  • The proxy exposes an OpenAI-compatible API at configurable host/port endpoints

Frequently Asked Questions

What is the difference between efficient_target and capable_target in Switchyard routes?

The efficient_target designates the model for cost-effective, high-throughput inference, while the capable_target specifies the model for complex requests requiring higher accuracy. When using the auto routing algorithm, Switchyard intelligently selects between these targets based on request characteristics defined in docs/reference/toml_schema.md.

How do I validate a Switchyard deployment without exposing network ports?

Run switchyard-server --config routes.toml --dry-run to perform schema validation and environment variable checks without binding a socket or accepting traffic. This command validates the TOML structure and ensures all referenced environment variables like OPENROUTER_API_KEY are present.

Can I deploy Switchyard without installing the standalone server binary?

Yes. Import the switchyard-libsy crate directly into your Rust application to embed the routing engine. This library path, documented in docs/getting_started.md, allows programmatic control over routing decisions without running a separate proxy process.

Where is the complete TOML schema for Switchyard deployments documented?

The authoritative specification resides in docs/reference/toml_schema.md within the NVIDIA-NeMo/Switchyard repository. This file documents all valid configuration sections, field types, and routing algorithm parameters for production deployments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →