How to Deploy Switchyard Models: Server and Library Configuration Guide
Deploy Switchyard models by defining LLM clients, targets, and routes in a TOML configuration file, then run the switchyard-server binary or embed the switchyard-libsy library in your Rust application.
Switchyard is an LLM traffic proxy maintained by NVIDIA that sits between client applications and model backends. According to the NVIDIA-NeMo/Switchyard repository, deploying Switchyard involves configuring a native TOML deployment file that specifies how to route requests between efficient and capable targets while translating provider formats. The stable client-facing API allows you to switch underlying models without changing application code.
Architecture Overview
Switchyard deployments consist of three distinct layers defined in docs/architecture.md. LLM clients define wire formats (openai_chat, openai_responses, anthropic_messages) and authentication for upstream providers like OpenAI or Anthropic. Targets map concrete model identifiers (e.g., openai/gpt-4o) to specific LLM clients. Routes declare the routing algorithm (auto, random, llm_classifier, stage_router) and designate which target is efficient versus capable for intelligent traffic splitting.
Prerequisites
Before you deploy Switchyard models, install the following dependencies:
- Git
- A C/C++ toolchain
- Rust (via
rustup)
These tools are required to compile the Rust-based server binary from the crates/switchyard-server source.
Server Path Deployment
The server path is the most common deployment method for production use. Install the standalone switchyard-server binary and configure it via TOML.
Install the Server Binary
Compile and install the server using Cargo:
cargo install --locked switchyard-server
This installs the binary from the crates/switchyard-server package as documented in crates/switchyard-server/README.md.
Create the Deployment TOML
Create a file named routes.toml that defines your LLM clients, targets, and routing logic:
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"
[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"
[routes.smart]
id = "switchyard"
type = "auto"
capable_target = "strong"
efficient_target = "weak"
This configuration routes traffic to OpenRouter, defining a cost-efficient mini model and a capable full-scale model, then applies the auto routing strategy.
Validate the Configuration
Validate your TOML file and environment variables without starting the server:
export OPENROUTER_API_KEY="your-openrouter-key"
switchyard-server --config routes.toml --dry-run
The --dry-run flag parses the configuration and checks for schema errors or missing environment variables without binding to a network port, as specified in docs/cli_reference.md.
Start the Proxy Server
Launch the server to expose an OpenAI-compatible API:
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
Optional flags include --tls-cert and --tls-key for TLS termination, and --routing-log for traffic audit trails.
Test the Deployment
Verify the deployment using any OpenAI-compatible client:
curl http://localhost:4000/v1/models
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"switchyard","messages":[{"role":"user","content":"Hello"}]}'
Library Path Deployment
For Rust applications requiring embedded routing logic, use the library path instead of the standalone server. Add the following dependencies to your Cargo.toml:
[dependencies]
async-trait = "0.1"
futures = "0.3"
switchyard-libsy = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
switchyard-protocol = { git = "https://github.com/NVIDIA-NeMo/Switchyard.git", tag = "v0.2.0" }
tokio = { version = "1", features = ["macros", "rt"] }
The switchyard-libsy crate provides the core routing algorithm implemented in crates/switchyard-libsy/src/, allowing you to drive traffic management directly from your application code. See docs/getting_started.md for integration examples.
Configuration Reference
The complete TOML schema specification resides in docs/reference/toml_schema.md. This document defines valid fields for:
- LLM clients: Provider endpoints, API key environment variables, and request formats
- Targets: Model IDs and client associations
- Routes: Algorithm types, target weighting, and fallback chains
For a practical integration example with LiteLLM, consult examples/litellm/README.md.
Summary
- Switchyard deployments require a TOML file with three layers: LLM clients, targets, and routes
- Use the server path (
switchyard-serverbinary) for standalone proxy deployments - Use the library path (
switchyard-libsycrate) to embed routing logic in Rust applications - Always validate configurations with
--dry-runbefore exposing network ports - The proxy exposes an OpenAI-compatible API at configurable host/port endpoints
Frequently Asked Questions
What is the difference between efficient_target and capable_target in Switchyard routes?
The efficient_target designates the model for cost-effective, high-throughput inference, while the capable_target specifies the model for complex requests requiring higher accuracy. When using the auto routing algorithm, Switchyard intelligently selects between these targets based on request characteristics defined in docs/reference/toml_schema.md.
How do I validate a Switchyard deployment without exposing network ports?
Run switchyard-server --config routes.toml --dry-run to perform schema validation and environment variable checks without binding a socket or accepting traffic. This command validates the TOML structure and ensures all referenced environment variables like OPENROUTER_API_KEY are present.
Can I deploy Switchyard without installing the standalone server binary?
Yes. Import the switchyard-libsy crate directly into your Rust application to embed the routing engine. This library path, documented in docs/getting_started.md, allows programmatic control over routing decisions without running a separate proxy process.
Where is the complete TOML schema for Switchyard deployments documented?
The authoritative specification resides in docs/reference/toml_schema.md within the NVIDIA-NeMo/Switchyard repository. This file documents all valid configuration sections, field types, and routing algorithm parameters for production deployments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →