How to Set Up Switchyard as an LLM Traffic Orchestration Proxy
Switchyard is a Rust-based proxy that intercepts OpenAI and Anthropic SDK requests, translates them into a provider-neutral internal representation, and routes them to optimal backend LLMs using configurable algorithms like random split, LLM-as-classifier, or escalation routing.
Setting up Switchyard as an LLM traffic orchestration proxy allows you to sit between client applications and backend providers to perform intelligent load balancing, cost optimization, and failover handling. The NVIDIA-NeMo/Switchyard repository provides both a standalone server binary for HTTP proxying and a library crate for embedded Rust applications.
What Is Switchyard?
Switchyard functions as a protocol-aware gateway that normalizes disparate LLM API formats into a unified routing layer. According to the architecture documentation in the NVIDIA-NeMo/Switchyard repository, the system performs three core functions: protocol translation between OpenAI Chat, OpenAI Responses, and Anthropic Messages formats; intelligent routing using pluggable algorithms; and metrics emission with automatic fallback handling.
Core Architecture Components
The proxy operates through a provider-neutral intermediate representation (IR) defined in crates/protocol/README.md. When a client sends a request, Switchyard:
- Translates the native format (OpenAI or Anthropic) into the internal IR
- Applies the configured routing algorithm to select a backend target
- Converts the request back to the target's native format
- Returns the response to the client in their original expected format
As implemented in crates/switchyard-server, the server automatically exposes health endpoints and Prometheus metrics at /metrics for observability into request counts, latencies, and token usage.
Execution Paths: Server vs. Library
Switchyard supports two distinct deployment patterns:
- Server path (
switchyard-serverbinary): Use this when you need a standalone HTTP proxy that any OpenAI-compatible client can call without code changes. This is the most common setup for LLM traffic orchestration. - Library path (
switchyard-libsycrate): Use this when building a custom Rust application that needs embedded routing logic without the HTTP layer overhead.
This guide focuses on the server path, which provides immediate integration with existing Python, JavaScript, or CLI-based LLM clients.
Prerequisites
Before installing Switchyard, ensure your environment meets the following requirements:
- Rust toolchain ≥ 1.96 with Cargo (install via rustup)
- API credentials for an OpenAI-compatible provider (OpenRouter, OpenAI, Anthropic, or Azure OpenAI)
- Network access to your chosen LLM backend endpoints
The crates/switchyard-server/README.md file specifies that the server binds to TCP ports by default and requires environment variables for API key injection rather than hardcoded secrets.
Installation and Basic Configuration
Install the Server Binary
Install the Switchyard server using Cargo's locked installation to ensure reproducible builds:
cargo install --locked switchyard-server
This command compiles the release binary and places it in ~/.cargo/bin. Verify the installation by checking the help output:
switchyard-server --help
Create the routes.toml Configuration
Switchyard reads a TOML configuration file that declares LLM clients, backend targets, and routing rules. Create a file named routes.toml in your working directory with the following structure:
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"
[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"
[routes.smart]
id = "switchyard"
type = "llm_classifier"
mode = "capability"
classifier_target = "weak"
strong_target = "strong"
weak_target = "weak"
base_threshold = 0.5
Key configuration parameters:
format: Specifies the upstream protocol (openai_chat,openai_responses, oranthropic_messages)api_key_env: Names the environment variable containing the provider secret (secrets never appear in the configuration file)type: Defines the routing algorithm (see routing options below)
Validate with Dry-Run
Before opening network sockets, validate your configuration syntax and environment variable presence:
export OPENROUTER_API_KEY="your-api-key-here"
switchyard-server --config routes.toml --dry-run
The --dry-run flag parses the TOML schema, checks that referenced environment variables exist, validates target references, and reports configuration errors without starting the HTTP server.
Running the LLM Traffic Orchestration Proxy
Start the Server
Launch the proxy with explicit host and port bindings:
switchyard-server --config routes.toml \
--host 127.0.0.1 \
--port 4000
The server now listens on http://127.0.0.1:4000 and is ready to accept OpenAI-compatible requests. For production deployments, bind to 0.0.0.0 or a specific interface as documented in docs/getting_started.md.
Verify the Proxy Endpoints
Test the running server using standard HTTP clients:
# Health check
curl http://localhost:4000/health
# Expected: {"status":"ok"}
# List available models (includes your route IDs)
curl http://localhost:4000/v1/models
# Chat completion through the routing layer
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"switchyard","messages":[{"role":"user","content":"hello"}]}'
A successful chat request returns a completion routed through your configured algorithm (in this example, the LLM classifier selects either the weak or strong target based on input complexity).
Configuring Routing Algorithms
The docs/routing_algorithms/overview.md file documents multiple strategies for LLM traffic orchestration. Update the type field in your routes.toml to switch between algorithms.
LLM-as-Classifier Routing
The llm_classifier type uses a lightweight model to determine whether to route to a cost-effective "weak" model or a high-capability "strong" model:
[routes.smart]
id = "switchyard"
type = "llm_classifier"
mode = "capability"
classifier_target = "weak"
strong_target = "strong"
weak_target = "weak"
base_threshold = 0.5
Random A/B Testing
Use random routing for fixed traffic splits between model versions or providers:
[routes.ab_test]
id = "switchyard"
type = "random"
weights = { weak = 0.7, strong = 0.3 }
weak_target = "weak"
strong_target = "strong"
This configuration sends 70% of traffic to the weak target and 30% to the strong target, ideal for benchmarking prompt performance across model tiers.
Stage Router and Escalation Patterns
For complex workflows involving tool use or error recovery, use stage_router to route based on intermediate results. For cost-sensitive escalation patterns, use llm_classifier with mode = "escalation" to run a cheap tier first, then conditionally retry with an expensive model based on a judge's evaluation.
Monitoring and Observability
Prometheus Metrics Integration
Switchyard automatically emits Prometheus-compatible metrics at the /metrics endpoint. Configure your Prometheus server to scrape http://localhost:4000/metrics to collect:
switchyard_requests_total: Counter of requests by route and targetswitchyard_latency_seconds: Histogram of end-to-end request latencyswitchyard_routing_overhead_seconds: Time spent in routing logic
Refer to examples/prometheus/README.md for complete scrape configuration examples and Grafana dashboard templates.
Library Path Integration
For Rust applications requiring embedded routing without the HTTP overhead, import the switchyard-libsy crate directly:
[dependencies]
switchyard-libsy = "0.1"
This provides programmatic access to the same routing algorithms available in the server path, allowing you to build custom LLM traffic orchestration into your application binary. The crates/libsy/README.md contains the API documentation for the Router and Target trait implementations.
Summary
- Switchyard provides protocol translation between OpenAI and Anthropic formats using a provider-neutral IR defined in
crates/protocol. - Install the server via
cargo install --locked switchyard-serverand configure routing logic in aroutes.tomlfile. - Choose from multiple routing algorithms including random splitting, LLM-based classification, and escalation patterns documented in
docs/routing_algorithms/overview.md. - Validate configurations safely using the
--dry-runflag before exposing network services. - Monitor traffic and latency through the built-in Prometheus metrics endpoint at
/metrics.
Frequently Asked Questions
What is the difference between the server path and library path in Switchyard?
The server path runs Switchyard as a standalone HTTP proxy that any OpenAI-compatible client can connect to via HTTP requests, making it ideal for polyglot environments with Python, JavaScript, or CLI tools. The library path (switchyard-libsy) embeds the routing logic directly into your Rust application as a library dependency, eliminating HTTP overhead but requiring Rust as the implementation language.
How does Switchyard handle API key security?
Switchyard never stores API keys in the routes.toml configuration file. Instead, you specify an api_key_env parameter that names an environment variable; the server reads the secret from the process environment at startup. This approach prevents credential leakage in version control and allows secure injection via Kubernetes secrets or Docker environment files.
Can Switchyard route between different LLM providers (e.g., OpenAI and Anthropic)?
Yes. Switchyard's protocol translation layer in crates/protocol normalizes requests into a provider-neutral format, allowing you to define targets with different format values (e.g., openai_chat for one target and anthropic_messages for another) within the same routes.toml configuration. The proxy automatically converts requests and responses between the client's expected format and each backend's native protocol.
What happens if a routed request fails?
Switchyard implements automatic fallback handling as documented in the operations guides. If a selected target returns an error or times out, the proxy can retry the request against a default fallback target defined in your configuration. Additionally, the system emits Prometheus metrics tracking failure rates by target, enabling alerting on backend health degradation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →