How to Set Up Switchyard as an LLM Traffic Orchestration Proxy

Switchyard is a Rust-based proxy that intercepts OpenAI and Anthropic SDK requests, translates them into a provider-neutral internal representation, and routes them to optimal backend LLMs using configurable algorithms like random split, LLM-as-classifier, or escalation routing.

Setting up Switchyard as an LLM traffic orchestration proxy allows you to sit between client applications and backend providers to perform intelligent load balancing, cost optimization, and failover handling. The NVIDIA-NeMo/Switchyard repository provides both a standalone server binary for HTTP proxying and a library crate for embedded Rust applications.

What Is Switchyard?

Switchyard functions as a protocol-aware gateway that normalizes disparate LLM API formats into a unified routing layer. According to the architecture documentation in the NVIDIA-NeMo/Switchyard repository, the system performs three core functions: protocol translation between OpenAI Chat, OpenAI Responses, and Anthropic Messages formats; intelligent routing using pluggable algorithms; and metrics emission with automatic fallback handling.

Core Architecture Components

The proxy operates through a provider-neutral intermediate representation (IR) defined in crates/protocol/README.md. When a client sends a request, Switchyard:

  1. Translates the native format (OpenAI or Anthropic) into the internal IR
  2. Applies the configured routing algorithm to select a backend target
  3. Converts the request back to the target's native format
  4. Returns the response to the client in their original expected format

As implemented in crates/switchyard-server, the server automatically exposes health endpoints and Prometheus metrics at /metrics for observability into request counts, latencies, and token usage.

Execution Paths: Server vs. Library

Switchyard supports two distinct deployment patterns:

  • Server path (switchyard-server binary): Use this when you need a standalone HTTP proxy that any OpenAI-compatible client can call without code changes. This is the most common setup for LLM traffic orchestration.
  • Library path (switchyard-libsy crate): Use this when building a custom Rust application that needs embedded routing logic without the HTTP layer overhead.

This guide focuses on the server path, which provides immediate integration with existing Python, JavaScript, or CLI-based LLM clients.

Prerequisites

Before installing Switchyard, ensure your environment meets the following requirements:

  • Rust toolchain ≥ 1.96 with Cargo (install via rustup)
  • API credentials for an OpenAI-compatible provider (OpenRouter, OpenAI, Anthropic, or Azure OpenAI)
  • Network access to your chosen LLM backend endpoints

The crates/switchyard-server/README.md file specifies that the server binds to TCP ports by default and requires environment variables for API key injection rather than hardcoded secrets.

Installation and Basic Configuration

Install the Server Binary

Install the Switchyard server using Cargo's locked installation to ensure reproducible builds:

cargo install --locked switchyard-server

This command compiles the release binary and places it in ~/.cargo/bin. Verify the installation by checking the help output:

switchyard-server --help

Create the routes.toml Configuration

Switchyard reads a TOML configuration file that declares LLM clients, backend targets, and routing rules. Create a file named routes.toml in your working directory with the following structure:

schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[routes.smart]
id = "switchyard"
type = "llm_classifier"
mode = "capability"
classifier_target = "weak"
strong_target = "strong"
weak_target = "weak"
base_threshold = 0.5

Key configuration parameters:

  • format: Specifies the upstream protocol (openai_chat, openai_responses, or anthropic_messages)
  • api_key_env: Names the environment variable containing the provider secret (secrets never appear in the configuration file)
  • type: Defines the routing algorithm (see routing options below)

Validate with Dry-Run

Before opening network sockets, validate your configuration syntax and environment variable presence:

export OPENROUTER_API_KEY="your-api-key-here"
switchyard-server --config routes.toml --dry-run

The --dry-run flag parses the TOML schema, checks that referenced environment variables exist, validates target references, and reports configuration errors without starting the HTTP server.

Running the LLM Traffic Orchestration Proxy

Start the Server

Launch the proxy with explicit host and port bindings:

switchyard-server --config routes.toml \
  --host 127.0.0.1 \
  --port 4000

The server now listens on http://127.0.0.1:4000 and is ready to accept OpenAI-compatible requests. For production deployments, bind to 0.0.0.0 or a specific interface as documented in docs/getting_started.md.

Verify the Proxy Endpoints

Test the running server using standard HTTP clients:


# Health check

curl http://localhost:4000/health

# Expected: {"status":"ok"}

# List available models (includes your route IDs)

curl http://localhost:4000/v1/models

# Chat completion through the routing layer

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard","messages":[{"role":"user","content":"hello"}]}'

A successful chat request returns a completion routed through your configured algorithm (in this example, the LLM classifier selects either the weak or strong target based on input complexity).

Configuring Routing Algorithms

The docs/routing_algorithms/overview.md file documents multiple strategies for LLM traffic orchestration. Update the type field in your routes.toml to switch between algorithms.

LLM-as-Classifier Routing

The llm_classifier type uses a lightweight model to determine whether to route to a cost-effective "weak" model or a high-capability "strong" model:

[routes.smart]
id = "switchyard"
type = "llm_classifier"
mode = "capability"
classifier_target = "weak"
strong_target = "strong"
weak_target = "weak"
base_threshold = 0.5

Random A/B Testing

Use random routing for fixed traffic splits between model versions or providers:

[routes.ab_test]
id = "switchyard"
type = "random"
weights = { weak = 0.7, strong = 0.3 }
weak_target = "weak"
strong_target = "strong"

This configuration sends 70% of traffic to the weak target and 30% to the strong target, ideal for benchmarking prompt performance across model tiers.

Stage Router and Escalation Patterns

For complex workflows involving tool use or error recovery, use stage_router to route based on intermediate results. For cost-sensitive escalation patterns, use llm_classifier with mode = "escalation" to run a cheap tier first, then conditionally retry with an expensive model based on a judge's evaluation.

Monitoring and Observability

Prometheus Metrics Integration

Switchyard automatically emits Prometheus-compatible metrics at the /metrics endpoint. Configure your Prometheus server to scrape http://localhost:4000/metrics to collect:

  • switchyard_requests_total: Counter of requests by route and target
  • switchyard_latency_seconds: Histogram of end-to-end request latency
  • switchyard_routing_overhead_seconds: Time spent in routing logic

Refer to examples/prometheus/README.md for complete scrape configuration examples and Grafana dashboard templates.

Library Path Integration

For Rust applications requiring embedded routing without the HTTP overhead, import the switchyard-libsy crate directly:

[dependencies]
switchyard-libsy = "0.1"

This provides programmatic access to the same routing algorithms available in the server path, allowing you to build custom LLM traffic orchestration into your application binary. The crates/libsy/README.md contains the API documentation for the Router and Target trait implementations.

Summary

  • Switchyard provides protocol translation between OpenAI and Anthropic formats using a provider-neutral IR defined in crates/protocol.
  • Install the server via cargo install --locked switchyard-server and configure routing logic in a routes.toml file.
  • Choose from multiple routing algorithms including random splitting, LLM-based classification, and escalation patterns documented in docs/routing_algorithms/overview.md.
  • Validate configurations safely using the --dry-run flag before exposing network services.
  • Monitor traffic and latency through the built-in Prometheus metrics endpoint at /metrics.

Frequently Asked Questions

What is the difference between the server path and library path in Switchyard?

The server path runs Switchyard as a standalone HTTP proxy that any OpenAI-compatible client can connect to via HTTP requests, making it ideal for polyglot environments with Python, JavaScript, or CLI tools. The library path (switchyard-libsy) embeds the routing logic directly into your Rust application as a library dependency, eliminating HTTP overhead but requiring Rust as the implementation language.

How does Switchyard handle API key security?

Switchyard never stores API keys in the routes.toml configuration file. Instead, you specify an api_key_env parameter that names an environment variable; the server reads the secret from the process environment at startup. This approach prevents credential leakage in version control and allows secure injection via Kubernetes secrets or Docker environment files.

Can Switchyard route between different LLM providers (e.g., OpenAI and Anthropic)?

Yes. Switchyard's protocol translation layer in crates/protocol normalizes requests into a provider-neutral format, allowing you to define targets with different format values (e.g., openai_chat for one target and anthropic_messages for another) within the same routes.toml configuration. The proxy automatically converts requests and responses between the client's expected format and each backend's native protocol.

What happens if a routed request fails?

Switchyard implements automatic fallback handling as documented in the operations guides. If a selected target returns an error or times out, the proxy can retry the request against a default fallback target defined in your configuration. Additionally, the system emits Prometheus metrics tracking failure rates by target, enabling alerting on backend health degradation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →