# How to Build and Run the Standalone Switchyard HTTP Proxy Server

> Learn to build and run the standalone Switchyard HTTP proxy server. Install via cargo or build from source, configure with TOML, and route OpenAI and Anthropic API requests using this native Rust binary.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-21

---

**The standalone Switchyard HTTP proxy server is a native Rust binary that you can install via `cargo install --locked switchyard-server` or build from source, configure using a TOML deployment file, and launch with the `switchyard-server` CLI to route OpenAI and Anthropic API requests.**

This guide covers the complete workflow for deploying the **switchyard-server** binary from the [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard) repository. You will learn how to compile the standalone HTTP proxy, structure the routing configuration, and start the server with the correct CLI flags.

## Architecture Overview

The `switchyard-server` binary implements a three-layer architecture designed for high-performance LLM request routing.

**HTTP Entry Point** – The binary in [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs) uses Tokio to parse CLI arguments, initialize observability hooks, and bind to the configured network interface. This entry point delegates request handling to the core library crates.

**Routing Core** – The `switchyard-libsy` library implements routing algorithms including random selection, LLM classifiers, and stage routers. The server loads your TOML configuration at startup, constructs an in-memory routing table, and invokes these algorithms for each incoming request based on the `model` parameter.

**Protocol Translation** – The `switchyard-protocol` crate defines provider-neutral request/response types, while `switchyard-translation` handles conversion between OpenAI, Anthropic, and native formats. This allows the proxy to accept OpenAI-compatible calls and forward them to Anthropic endpoints (or vice versa) without client-side changes.

The binary exposes the following endpoints by default:

- `POST /v1/chat/completions` – OpenAI Chat Completions
- `POST /v1/messages` – Anthropic Messages  
- `POST /v1/responses` – OpenAI Responses
- `GET /v1/models` – List configured routes
- `GET /metrics` – Prometheus metrics
- `GET /health` – Liveness probe

## Prerequisites and Installation

Switchyard requires a Rust toolchain. If you do not have Rust installed, use `rustup`:

```bash
sudo apt-get update && sudo apt-get install -y build-essential curl git
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
source "$HOME/.cargo/env"
rustc --version && cargo --version

```

### Install from Crates.io

The simplest method uses Cargo to fetch and compile the locked release:

```bash
cargo install --locked switchyard-server

```

This command places the `switchyard-server` binary in `~/.cargo/bin`. Ensure this directory is in your `PATH` before proceeding.

### Build from Source

To modify the server or use the latest development version, clone the repository and build the release target:

```bash
git clone https://github.com/NVIDIA-NeMo/Switchyard.git
cd Switchyard
cargo build --release -p switchyard-server

```

The compiled binary will be available at `target/release/switchyard-server`. Building from source requires the same Rust toolchain prerequisites and takes longer due to full compilation of the workspace crates.

## Configuring the Proxy Server

The server requires a TOML deployment file that defines **LLM clients**, **targets**, and **routes**. Create a file named [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) with the following structure:

```toml
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[routes.smart]
id = "switchyard"
type = "llm_classifier"
mode = "capability"
classifier_target = "weak"
strong_target = "strong"
weak_target = "weak"
base_threshold = 0.5

```

Key configuration concepts:

- **`llm_clients`** – Defines provider connections and the environment variable containing the API key (never hardcode secrets).
- **`targets`** – Maps to specific model IDs (e.g., `openai/gpt-4o`) that the router can select.
- **`routes`** – Defines routing logic. The `id` field (e.g., `switchyard`) is the value clients must pass in the `model` parameter.

For the complete schema specification, see the [TOML Schema documentation](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md) and the server README at [`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md).

## Running the Standalone Server

Before starting the server, export the API key referenced in your TOML configuration:

```bash
export OPENROUTER_API_KEY="your-openrouter-key"

```

### Validate Configuration

Run a dry-run check to verify the TOML syntax and routing table without binding to a port:

```bash
switchyard-server --config routes.toml --dry-run

```

This command parses the deployment file, initializes clients, and reports errors without starting the HTTP listener.

### Start the Server

Launch the proxy with explicit host and port bindings. The default host is `0.0.0.0`, but you can restrict it to localhost for development:

```bash
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

For production deployments, bind to all interfaces and enable TLS if required. Additional flags for certificate paths and shutdown behavior are documented in the [CLI Reference](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/cli_reference.md).

### Verify the Deployment

Test the health endpoint and model listing in a separate terminal:

```bash
curl http://localhost:4000/health
curl http://localhost:4000/v1/models

```

Send a test chat completion request to verify end-to-end routing:

```bash
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"switchyard","messages":[{"role":"user","content":"Explain the architecture"}]}'

```

If configured correctly, the server routes this request according to the classifier logic defined in [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) and returns the response in OpenAI-compatible JSON format.

## Summary

- The **switchyard-server** binary is a Rust-native HTTP proxy that translates between OpenAI and Anthropic API formats.
- Install via `cargo install --locked switchyard-server` or build from source using `cargo build --release -p switchyard-server`.
- Configuration uses a TOML file defining clients, targets, and routes; secrets are passed via environment variables referenced by `api_key_env`.
- Validate configurations with `--dry-run` before starting the server to catch schema errors early.
- The entry point at [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs) handles CLI parsing and delegates to the routing core in `switchyard-libsy`.

## Frequently Asked Questions

### What are the system requirements for building switchyard-server?

You need a stable Rust toolchain (1.70 or later recommended) and standard build tools (`build-essential`, `git`, `curl`). The project uses Cargo workspaces, so the build process compiles multiple crates including `switchyard-protocol`, `switchyard-translation`, and `switchyard-libsy` before producing the final binary.

### How does the TOML configuration define API authentication?

The TOML file specifies an `api_key_env` field in the `[llm_clients]` section that names an environment variable. The server reads the actual key from this variable at runtime, keeping credentials out of version control. For example, `api_key_env = "OPENROUTER_API_KEY"` requires you to run `export OPENROUTER_API_KEY="sk-..."` before starting the server.

### Can the proxy handle both OpenAI and Anthropic requests simultaneously?

Yes. The `switchyard-translation` crate automatically converts request and response formats based on the `format` field defined in the client configuration. You can define multiple clients (one for OpenAI, one for Anthropic) and route requests to either provider using the same OpenAI-compatible endpoint structure on the client side.

### Does switchyard-server support HTTPS/TLS termination?

Yes, though the documentation recommends using a reverse proxy (like Nginx or Envoy) for TLS termination in production. The server CLI supports flags for TLS certificate and key paths as documented in the CLI reference. Alternatively, you can run the server behind a load balancer that handles TLS, exposing only HTTP internally.