# How the switchyard-server Crate Functions as an HTTP Server for LLM Traffic

> Discover how the switchyard-server Rust crate acts as an HTTP server, efficiently proxying and routing LLM traffic with OpenAI/Anthropic API compatibility and Prometheus metrics.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-22

---

**The switchyard-server crate is a standalone Rust binary that exposes an HTTP proxy server translating between OpenAI/Anthropic API formats and backend-specific protocols, routing LLM traffic through configurable algorithms while emitting Prometheus metrics.**

The `switchyard-server` crate serves as the primary HTTP gateway for the NVIDIA-NeMo/Switchyard project, enabling seamless proxying of Large Language Model (LLM) requests between clients and diverse backends. Acting as a translator-router-proxy, it accepts standard OpenAI and Anthropic API requests and forwards them to configured model endpoints. This article examines the crate’s architecture, request lifecycle, and configuration based on the actual source implementation.

## Core Architecture and Entry Points

The crate operates as an asynchronous HTTP server built on the Axum framework, initializing through two primary source files.

### Server Initialization in main.rs

The entrypoint resides in [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs). This file parses command-line flags such as `--config`, `--host`, and `--port`, validates the [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration (optionally via `--dry-run`), and launches the Axum HTTP server. It binds to the configured address—typically `127.0.0.1:4000`—and prepares the routing table for incoming LLM traffic.

### Request Handling Logic in server.rs

The core request processing occurs in [`crates/switchyard-server/src/server.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/server.rs). This module implements the HTTP handler functions that intercept incoming requests, delegate routing decisions to the `switchyard-libsy` crate, and coordinate translation via `switchyard-translation`. It manages the full lifecycle from JSON deserialization to backend forwarding and response encoding.

## The Seven-Stage LLM Request Lifecycle

The server processes each HTTP request through a deterministic pipeline that bridges client APIs with backend implementations.

1. **Expose the HTTP Endpoint**: The server listens on a configurable host and port (default `4000`), presenting standard endpoints like `/v1/chat/completions`, `/v1/responses`, and `/messages` compatible with OpenAI and Anthropic specifications.

2. **Parse Provider-Neutral Requests**: Incoming JSON payloads deserialize into provider-neutral request types defined in [`crates/switchyard-protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-protocol/src/lib.rs). This abstraction decouples the HTTP interface from specific backend formats.

3. **Select Backend via Routing Algorithm**: Based on the [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration, the server invokes routing algorithms—`random`, `llm_classifier`, or `stage_router`—implemented in [`crates/switchyard-libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-libsy/src/lib.rs) to select the appropriate model backend.

4. **Translate to Backend Format**: The `switchyard-translation` crate converts the provider-neutral request into the backend-specific format required by targets such as vLLM, NVIDIA NIM, or Ollama. Translation logic resides in [`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs).

5. **Forward Asynchronous HTTP Calls**: The server transmits the translated request to the selected backend using an async HTTP client, maintaining streaming support for chunked responses.

6. **Translate and Encode Responses**: Backend responses (including streamed chunks) convert back to the provider-neutral response type, then re-encode into the original OpenAI or Anthropic JSON shape expected by the caller.

7. **Emit Prometheus Metrics**: Every request updates observability counters exposed at `/metrics`, including `switchyard_requests_total`, `switchyard_latency_seconds`, and `switchyard_routing_overhead_seconds`.

## Configuration and Routing Algorithms

### Defining Traffic Rules with routes.toml

The server relies on a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration file to map incoming requests to backend pools. This file specifies routing algorithms, backend URLs, and model aliases. The `--config` CLI flag points to this file during startup, with optional validation available through `--dry-run`.

### Available Routing Strategies

The `switchyard-libsy` crate provides several algorithmic options for traffic distribution:

- **random**: Distributes requests uniformly across available backends.
- **llm_classifier**: Routes based on content classification or model capabilities.
- **stage_router**: Implements multi-stage routing logic for complex deployment scenarios.

## API Compatibility and Protocol Translation

The server maintains compatibility with three major API specifications: OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages.

The `switchyard-protocol` crate defines the internal structures that normalize these disparate formats, while `switchyard-translation` handles bidirectional conversion. This architecture allows clients to use native SDKs without modification while Switchyard manages backend heterogeneity.

## Running the HTTP Proxy

### Installation and Startup

Install the binary using Cargo and launch the server with your routing configuration:

```bash
cargo install --locked switchyard-server
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

Validate configuration without serving traffic:

```bash
switchyard-server --config routes.toml --dry-run

```

### Health and Metrics Endpoints

Verify server status via the health endpoint:

```bash
curl http://localhost:4000/health

# {"status":"ok"}

```

Access Prometheus metrics for observability:

```bash
curl http://localhost:4000/metrics

```

### Sending LLM Requests

Test the proxy with standard OpenAI client calls:

```bash
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4",
        "messages": [{"role":"user","content":"Hello, world!"}]
      }'

```

## Summary

- The `switchyard-server` crate acts as an HTTP proxy accepting OpenAI and Anthropic API requests on configurable ports.
- It deserializes traffic into provider-neutral types defined in `switchyard-protocol`, then applies routing algorithms from `switchyard-libsy`.
- Request translation occurs via `switchyard-translation` to support diverse backends including vLLM and NVIDIA NIM.
- The server emits Prometheus metrics at `/metrics` and exposes health checks at `/health`.
- Configuration is driven by [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) and validated through CLI flags like `--dry-run`.

## Frequently Asked Questions

### What HTTP frameworks does switchyard-server use?

The crate utilizes the **Axum** framework for its HTTP server implementation, providing async request handling and middleware support as defined in [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs).

### How does switchyard-server handle different LLM API formats?

The server normalizes incoming requests into provider-neutral structures using `switchyard-protocol`, then translates them to backend-specific formats via `switchyard-translation`, enabling compatibility with OpenAI, Anthropic, and proprietary backends without client-side changes.

### What routing options are available in routes.toml?

The configuration supports algorithmic routing including `random` distribution, `llm_classifier` for intelligent content-based routing, and `stage_router` for complex multi-stage deployment scenarios, all implemented in the `switchyard-libsy` crate.

### How can I monitor switchyard-server performance?

The server exposes Prometheus-compatible metrics at the `/metrics` endpoint, tracking total requests, latency percentiles, routing overhead, and error rates, accessible via standard monitoring tools or curl commands.