How the switchyard-server Crate Functions as an HTTP Server for LLM Traffic

The switchyard-server crate is a standalone Rust binary that exposes an HTTP proxy server translating between OpenAI/Anthropic API formats and backend-specific protocols, routing LLM traffic through configurable algorithms while emitting Prometheus metrics.

The switchyard-server crate serves as the primary HTTP gateway for the NVIDIA-NeMo/Switchyard project, enabling seamless proxying of Large Language Model (LLM) requests between clients and diverse backends. Acting as a translator-router-proxy, it accepts standard OpenAI and Anthropic API requests and forwards them to configured model endpoints. This article examines the crate’s architecture, request lifecycle, and configuration based on the actual source implementation.

Core Architecture and Entry Points

The crate operates as an asynchronous HTTP server built on the Axum framework, initializing through two primary source files.

Server Initialization in main.rs

The entrypoint resides in crates/switchyard-server/src/main.rs. This file parses command-line flags such as --config, --host, and --port, validates the routes.toml configuration (optionally via --dry-run), and launches the Axum HTTP server. It binds to the configured address—typically 127.0.0.1:4000—and prepares the routing table for incoming LLM traffic.

Request Handling Logic in server.rs

The core request processing occurs in crates/switchyard-server/src/server.rs. This module implements the HTTP handler functions that intercept incoming requests, delegate routing decisions to the switchyard-libsy crate, and coordinate translation via switchyard-translation. It manages the full lifecycle from JSON deserialization to backend forwarding and response encoding.

The Seven-Stage LLM Request Lifecycle

The server processes each HTTP request through a deterministic pipeline that bridges client APIs with backend implementations.

  1. Expose the HTTP Endpoint: The server listens on a configurable host and port (default 4000), presenting standard endpoints like /v1/chat/completions, /v1/responses, and /messages compatible with OpenAI and Anthropic specifications.

  2. Parse Provider-Neutral Requests: Incoming JSON payloads deserialize into provider-neutral request types defined in crates/switchyard-protocol/src/lib.rs. This abstraction decouples the HTTP interface from specific backend formats.

  3. Select Backend via Routing Algorithm: Based on the routes.toml configuration, the server invokes routing algorithms—random, llm_classifier, or stage_router—implemented in crates/switchyard-libsy/src/lib.rs to select the appropriate model backend.

  4. Translate to Backend Format: The switchyard-translation crate converts the provider-neutral request into the backend-specific format required by targets such as vLLM, NVIDIA NIM, or Ollama. Translation logic resides in crates/switchyard-translation/src/lib.rs.

  5. Forward Asynchronous HTTP Calls: The server transmits the translated request to the selected backend using an async HTTP client, maintaining streaming support for chunked responses.

  6. Translate and Encode Responses: Backend responses (including streamed chunks) convert back to the provider-neutral response type, then re-encode into the original OpenAI or Anthropic JSON shape expected by the caller.

  7. Emit Prometheus Metrics: Every request updates observability counters exposed at /metrics, including switchyard_requests_total, switchyard_latency_seconds, and switchyard_routing_overhead_seconds.

Configuration and Routing Algorithms

Defining Traffic Rules with routes.toml

The server relies on a routes.toml configuration file to map incoming requests to backend pools. This file specifies routing algorithms, backend URLs, and model aliases. The --config CLI flag points to this file during startup, with optional validation available through --dry-run.

Available Routing Strategies

The switchyard-libsy crate provides several algorithmic options for traffic distribution:

  • random: Distributes requests uniformly across available backends.
  • llm_classifier: Routes based on content classification or model capabilities.
  • stage_router: Implements multi-stage routing logic for complex deployment scenarios.

API Compatibility and Protocol Translation

The server maintains compatibility with three major API specifications: OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages.

The switchyard-protocol crate defines the internal structures that normalize these disparate formats, while switchyard-translation handles bidirectional conversion. This architecture allows clients to use native SDKs without modification while Switchyard manages backend heterogeneity.

Running the HTTP Proxy

Installation and Startup

Install the binary using Cargo and launch the server with your routing configuration:

cargo install --locked switchyard-server
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

Validate configuration without serving traffic:

switchyard-server --config routes.toml --dry-run

Health and Metrics Endpoints

Verify server status via the health endpoint:

curl http://localhost:4000/health

# {"status":"ok"}

Access Prometheus metrics for observability:

curl http://localhost:4000/metrics

Sending LLM Requests

Test the proxy with standard OpenAI client calls:

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4",
        "messages": [{"role":"user","content":"Hello, world!"}]
      }'

Summary

  • The switchyard-server crate acts as an HTTP proxy accepting OpenAI and Anthropic API requests on configurable ports.
  • It deserializes traffic into provider-neutral types defined in switchyard-protocol, then applies routing algorithms from switchyard-libsy.
  • Request translation occurs via switchyard-translation to support diverse backends including vLLM and NVIDIA NIM.
  • The server emits Prometheus metrics at /metrics and exposes health checks at /health.
  • Configuration is driven by routes.toml and validated through CLI flags like --dry-run.

Frequently Asked Questions

What HTTP frameworks does switchyard-server use?

The crate utilizes the Axum framework for its HTTP server implementation, providing async request handling and middleware support as defined in crates/switchyard-server/src/main.rs.

How does switchyard-server handle different LLM API formats?

The server normalizes incoming requests into provider-neutral structures using switchyard-protocol, then translates them to backend-specific formats via switchyard-translation, enabling compatibility with OpenAI, Anthropic, and proprietary backends without client-side changes.

What routing options are available in routes.toml?

The configuration supports algorithmic routing including random distribution, llm_classifier for intelligent content-based routing, and stage_router for complex multi-stage deployment scenarios, all implemented in the switchyard-libsy crate.

How can I monitor switchyard-server performance?

The server exposes Prometheus-compatible metrics at the /metrics endpoint, tracking total requests, latency percentiles, routing overhead, and error rates, accessible via standard monitoring tools or curl commands.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →