What Is NVIDIA Switchyard? A Rust‑Based Proxy for LLM Traffic Orchestration

Switchyard is a Rust‑based proxy and library that sits between client applications and large‑language‑model backends, solving production challenges including protocol translation, intelligent multi‑backend routing, and operational observability.

Switchyard is an open‑source project maintained by NVIDIA within the NeMo ecosystem. It acts as a unified traffic‑orchestration layer that accepts requests in standard formats like OpenAI Chat and Anthropic Messages, then routes them to heterogeneous backend services while providing Prometheus‑grade metrics and extensible routing logic.

Protocol Translation for Heterogeneous LLM Ecosystems

Modern LLM applications face a protocol mismatch problem: clients often speak native OpenAI Chat, OpenAI Responses, or Anthropic Messages formats, while backends may use provider‑specific APIs. Switchyard addresses this through a dedicated translation layer located in crates/switchyard-translation.

The proxy intercepts incoming requests on standard HTTP endpoints such as /v1/chat/completions and /v1/messages, then converts the payload on‑the‑fly into the target backend’s native format. When the upstream model returns a response, Switchyard translates it back to the client’s expected format before streaming or aggregating the result.

Intelligent Multi‑Backend Routing Algorithms

Production LLM traffic requires sophisticated routing for cost optimization, A/B testing, and capability‑based escalation. Switchyard implements multiple routing algorithms in the Rust crate switchyard-libsy, exposing them to Python via switchyard/libsy/algorithms.py.

Available algorithms include:

  • Random routing – Distributes traffic across model pools with configurable weights.
  • LLM classifier – Uses a lightweight judge model to select the appropriate backend based on request complexity.
  • Stage router – Reuses intermediate signals (tool results, system prompts) to avoid redundant classifier calls and reduce latency.

Because routing decisions occur before the upstream call, Switchyard can perform tiered escalation, sending simple queries to cost‑efficient models and promoting complex queries to premium tiers only when classifiers deem it necessary.

Production Observability and Metrics

Switchyard emits Prometheus‑compatible metrics for every routed call, providing visibility into request counts, end‑to‑end latency, token usage, and routing overhead. Unlike opaque proxy solutions, Switchyard exposes granular data on classifier outcomes and algorithmic decision latency, enabling SRE teams to optimize their routing strategies based on real‑time production data.

Extensible Architecture: Library vs. Server Path

Switchyard offers two deployment modes to accommodate different integration requirements:

Library Path (switchyard-libsy) – Link the Rust crate directly into existing binaries for embedded routing logic. This path exposes a stable API through switchyard_rust.server that can be instantiated from Rust or Python.

Server Path (switchyard-server) – Deploy Switchyard as a standalone, stateless binary that operates as a drop‑in proxy. The server loads routing rules from a TOML configuration file and requires no client‑side code changes.

Both paths maintain stateless operation except for optional session‑affinity logs, simplifying horizontal scaling across Kubernetes or VM fleets.

How Switchyard Processes LLM Requests

The request lifecycle follows a deterministic pipeline implemented in crates/switchyard-server:

  1. Ingress – Client sends a request to Switchyard’s HTTP endpoint in their native format (OpenAI or Anthropic).
  2. Routing – Switchyard selects a target model using the configured algorithm (e.g., algorithms.random or algorithms.llm_classifier).
  3. Translation – The switchyard-translation crate converts the request to the provider’s native format.
  4. Optional Judging – For classifier or stage‑router algorithms, Switchyard may call an intermediate judge model before the final LLM call.
  5. Upstream – Switchyard forwards the request to the chosen backend, collects the response, translates it back, and returns it to the client.

This architecture enables provider‑agnostic operation, working seamlessly with vLLM, NVIDIA NIM, Ollama, or any OpenAI‑compatible endpoint.

Code Examples

Using the Python Library for Custom Routing

The file examples/libsy.py demonstrates how to drive routing algorithms directly from Python without running the full server. Below is a minimal implementation using the random algorithm:

import asyncio
from switchyard.libsy import LlmResponse, Step, algorithms

class EchoClient:
    """Mock LLM backend for demonstration."""
    async def call(self, request, model) -> LlmResponse.Agg | LlmResponse.Stream:
        if request.get("stream"):
            async def events():
                yield {"normalized": [{"MessageStart": {"id": "echo", "model": model}}]}
                yield {"normalized": [{"TextDelta": {"index": 0, "text": "Hello"}}]}
                yield {"normalized": [{"MessageStop": {"reason": "end_turn"}}]}
            return LlmResponse.Stream(events())
        return LlmResponse.Agg({
            "model": model,
            "outputs": [{"role": "assistant", "content": [{"type": "text", "text": "Hello"}]}]
        })

async def main():
    request = {
        "model": "auto",
        "stream": True,
        "messages": [{"role": "user", "content": [{"type": "text", "text": "Hello"}]}]
    }
    client = EchoClient()
    
    # Initialize random router with weighted distribution

    algorithm = algorithms.random(["fast", "quality"], weights=[1, 3], seed=42)
    
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                call.respond(await client.call(call.request, call.models[0]))
            case Step.Done(outcome):
                print(f"Selected model: {outcome.selected_model_id}")

if __name__ == "__main__":
    asyncio.run(main())

This example shows streaming aggregation and routing decision introspection using the Python façade over Rust implementations.

Deploying the Standalone Server

For production deployments, instantiate the Rust server via the Python wrapper in switchyard_rust/server.py:

from switchyard_rust.server import Server

# routes.toml defines backend endpoints and routing rules

server = Server("routes.toml", port=4000)
print(f"Switchyard listening at {server.base_url}")

# Server runs in background until interrupted

The Server class dynamically loads the native Rust binary and parses the TOML configuration referenced in crates/switchyard-server/README.md, exposing the proxy on the specified port.

Summary

  • Switchyard is a Rust‑based proxy and library from NVIDIA NeMo that solves LLM traffic routing challenges through protocol‑agnostic request handling.
  • The translation layer in crates/switchyard-translation converts between OpenAI, Anthropic, and provider‑native formats on the fly.
  • Routing algorithms (random, LLM‑classifier, stage‑router) defined in switchyard-libsy enable cost‑aware traffic splitting and intelligent escalation without client‑side changes.
  • Prometheus metrics provide production visibility into latency, token usage, and routing decisions.
  • Dual deployment modes (embedded library via switchyard_rust.server or standalone binary via switchyard-server) offer flexibility for both custom applications and drop‑in proxy scenarios.

Frequently Asked Questions

What makes Switchyard different from other LLM proxies?

Switchyard combines protocol translation, composable routing algorithms, and Prometheus observability in a single Rust codebase. Unlike simple reverse proxies, it includes an LLM‑classifier routing engine and stage‑router logic for reusing intermediate computation, all exposed through both Rust and Python APIs as seen in switchyard/libsy/algorithms.py.

Can I implement custom routing logic with Switchyard?

Yes. The library path (switchyard-libsy) allows you to write custom Rust algorithms and expose them through the same Python façade used by built‑in methods. The examples/libsy.py file demonstrates the Step and LlmResponse interfaces required to integrate custom decision logic.

How does Switchyard handle streaming responses?

Switchyard natively supports streaming and aggregated responses through the LlmResponse.Stream and LlmResponse.Agg types defined in the Rust core. The proxy maintains the SSE (Server‑Sent Events) stream from the backend while translating chunks to the client’s expected format, ensuring low‑latency delivery without buffering entire responses.

Is Switchyard limited to NVIDIA models or infrastructure?

No. Switchyard is provider‑agnostic and works with any OpenAI‑compatible endpoint, including vLLM, Ollama, and cloud‑hosted APIs. The routes.toml configuration file used by switchyard_rust.server.Server accepts arbitrary base URLs and authentication headers, making it compatible with heterogeneous model deployments regardless of hardware vendor.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →