# What Is NVIDIA Switchyard? A Rust‑Based Proxy for LLM Traffic Orchestration

> Explore NVIDIA Switchyard, a Rust proxy for LLM traffic. It simplifies protocol translation, intelligent routing, and observability for seamless backend integration.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: getting-started
- Published: 2026-08-22

---

**Switchyard is a Rust‑based proxy and library that sits between client applications and large‑language‑model backends, solving production challenges including protocol translation, intelligent multi‑backend routing, and operational observability.**

Switchyard is an open‑source project maintained by NVIDIA within the NeMo ecosystem. It acts as a unified traffic‑orchestration layer that accepts requests in standard formats like OpenAI Chat and Anthropic Messages, then routes them to heterogeneous backend services while providing Prometheus‑grade metrics and extensible routing logic.

## Protocol Translation for Heterogeneous LLM Ecosystems

Modern LLM applications face a **protocol mismatch problem**: clients often speak native OpenAI Chat, OpenAI Responses, or Anthropic Messages formats, while backends may use provider‑specific APIs. Switchyard addresses this through a dedicated translation layer located in `crates/switchyard-translation`. 

The proxy intercepts incoming requests on standard HTTP endpoints such as `/v1/chat/completions` and `/v1/messages`, then converts the payload on‑the‑fly into the target backend’s native format. When the upstream model returns a response, Switchyard translates it back to the client’s expected format before streaming or aggregating the result.

## Intelligent Multi‑Backend Routing Algorithms

Production LLM traffic requires **sophisticated routing** for cost optimization, A/B testing, and capability‑based escalation. Switchyard implements multiple routing algorithms in the Rust crate `switchyard-libsy`, exposing them to Python via [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py).

Available algorithms include:

- **Random routing** – Distributes traffic across model pools with configurable weights.
- **LLM classifier** – Uses a lightweight judge model to select the appropriate backend based on request complexity.
- **Stage router** – Reuses intermediate signals (tool results, system prompts) to avoid redundant classifier calls and reduce latency.

Because routing decisions occur *before* the upstream call, Switchyard can perform **tiered escalation**, sending simple queries to cost‑efficient models and promoting complex queries to premium tiers only when classifiers deem it necessary.

## Production Observability and Metrics

Switchyard emits **Prometheus‑compatible metrics** for every routed call, providing visibility into request counts, end‑to‑end latency, token usage, and routing overhead. Unlike opaque proxy solutions, Switchyard exposes granular data on classifier outcomes and algorithmic decision latency, enabling SRE teams to optimize their routing strategies based on real‑time production data.

## Extensible Architecture: Library vs. Server Path

Switchyard offers two deployment modes to accommodate different integration requirements:

**Library Path (`switchyard-libsy`)** – Link the Rust crate directly into existing binaries for embedded routing logic. This path exposes a stable API through `switchyard_rust.server` that can be instantiated from Rust or Python.

**Server Path (`switchyard-server`)** – Deploy Switchyard as a standalone, stateless binary that operates as a drop‑in proxy. The server loads routing rules from a TOML configuration file and requires no client‑side code changes.

Both paths maintain **stateless operation** except for optional session‑affinity logs, simplifying horizontal scaling across Kubernetes or VM fleets.

## How Switchyard Processes LLM Requests

The request lifecycle follows a deterministic pipeline implemented in `crates/switchyard-server`:

1. **Ingress** – Client sends a request to Switchyard’s HTTP endpoint in their native format (OpenAI or Anthropic).
2. **Routing** – Switchyard selects a target model using the configured algorithm (e.g., `algorithms.random` or `algorithms.llm_classifier`).
3. **Translation** – The `switchyard-translation` crate converts the request to the provider’s native format.
4. **Optional Judging** – For classifier or stage‑router algorithms, Switchyard may call an intermediate judge model before the final LLM call.
5. **Upstream** – Switchyard forwards the request to the chosen backend, collects the response, translates it back, and returns it to the client.

This architecture enables **provider‑agnostic** operation, working seamlessly with vLLM, NVIDIA NIM, Ollama, or any OpenAI‑compatible endpoint.

## Code Examples

### Using the Python Library for Custom Routing

The file [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py) demonstrates how to drive routing algorithms directly from Python without running the full server. Below is a minimal implementation using the random algorithm:

```python
import asyncio
from switchyard.libsy import LlmResponse, Step, algorithms

class EchoClient:
    """Mock LLM backend for demonstration."""
    async def call(self, request, model) -> LlmResponse.Agg | LlmResponse.Stream:
        if request.get("stream"):
            async def events():
                yield {"normalized": [{"MessageStart": {"id": "echo", "model": model}}]}
                yield {"normalized": [{"TextDelta": {"index": 0, "text": "Hello"}}]}
                yield {"normalized": [{"MessageStop": {"reason": "end_turn"}}]}
            return LlmResponse.Stream(events())
        return LlmResponse.Agg({
            "model": model,
            "outputs": [{"role": "assistant", "content": [{"type": "text", "text": "Hello"}]}]
        })

async def main():
    request = {
        "model": "auto",
        "stream": True,
        "messages": [{"role": "user", "content": [{"type": "text", "text": "Hello"}]}]
    }
    client = EchoClient()
    
    # Initialize random router with weighted distribution

    algorithm = algorithms.random(["fast", "quality"], weights=[1, 3], seed=42)
    
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                call.respond(await client.call(call.request, call.models[0]))
            case Step.Done(outcome):
                print(f"Selected model: {outcome.selected_model_id}")

if __name__ == "__main__":
    asyncio.run(main())

```

This example shows **streaming aggregation** and **routing decision introspection** using the Python façade over Rust implementations.

### Deploying the Standalone Server

For production deployments, instantiate the Rust server via the Python wrapper in [`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py):

```python
from switchyard_rust.server import Server

# routes.toml defines backend endpoints and routing rules

server = Server("routes.toml", port=4000)
print(f"Switchyard listening at {server.base_url}")

# Server runs in background until interrupted

```

The `Server` class dynamically loads the native Rust binary and parses the TOML configuration referenced in [`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md), exposing the proxy on the specified port.

## Summary

- **Switchyard** is a Rust‑based proxy and library from NVIDIA NeMo that solves **LLM traffic routing** challenges through protocol‑agnostic request handling.
- The **translation layer** in `crates/switchyard-translation` converts between OpenAI, Anthropic, and provider‑native formats on the fly.
- **Routing algorithms** (random, LLM‑classifier, stage‑router) defined in `switchyard-libsy` enable cost‑aware traffic splitting and intelligent escalation without client‑side changes.
- **Prometheus metrics** provide production visibility into latency, token usage, and routing decisions.
- Dual deployment modes (embedded library via `switchyard_rust.server` or standalone binary via `switchyard-server`) offer flexibility for both custom applications and drop‑in proxy scenarios.

## Frequently Asked Questions

### What makes Switchyard different from other LLM proxies?

Switchyard combines **protocol translation**, **composable routing algorithms**, and **Prometheus observability** in a single Rust codebase. Unlike simple reverse proxies, it includes an LLM‑classifier routing engine and stage‑router logic for reusing intermediate computation, all exposed through both Rust and Python APIs as seen in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py).

### Can I implement custom routing logic with Switchyard?

Yes. The library path (`switchyard-libsy`) allows you to write custom Rust algorithms and expose them through the same Python façade used by built‑in methods. The [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py) file demonstrates the `Step` and `LlmResponse` interfaces required to integrate custom decision logic.

### How does Switchyard handle streaming responses?

Switchyard natively supports **streaming and aggregated responses** through the `LlmResponse.Stream` and `LlmResponse.Agg` types defined in the Rust core. The proxy maintains the SSE (Server‑Sent Events) stream from the backend while translating chunks to the client’s expected format, ensuring low‑latency delivery without buffering entire responses.

### Is Switchyard limited to NVIDIA models or infrastructure?

No. Switchyard is **provider‑agnostic** and works with any OpenAI‑compatible endpoint, including vLLM, Ollama, and cloud‑hosted APIs. The [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration file used by `switchyard_rust.server.Server` accepts arbitrary base URLs and authentication headers, making it compatible with heterogeneous model deployments regardless of hardware vendor.