# Understanding the Switchyard Architecture: NVIDIA's LLM Traffic Proxy Explained

> Understand Switchyard, NVIDIA's LLM traffic proxy. Normalize API requests, route them with algorithms, and translate responses seamlessly for provider-agnostic LLM interactions.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: architecture
- Published: 2026-08-23

---

**Switchyard is an LLM traffic proxy that normalizes OpenAI and Anthropic API requests into a provider-agnostic protocol, routes them through configurable algorithms, and translates responses back to the original client format.**

The Switchyard architecture, implemented in the NVIDIA-NeMo/Switchyard repository, provides a high-performance routing layer between client applications and LLM backends. Written primarily in Rust with Python bindings, this architecture enables stable client APIs while supporting dynamic backend selection, format translation, and fallback handling. Understanding the Switchyard architecture requires examining its four logical layers and the request lifecycle that flows through them.

## Four Logical Layers of the Switchyard Architecture

The Switchyard architecture consists of four distinct logical layers that map directly to the source tree structure.

### 1. Client-Facing Proxy Layer

The **client-facing proxy** accepts requests in OpenAI or Anthropic API formats and exposes a local HTTP endpoint via the `switchyard-server` binary. This layer mimics public LLM APIs to ensure compatibility with existing SDKs and client applications. According to the documentation in [`docs/architecture.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/architecture.md), this proxy maintains a stable interface while the backend infrastructure changes underneath.

### 2. Normalization and Protocol Layer

The **normalization layer** transforms incoming requests into provider-independent protocol objects defined in the `switchyard-protocol` crate. This decouples routing logic from vendor-specific payloads. The protocol definitions live in the `crates/protocol` directory and provide neutral request/response models used throughout the routing pipeline.

### 3. Routing and Algorithms Layer

The **routing layer** applies policies such as weights, classifiers, and staged routing to select appropriate backend endpoints. Algorithms are implemented in Rust within the **libsy** crate and exposed to Python via thin factories in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py). Key algorithms include `llm_classifier` and `stage_router`, which determine backend selection based on request content or predefined rules.

### 4. Backend Execution and Translation Layer

The **execution layer** invokes the chosen backend using the configured wire format—such as `openai_chat` or `anthropic_messages`—and translates responses back to the original client format. This functionality resides in the `switchyard-translation` crate, which handles codecs for various provider formats.

## Request Lifecycle in Switchyard

The flow of a single request through the Switchyard architecture follows five distinct stages, as visualized in the Mermaid diagrams within [`docs/architecture.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/architecture.md):

1. **Receive** – The proxy accepts a client request in OpenAI or Anthropic-compatible JSON payload format.
2. **Normalize** – The request transforms into a provider-agnostic protocol object defined in the `switchyard-protocol` crate.
3. **Route** – Routing logic (e.g., `llm_classifier` or `stage_router`) selects the appropriate backend endpoint based on configured policies.
4. **Execute** – The system invokes the selected backend using the proper wire format, applying any configured fallback mechanisms if the primary backend fails.
5. **Respond** – The raw backend response translates back to the client’s original expected format and streams to the client.

## Core Components and Source Code Structure

The Switchyard architecture relies on several discrete components that work together to provide a unified Python API while leveraging high-performance Rust internals:

- **`switchyard-server`** (Rust) – The native HTTP server implementing proxy endpoints, located in `crates/switchyard-server`.
- **`switchyard-translation`** (Rust) – Codecs that convert between internal protocol objects and provider wire formats, found in `crates/switchyard-translation`.
- **`switchyard-libsy`** (Rust) – Contains the core routing algorithms; Python factory functions reside in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py).
- **`switchyard-rust`** (Python façade) – PyO3 bindings that expose Rust implementations to the Python runtime.
- **`switchyard-protocol`** – Provider-neutral request/response models used across the routing layer.

## Practical Implementation Examples

Below are practical code snippets demonstrating common interactions with the Switchyard architecture.

### Running the Local Switchyard Server

Launch the native server using a YAML configuration file that defines backends and routing rules:

```bash

# Launch the server with a deployment configuration

switchyard-server --config examples/prometheus/switchyard.rules.yaml --port 4000

```

The configuration file specifies client types (e.g., `openai_chat`), associated backend URLs, routing weights, and classifier settings. Example configurations are available in [`examples/prometheus/switchyard.rules.yaml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/prometheus/switchyard.rules.yaml).

### Sending Chat Requests from Python

Interact with the proxy using the OpenAI-compatible client:

```python
import switchyard as sy

# Initialize client targeting the local Switchyard proxy

client = sy.OpenAIChatClient(base_url="http://localhost:4000/v1")

response = client.chat(
    model="gpt-4",
    messages=[{"role": "user", "content": "Explain the Switchyard architecture"}],
    temperature=0.2,
)

print(response.choices[0].message["content"])

```

This client automatically utilizes the normalization, routing, and execution pipeline defined by the server configuration.

### Configuring Custom Routing Algorithms

Implement dynamic routing using the libsy algorithm factories:

```python
import switchyard.libsy as libsy
import requests

# Create a classifier router that routes by language

router = libsy.llm_classifier(
    classifier="language",
    models={"en": "gpt-4", "fr": "gemini-pro"},
)

# Register with the server via the admin REST API

requests.post(
    "http://localhost:4000/admin/router",
    json=router.to_json(),
)

```

The `llm_classifier` function in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) provides a thin Python wrapper around the Rust implementation found in `switchyard_rust/libsy/llm_classifier`.

## Summary

- The Switchyard architecture comprises four logical layers: Client-Facing Proxy, Normalization & Protocol, Routing & Algorithms, and Backend Execution.
- Request processing follows a five-stage lifecycle: Receive, Normalize, Route, Execute, and Respond.
- Core routing algorithms like `llm_classifier` and `stage_router` are implemented in Rust within the libsy crate and exposed via [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py).
- The system uses provider-agnostic protocol objects defined in the `switchyard-protocol` crate to decouple client APIs from backend implementations.
- Deployment requires configuring the `switchyard-server` binary with YAML rule files that define backend endpoints and routing policies.

## Frequently Asked Questions

### What is the primary purpose of Switchyard?

Switchyard serves as an LLM traffic proxy that maintains stable client-facing APIs while providing flexible backend routing, format translation, and fallback handling. It allows organizations to switch between different LLM providers or deploy multiple backends without modifying client application code.

### How does Switchyard normalize different LLM API formats?

The system uses the `switchyard-protocol` crate to transform incoming OpenAI or Anthropic API requests into provider-neutral protocol objects. This normalization occurs in the second architectural layer, enabling the routing engine to process requests generically without vendor-specific logic. The `switchyard-translation` crate then converts these internal representations back to the appropriate wire format when calling specific backends.

### Where are routing algorithms implemented in the Switchyard codebase?

Routing algorithms are implemented in Rust within the `libsy` crate and exposed to Python through factory functions in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py). These include algorithms such as `llm_classifier` for content-based routing and `stage_router` for staged fallbacks. The Python wrappers provide a convenient interface while the performance-critical logic executes in compiled Rust code.

### How do I configure backends in Switchyard?

Backend configuration occurs through YAML deployment files passed to the `switchyard-server` binary via the `--config` flag. These files define client types (such as `openai_chat`), backend URLs, routing weights, and algorithm-specific parameters. The repository provides example configurations in the `examples/` directory, including [`examples/prometheus/switchyard.rules.yaml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/prometheus/switchyard.rules.yaml), which demonstrates production-ready routing rules and monitoring setup.