Understanding the Switchyard Architecture: NVIDIA's LLM Traffic Proxy Explained

Switchyard is an LLM traffic proxy that normalizes OpenAI and Anthropic API requests into a provider-agnostic protocol, routes them through configurable algorithms, and translates responses back to the original client format.

The Switchyard architecture, implemented in the NVIDIA-NeMo/Switchyard repository, provides a high-performance routing layer between client applications and LLM backends. Written primarily in Rust with Python bindings, this architecture enables stable client APIs while supporting dynamic backend selection, format translation, and fallback handling. Understanding the Switchyard architecture requires examining its four logical layers and the request lifecycle that flows through them.

Four Logical Layers of the Switchyard Architecture

The Switchyard architecture consists of four distinct logical layers that map directly to the source tree structure.

1. Client-Facing Proxy Layer

The client-facing proxy accepts requests in OpenAI or Anthropic API formats and exposes a local HTTP endpoint via the switchyard-server binary. This layer mimics public LLM APIs to ensure compatibility with existing SDKs and client applications. According to the documentation in docs/architecture.md, this proxy maintains a stable interface while the backend infrastructure changes underneath.

2. Normalization and Protocol Layer

The normalization layer transforms incoming requests into provider-independent protocol objects defined in the switchyard-protocol crate. This decouples routing logic from vendor-specific payloads. The protocol definitions live in the crates/protocol directory and provide neutral request/response models used throughout the routing pipeline.

3. Routing and Algorithms Layer

The routing layer applies policies such as weights, classifiers, and staged routing to select appropriate backend endpoints. Algorithms are implemented in Rust within the libsy crate and exposed to Python via thin factories in switchyard/libsy/algorithms.py. Key algorithms include llm_classifier and stage_router, which determine backend selection based on request content or predefined rules.

4. Backend Execution and Translation Layer

The execution layer invokes the chosen backend using the configured wire format—such as openai_chat or anthropic_messages—and translates responses back to the original client format. This functionality resides in the switchyard-translation crate, which handles codecs for various provider formats.

Request Lifecycle in Switchyard

The flow of a single request through the Switchyard architecture follows five distinct stages, as visualized in the Mermaid diagrams within docs/architecture.md:

  1. Receive – The proxy accepts a client request in OpenAI or Anthropic-compatible JSON payload format.
  2. Normalize – The request transforms into a provider-agnostic protocol object defined in the switchyard-protocol crate.
  3. Route – Routing logic (e.g., llm_classifier or stage_router) selects the appropriate backend endpoint based on configured policies.
  4. Execute – The system invokes the selected backend using the proper wire format, applying any configured fallback mechanisms if the primary backend fails.
  5. Respond – The raw backend response translates back to the client’s original expected format and streams to the client.

Core Components and Source Code Structure

The Switchyard architecture relies on several discrete components that work together to provide a unified Python API while leveraging high-performance Rust internals:

  • switchyard-server (Rust) – The native HTTP server implementing proxy endpoints, located in crates/switchyard-server.
  • switchyard-translation (Rust) – Codecs that convert between internal protocol objects and provider wire formats, found in crates/switchyard-translation.
  • switchyard-libsy (Rust) – Contains the core routing algorithms; Python factory functions reside in switchyard/libsy/algorithms.py.
  • switchyard-rust (Python façade) – PyO3 bindings that expose Rust implementations to the Python runtime.
  • switchyard-protocol – Provider-neutral request/response models used across the routing layer.

Practical Implementation Examples

Below are practical code snippets demonstrating common interactions with the Switchyard architecture.

Running the Local Switchyard Server

Launch the native server using a YAML configuration file that defines backends and routing rules:


# Launch the server with a deployment configuration

switchyard-server --config examples/prometheus/switchyard.rules.yaml --port 4000

The configuration file specifies client types (e.g., openai_chat), associated backend URLs, routing weights, and classifier settings. Example configurations are available in examples/prometheus/switchyard.rules.yaml.

Sending Chat Requests from Python

Interact with the proxy using the OpenAI-compatible client:

import switchyard as sy

# Initialize client targeting the local Switchyard proxy

client = sy.OpenAIChatClient(base_url="http://localhost:4000/v1")

response = client.chat(
    model="gpt-4",
    messages=[{"role": "user", "content": "Explain the Switchyard architecture"}],
    temperature=0.2,
)

print(response.choices[0].message["content"])

This client automatically utilizes the normalization, routing, and execution pipeline defined by the server configuration.

Configuring Custom Routing Algorithms

Implement dynamic routing using the libsy algorithm factories:

import switchyard.libsy as libsy
import requests

# Create a classifier router that routes by language

router = libsy.llm_classifier(
    classifier="language",
    models={"en": "gpt-4", "fr": "gemini-pro"},
)

# Register with the server via the admin REST API

requests.post(
    "http://localhost:4000/admin/router",
    json=router.to_json(),
)

The llm_classifier function in switchyard/libsy/algorithms.py provides a thin Python wrapper around the Rust implementation found in switchyard_rust/libsy/llm_classifier.

Summary

  • The Switchyard architecture comprises four logical layers: Client-Facing Proxy, Normalization & Protocol, Routing & Algorithms, and Backend Execution.
  • Request processing follows a five-stage lifecycle: Receive, Normalize, Route, Execute, and Respond.
  • Core routing algorithms like llm_classifier and stage_router are implemented in Rust within the libsy crate and exposed via switchyard/libsy/algorithms.py.
  • The system uses provider-agnostic protocol objects defined in the switchyard-protocol crate to decouple client APIs from backend implementations.
  • Deployment requires configuring the switchyard-server binary with YAML rule files that define backend endpoints and routing policies.

Frequently Asked Questions

What is the primary purpose of Switchyard?

Switchyard serves as an LLM traffic proxy that maintains stable client-facing APIs while providing flexible backend routing, format translation, and fallback handling. It allows organizations to switch between different LLM providers or deploy multiple backends without modifying client application code.

How does Switchyard normalize different LLM API formats?

The system uses the switchyard-protocol crate to transform incoming OpenAI or Anthropic API requests into provider-neutral protocol objects. This normalization occurs in the second architectural layer, enabling the routing engine to process requests generically without vendor-specific logic. The switchyard-translation crate then converts these internal representations back to the appropriate wire format when calling specific backends.

Where are routing algorithms implemented in the Switchyard codebase?

Routing algorithms are implemented in Rust within the libsy crate and exposed to Python through factory functions in switchyard/libsy/algorithms.py. These include algorithms such as llm_classifier for content-based routing and stage_router for staged fallbacks. The Python wrappers provide a convenient interface while the performance-critical logic executes in compiled Rust code.

How do I configure backends in Switchyard?

Backend configuration occurs through YAML deployment files passed to the switchyard-server binary via the --config flag. These files define client types (such as openai_chat), backend URLs, routing weights, and algorithm-specific parameters. The repository provides example configurations in the examples/ directory, including examples/prometheus/switchyard.rules.yaml, which demonstrates production-ready routing rules and monitoring setup.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →