# Switchyard Components: The Complete Architecture of NVIDIA’s LLM Router

> Explore the core components of Switchyard a Rust-based LLM router. Understand its architecture including libsy llm-client runner server Python bindings protocol and translation for efficient provider-agnostic routing.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: architecture
- Published: 2026-09-11

---

**Switchyard is a modular LLM router written in Rust with Python bindings, built around seven core crates—`libsy`, `llm-client`, `runner`, `server`, Python bindings, `protocol`, and `translation`—that together enable intelligent routing across multiple LLM providers while maintaining provider-agnostic data structures.**

Developed under the NVIDIA-NeMo organization, Switchyard separates routing logic from transport concerns through a strict crate-based architecture. Each component in the `crates/` directory handles a single responsibility, from core algorithmic decision-making to HTTP proxying and language interoperability.

## Core Routing Engine: switchyard-libsy

The **`switchyard-libsy`** crate forms the algorithmic heart of the system. Located at [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs), this crate exposes the `Algorithm` trait and its primary entry point `Algorithm::run_stream`, which drives the routing decision process.

According to the Switchyard source code, the core algorithm loop is implemented in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs). This crate operates entirely on normalized data structures, remaining agnostic to specific LLM providers or transport protocols.

## Provider Communication: switchyard-llm-client

The **`switchyard-llm-client`** crate handles HTTP communication with upstream LLM providers. Its main entry point is the `run` function defined in [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs).

This component feeds normalized responses back into the routing algorithms. It abstracts away provider-specific networking, retries, and credential management, allowing the core `libsy` algorithms to focus purely on routing logic.

## Configuration and Orchestration: switchyard-runner

The **`switchyard-runner`** crate serves as the orchestration layer, parsing TOML routing configurations and wiring the client to the algorithm. The primary implementation resides in [`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs).

The runner drives the routing process until a model is selected, managing the lifecycle of algorithm execution and coordinating between the configuration file and runtime components.

## HTTP Interface: switchyard-server

The **`switchyard-server`** crate provides a thin HTTP proxy that exposes OpenAI and Anthropic-compatible endpoints. The server implementation is located in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs).

This component delegates incoming requests to the runner, effectively wrapping the entire routing pipeline behind a standard HTTP interface. Any OpenAI-compatible client can connect without modification.

## Python Interoperability: switchyard-py

The **`switchyard-py`** crate provides Python bindings that expose the same routing primitives available in Rust. The Foreign Function Interface (FFI) layer is implemented in [`crates/switchyard-py/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/lib.rs).

These bindings allow Python applications to embed Switchyard directly, utilizing the `stage_router` algorithm and other components without leaving the Python runtime environment.

## Shared Infrastructure

### Protocol Crate

The **`protocol`** crate defines provider-neutral data structures including requests, responses, and streaming events. Defined in [`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs), these structures serve as the internal intermediate representation (IR) used across all Switchyard components.

### Translation Layer

The **`switchyard-translation`** crate handles bidirectional conversion between vendor-specific JSON formats (OpenAI, Anthropic) and the internal protocol IR. The translation logic resides in [`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs).

This isolation ensures that vendor API changes affect only the translation layer, leaving core algorithms and routing logic unchanged.

## Integration Ecosystem

Beyond the core seven crates, Switchyard provides optional integration plugins:

- **NeMo Relay Plugin**: Embeds Switchyard directly into NeMo Relay deployments (`crates/switchyard-nemo-relay-plugin/`)
- **LiteLLM Routing Plugin**: Allows LiteLLM to use Switchyard as its decision engine (`examples/litellm/`)

## Component Interaction Flow

The Switchyard components operate in a specific pipeline:

1. **Request Normalization**: `switchyard-py` or direct Rust code builds a `protocol::Request` using shared data structures from the `protocol` crate.

2. **Configuration Loading**: The `switchyard-runner` parses the TOML configuration and instantiates the appropriate algorithm from `libsy`.

3. **Algorithm Execution**: The algorithm in `switchyard-libsy` may invoke the LLM client (`libsy-llm-client`) multiple times, applying judges, stages, or escalation logic.

4. **Response Translation**: `switchyard-translation` converts provider-specific JSON into the internal IR before returning it to the algorithm.

5. **HTTP Serving**: When using `switchyard-server`, the entire pipeline is exposed behind OpenAI-compatible endpoints.

## Implementation Examples

### Embedding Switchyard in Python

```python
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# Build an algorithm (stage router)

algorithm = stage_router(
    capable="capable",
    efficient="efficient",
    picker="efficient_first",
    confidence_threshold=0.5,
)

# Helper to call a real model

async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
    for model in models:
        try:
            return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
        except Exception:
            continue
    raise RuntimeError("All candidates failed")

# Drive the algorithm

async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                try:
                    call.respond(await call_with_fallback(call.request, call.models, clients))
                except Exception as e:
                    call.fail(e)
            case Step.Done(outcome):
                if outcome.response:
                    return outcome.response
                return await call_with_fallback(
                    outcome.request, 
                    outcome.selected_model_ids, 
                    clients
                )
    raise RuntimeError("Algorithm ended without a decision")

```

### Running the Standalone Proxy

```bash

# Install the server

cargo install --locked switchyard-server

# Create a TOML configuration

cat > routes.toml <<'TOML'
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
TOML

# Start the proxy

export OPENROUTER_API_KEY="your_key_here"
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

## Summary

- **switchyard-libsy** ([`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs)) provides the core `Algorithm::run_stream` logic for routing decisions.
- **switchyard-llm-client** ([`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs)) manages HTTP communication with upstream providers via the `run` function.
- **switchyard-runner** ([`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs)) parses TOML configurations and orchestrates the routing pipeline.
- **switchyard-server** ([`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs)) exposes OpenAI/Anthropic-compatible HTTP endpoints.
- **switchyard-py** ([`crates/switchyard-py/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/lib.rs)) enables Python embedding through FFI bindings.
- **protocol** ([`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs)) defines provider-neutral request/response structures.
- **switchyard-translation** ([`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs)) normalizes vendor-specific JSON formats to internal IR.

## Frequently Asked Questions

### What is the primary purpose of Switchyard?

Switchyard serves as a modular LLM router that intelligently directs requests across multiple AI providers based on configurable algorithms. It separates routing logic from transport concerns, allowing teams to implement complex fallback strategies, cost optimization, and capability-based routing without modifying application code.

### How does Switchyard handle different LLM provider APIs?

The `switchyard-translation` crate handles all provider-specific format conversions. It translates vendor-specific JSON from OpenAI, Anthropic, and other providers into a provider-neutral intermediate representation defined in the `protocol` crate. This ensures that core routing algorithms remain agnostic to specific API formats.

### Can Switchyard be used without the HTTP server component?

Yes. While `switchyard-server` provides a convenient OpenAI-compatible HTTP interface, you can embed Switchyard directly in Rust or Python applications using `switchyard-libsy` and `switchyard-py`. The `switchyard-runner` can be invoked programmatically without the HTTP layer, making it suitable for internal service integration.

### Which files contain the most critical routing logic?

The algorithmic core resides in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs), while the public API is exposed through [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs). Request processing and provider communication are handled in [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs), and the orchestration layer that binds these together is located in [`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs).