# What Are the Core Components of NVIDIA Switchyard? A Complete Architecture Breakdown

> Explore the core components of NVIDIA Switchyard. Understand the architecture of this Rust and Python LLM routing layer including server, libsy, protocol, translation, and llm client.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: architecture
- Published: 2026-08-23

---

**NVIDIA Switchyard is built on six tightly-integrated Rust crates and a thin Python façade that together form a provider-neutral LLM routing layer: `switchyard-server`, `switchyard-libsy`, `switchyard-protocol`, `switchyard-translation`, `switchyard-llm-client`, and the `switchyard_rust` Python package.**

NVIDIA Switchyard, available in the [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard) repository, is an open-source routing framework designed to direct LLM traffic across multiple backends without vendor lock-in. Its architecture cleanly separates routing logic, protocol contracts, format conversion, and network I/O into independently swappable Rust crates. This article breaks down each core component of NVIDIA Switchyard, explains how they interact, and provides runnable code examples from the source.

## The Six Core Components of NVIDIA Switchyard

Each crate in the Switchyard workspace has a single responsibility, which makes the system modular, testable, and easy to extend. The table below summarizes the components before we dive into each one.

| Component | Role | Key Source |
|-----------|------|------------|
| **`switchyard-server`** | Stand-alone HTTP proxy translating OpenAI/Anthropic requests to the neutral IR, routing, and forwarding. | [[`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md) |
| **`switchyard-libsy`** | Embeddable routing core (`Algorithm` trait) with no network calls. | [[`crates/libsy/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md) |
| **`switchyard-protocol`** | Provider-neutral type definitions for requests, responses, streaming, and metadata. | [[`crates/protocol/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/README.md)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/README.md) |
| **`switchyard-translation`** | Codec converting between the neutral IR and OpenAI/Anthropic wire formats. | [[`crates/switchyard-translation/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/README.md)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/README.md) |
| **`switchyard-llm-client`** | Optional HTTP client that consumes the neutral IR, selects a backend, and performs the call. | [[`crates/libsy-llm-client/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/README.md)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/README.md) |
| **Python façade (`switchyard_rust`)** | Exposes Rust crates to Python for server or embedded algorithm usage. | [[`switchyard_rust/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/__init__.py)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/__init__.py) |

### `switchyard-server`: The Stand-Alone HTTP Proxy

The `switchyard-server` crate is the deployment-ready binary. It accepts OpenAI-compatible or Anthropic-compatible HTTP requests, translates them into the neutral intermediate representation (IR), runs the configured routing algorithm, forwards the call to an upstream model, and translates the response back to the original provider format.

Key capabilities include TOML-based route configuration, dry-run validation, and metric collection. You can install and run it directly from the crate as follows:

```bash

# Install the binary (requires Rust)

cargo install --locked switchyard-server

# Validate a TOML deployment (dry-run)

switchyard-server --config routes.toml --dry-run

# Start the proxy

switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

The server is the recommended entry point for production deployments that want a standalone network process.

### `switchyard-libsy`: The Embeddable Routing Core

**`switchyard-libsy`** is the intelligence center of Switchyard. It implements the `Algorithm` trait and performs no network operations — the host application is responsible for making actual model calls. All routing algorithms, including `random`, `llm_classifier`, and `stage_router`, live here.

Because `libsy` is a pure library, you can embed it inside any Rust or Python process without starting a separate server. This makes it the best choice for integrations that need programmatic routing without HTTP overhead.

### `switchyard-protocol`: The Shared Neutral IR

The `switchyard-protocol` crate defines the **provider-neutral type layer** — requests, responses, streaming events, metadata, and routing I/O. Every algorithm in `libsy` and every translation in `switchyard-translation` operates on these types rather than on vendor-specific formats. This is what makes Switchyard extensible: adding a new provider only requires a new translation codec, not changes to the routing core.

### `switchyard-translation`: The Wire-Format Codec

**`switchyard-translation`** is the codec component that converts between the neutral IR and the concrete wire formats of **OpenAI Chat**, **OpenAI Responses**, and **Anthropic Messages**. It contains the conversion logic needed to receive a request in one format, convert it to the IR, run the algorithm, and then re-encode the response into the original format for the client. The code belongs under `crates/switchyard-translation/` and is the only component that needs to change when a new provider format is supported.

### `switchyard-llm-client`: The Network I/O Layer

The **`switchyard-llm-client`** crate is an optional HTTP client that consumes the neutral IR, selects a backend from its configuration, encodes the request, performs the HTTP call, and decodes the reply. It pairs with `switchyard-libsy` to drive an algorithm end-to-end. With this crate, you can run routing and model calls from a single Rust process without deploying the standalone server.

### Python Façade (`switchyard_rust`)

For Python users, the `switchyard_rust` package exposes the Rust crates through the [`switchyard_rust/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/__init__.py) fileio. It enables two workflows:

- Running the server (`switchyard-server`) from Python.
- Embedding `switchyard-libsy` algorithms via the `switchyard.libsy` module.

This is the primary surface for Python-based agent orchestration frameworks and integrations.

## How the Core Components Fit Together

Understanding how the components form a request pipeline reveals the architecture's beauty. The **NVIDIA Switchyard request lifecycle** follows five steps:

1. **Client request** — A user or agent sends an OpenAI/Anthropic-compatible HTTP request to the `switchyard-server` (or directly to a Python-level façade).
2. **Translation** — `switchyard-translation` converts the core payload into the neutral `switchyard-protocol::Request`.
3. **Routing** — `switchyard-libsy` runs the selected `Algorithm` (e.g., `random`, `llm_classifier`, `stage_router`) on the neutral request, producing a stream of `Step`s indicating which target(s) to call.
4. **Model call** — If a `CallModel` step is emitted, either the built-in `switchyard-llm-client` or a custom host implementation forwards the request to the upstream backend using that backend's wire format.
5. **Response translation** — The backend response is decoded back into the neutral IR, then `switchyard-translation` re-encodes it into the original provider format for the client.

This separation lets you swap the routing algorithm, add a new LLM provider, or replace the HTTP client without touching the other layers. All routing logic lives in `libsy`, all protocol contracts live in `protocol`, all conversions between formats live in `translation`, and all network I/O lives in `llm-client`.

### Routing Requests in Python with libsy

Here is a minimal Python example from [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py) that embeds a routing algorithm directly:

```python
#!/usr/bin/env python3
import asyncio
from switchyard.libsy import algorithms, Step, LlmResponse

class EchoClient:
    async def call(self, request, model):
        # Return a fixed aggregation response (no streaming)

        return LlmResponse.Agg({
            "model": model,
            "outputs": [{"role": "assistant", "content": [{"type": "text", "text": "Hello"}]}],
        })

async def main() -> None:
    request = {
        "model": "auto",
        "stream": True,
        "messages": [{"role": "user", "content": [{"type": "text", "text": "Hello"}]}],
    }
    client = EchoClient()
    # Random routing among two targets

    algo = algorithms.random(["fast", "quality"], weights=[1, 3], seed=42)

    async for step in algo.run_stream(request):
        match step:
            case Step.CallModel(call):
                call.respond(await client.call(call.request, call.models[0]))
            case Step.Done(outcome):
                print("Chosen model:", outcome.selected_model_id)

if __name__ == "__main__":
    asyncio.run(main())

```

### Driving an Algorithm from Rust with the `run` Helper

The Rust API offers a `run` convenience function that combines an algorithm with an HTTP client router:

```rust
use std::sync::Arc;
use switchyard_libsy::Algorithm;
use switchyard_llm_client::{ClientRouter, TranslatingLlmClient};
use switchyard_protocol::Request;

// Build a translating client (OpenAI Chat backend)
let client = TranslatingLlmClient::new(&[ModelConfig::new(
    "gpt-4o-mini",
    Backend::OpenAiChat(HttpBackendConfig {
        base_url: "https://api.openai.com/v1".into(),
        api_key: std::env::var("OPENAI_API_KEY").ok(),
        ..Default::default()
    }),
    None,
)])?;

// Choose any libsy algorithm (e.g., random)
let algorithm: Arc<dyn Algorithm> = Arc::new(algorithms::random(
    vec!["fast".into(), "quality".into()], vec![1, 3], None,
));

// Route a request
let router = ClientRouter::single(Arc::new(client));
let request = Request::default(); // fill with desired LLMRequest …
let (selected_model, _response) = switchyard_llm_client::run(algorithm, router, request, None).await?;
println!("Model selected by algorithm: {}", selected_model);

```

## Key Files and Documentation Guide

To deepen your familiarity with the NVIDIA Switchyard codebase, focus on the following paths:

| Path | Description |
|------|-------------|
| [`README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/README.md) | High-level overview, feature list, quick-start instructions. |
| [`crates/switchyard-server/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md) | Server binary usage, TOML schema, metric collection. |
| [`crates/libsy/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md) | Core routing algorithms and their interaction with the neutral protocol. |
| [`crates/protocol/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/README.md) | Provider-neutral type definitions. |
| [`crates/switchyard-translation/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/README.md) | Codec for OpenAI Chat, OpenAI Responses, Anthropic Messages. |
| [`crates/libsy-llm-client/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/README.md) | HTTP client consuming the neutral IR. |
| [`switchyard_rust/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/__init__.py) | Python entry point exposing the Rust crates. |
| [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py) | Minimal Python example driving a libsy algorithm. |
| [`docs/core_concepts.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/core_concepts.md) | Runtime surfaces, request flow, and routing algorithms. |
| [`docs/architecture.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/architecture.md) | Mermaid diagram of the end-to-end lifecycle. |

## Summary

- **NVIDIA Switchyard core components** are `switchyard-server`, `switchyard-libsy`, `switchyard-protocol`, `switchyard-translation`, `switchyard-llm-client`, and the `switchyard_rust` Python package.
- The design separates **routing logic** (`libsy`), **protocol contracts** (`protocol`), **format conversion** (`translation`), and **network I/O** (`llm-client`).
- You can deploy Switchyard as a standalone proxy or embed it directly in a Python or Rust process.
- Translation layers support OpenAI Chat, OpenAI Responses, and Anthropic Messages out of the box.
- The neutral IR implementation ensures adding a new LLM provider doesn't require changes to the routing core.

## Frequently Asked Questions

### What is the difference between `switchyard-server` and `switchyard-libsy`?

`switchyard-server` is a network-standing HTTP proxy that accepts requests, translates them. It uses `switchyard-libsy` for routing logic but also manages HTTP serving, metrics, and deployments. `switchyard-libsy` is an embeddable Rust library that implements the `Algorithm` trait and performs no network I/O — you can compile it into any process and drive it directly from code.

### Which routing algorithms are included in Switchyard?

The `switchyard-libsy` crate ships `random`, `llm_classifier`, and `stage_router` algorithms. The `random` algorithm routes based on weights, `llm_classifier` uses a separate LLM to classify requests, and `stage_router` supports multi-stage routing pipelines. You can implement the `Algorithm` trait to add custom strategies.

### How does Switchyard handle different LLM providers' wire formats?

All providers are handled by the **translation** layer (`switchyard-translation`). This crate converts between the neutral IR and the concrete wire formats of OpenAI Chat, OpenAI Responses, and Anthropic Messages. Protocols like Anthropic use a single codec, while OpenAI's two product APIs map to two codecs.

### Can I use Switchyard without the HTTP server?

Yes. You can embed `switchyard-libsy` directly in a Python or Rust process and implement your own model-call callbacks. The Python package `switchyard_rust` exposes the libsy module, and the Rust `run` helper pairs an algorithm with an HTTP client to drive routing end-to-end without the server binary.