# Main Crates in the NVIDIA Switchyard Project: Complete Architecture Guide

> Explore the main Rust crates in the NVIDIA Switchyard project, an LLM traffic orchestration stack. Understand the architecture from the core engine to servers and Python bindings.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: architecture
- Published: 2026-08-22

---

**The NVIDIA Switchyard project comprises eight independent Rust crates that form a layered LLM traffic orchestration stack, ranging from the core `switchyard-libsy` algorithm engine to HTTP servers (`switchyard-server`), protocol translation layers (`switchyard-translation`), and Python bindings (`switchyard-py`).**

The NVIDIA-NeMo/Switchyard repository implements a provider-neutral routing system for Large Language Model traffic through a modular crate architecture. Understanding the main crates in the Switchyard project and their distinct responsibilities is essential for developers embedding the routing core, extending algorithms, or deploying the standalone server. Each crate owns a specific layer of the stack—from pure Rust orchestration logic to HTTP-compatible endpoints and Python facades.

## Core Orchestration Crates

### switchyard-libsy: The Routing Engine

The `switchyard-libsy` crate contains the **provider-neutral orchestration library** that implements the core routing logic. It defines the [`Algorithm`](crates/libsy/src/core/algorithm.rs) trait, which determines which model targets to invoke, in what order, and how to combine their results. This crate contains pure Rust logic with no network I/O, making it suitable for embedding in custom applications.

Key source files include [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs), [`src/algorithms.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/algorithms.rs), and [`src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/core/algorithm.rs).

### switchyard-protocol: Shared Data Contracts

The `switchyard-protocol` crate provides the **shared data model** for requests, responses, usage statistics, and streaming envelopes. It defines protocol-neutral contracts such as `Request`, `Response`, and `LlmResponse` used by all downstream crates. Located in [`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs), this crate ensures type safety across the asynchronous boundaries between the routing engine and client implementations.

### switchyard-llm-client: Terminal HTTP Execution

The `switchyard-llm-client` crate serves as the ready-made consumer for `switchyard-libsy` streams. It drives the algorithm's step stream, performs the **terminal HTTP call** to upstream LLM providers, and handles retries, fallbacks, and usage aggregation. Found in [`crates/libsy-llm-client/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/lib.rs), this crate bridges the decision-making logic with actual network operations.

## Interface and Translation Crates

### switchyard-translation: Provider Format Bridging

The `switchyard-translation` crate acts as a **codec layer** translating between three supported provider formats—`openai_chat`, `openai_responses`, and `anthropic_messages`—and the internal protocol types. Located in [`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs), it enables the server and Python bindings to accept any of these three API formats while internally using the unified protocol.

### switchyard-server: HTTP API Exposure

The `switchyard-server` crate provides a **stand-alone HTTP server** exposing routing algorithms as OpenAI-compatible endpoints (`/v1/chat/completions`, `/v1/messages`, etc.). It parses a TOML configuration declaring LLM clients, targets, and algorithms, then routes each request through `switchyard-libsy`. The entry point resides in [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs), with configuration handling in [`src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/config.rs).

### switchyard-py: Python FFI Bindings

The `switchyard-py` crate uses **PyO3 bindings** to expose the Rust core to Python applications. It provides the `switchyard` package (`switchyard.libsy`, `switchyard.server`, etc.), allowing Python applications to embed the same routing logic without launching the native server. The bindings are implemented in [`crates/switchyard-py/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/lib.rs) and [`src/server_bindings.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/server_bindings.rs).

## Specialized and Testing Crates

### switchyard-skill-distillation: Workflow-Specific Algorithms

The optional `switchyard-skill-distillation` crate implements a **skill distillation algorithm** usable as a routing target. It defines specific ports, model identifiers, and error handling logic for distillation workflows. Source files include [`crates/switchyard-skill-distillation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-skill-distillation/src/lib.rs) and [`src/model.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/model.rs).

### switchyard-soak: Load Testing Utilities

The `switchyard-soak` crate provides a **mock server and integration-test harness** for load-testing (soaking) the routing stack. While not part of the production runtime, it offers essential performance validation tools via [`examples/mock_server.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/mock_server.rs) and supporting test utilities.

## Integration Patterns and Code Examples

Understanding how these crates interact requires examining practical usage patterns. The following examples demonstrate the three primary deployment modes.

### Running the Standalone Server

To deploy the full HTTP interface using `switchyard-server`:

```bash
export OPENAI_API_KEY="sk-..."
cargo install --locked switchyard-server
switchyard-server --config routes.toml

```

The server reads the TOML configuration, instantiates `switchyard-libsy` algorithms, translates incoming OpenAI-style payloads via `switchyard-translation`, and delegates terminal calls to `switchyard-llm-client`.

### Embedding in Rust Applications

For custom binaries using the core crates directly:

```rust
use switchyard_libsy::{Algorithm, run};
use switchyard_protocol::Request;
use switchyard_llm_client::client::HttpClient;

let alg = switchyard_libsy::algorithms::Random::new(vec!["model_a", "model_b"]);
let req: Request = /* build a protocol request */;
let http = HttpClient::new();               // from switchyard-llm-client
let outcome = run(alg, req, http).await?;
println!("Selected model: {}", outcome.selected_model);

```

This pattern imports the `Algorithm` trait from `switchyard-libsy`, constructs protocol requests via `switchyard-protocol`, and executes them through the `HttpClient` in `switchyard-llm-client`.

### Calling from Python

To use the routing logic without the HTTP server:

```python
from switchyard import server, libsy

# Load a TOML configuration file

svc = server.Server(config_path="routes.toml")

# Make an OpenAI‑style request

response = svc.chat_completion(
    model="switchyard/general",
    messages=[{"role": "user", "content": "Hello"}]
)
print(response["choices"][0]["message"]["content"])

```

The `switchyard-py` bindings expose the same TOML-based configuration and routing capabilities available in the native server.

## Summary

- **`switchyard-libsy`** provides the core **routing algorithms** and decision logic without network dependencies.
- **`switchyard-protocol`** defines the **shared data contracts** ensuring type safety across crate boundaries.
- **`switchyard-llm-client`** executes **terminal HTTP calls** to upstream providers with retry and fallback logic.
- **`switchyard-translation`** converts between **OpenAI and Anthropic API formats** and internal types.
- **`switchyard-server`** exposes the routing stack as **OpenAI-compatible HTTP endpoints**.
- **`switchyard-py`** offers **Python bindings** for embedding the engine in Python applications.
- **`switchyard-skill-distillation`** implements **specialized algorithms** for distillation workflows.
- **`switchyard-soak`** provides **load-testing utilities** for performance validation.

## Frequently Asked Questions

### What is the relationship between switchyard-libsy and switchyard-server?

The `switchyard-libsy` crate contains the pure Rust **algorithm implementation** and routing logic, while `switchyard-server` provides the **HTTP interface** and configuration management. The server crate depends on `switchyard-libsy` to execute routing decisions but adds network I/O, TOML parsing, and OpenAI-compatible endpoints. You can use `switchyard-libsy` independently in embedded applications without the server overhead.

### Can I use Switchyard from Python without running the HTTP server?

Yes. The `switchyard-py` crate provides **native Python bindings** via PyO3 that expose the underlying Rust functionality directly. You can instantiate a `Server` object from Python, load TOML configurations, and call `chat_completion()` methods without spawning a separate HTTP process. This approach reduces latency for Python applications that want to embed the routing logic internally.

### How does switchyard-translation handle different LLM provider formats?

The `switchyard-translation` crate implements a **codec pattern** that maps between three external formats—`openai_chat`, `openai_responses`, and `anthropic_messages`—and the internal `switchyard-protocol` types. When the server receives a request, it uses the translation layer to normalize the payload into protocol-native structs before passing them to the `switchyard-libsy` algorithm. Responses undergo the reverse transformation before returning to the client.

### Where are the routing algorithms defined in the codebase?

Routing algorithms are defined in the **`switchyard-libsy`** crate, specifically within [`crates/libsy/src/algorithms.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms.rs) and [`src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/core/algorithm.rs). The [`Algorithm`](crates/libsy/src/core/algorithm.rs) trait in these files specifies how to select model targets, sequence calls, and combine results. Concrete implementations like the random selection algorithm reside alongside the trait definition, with additional specialized algorithms available in the optional `switchyard-skill-distillation` crate.