Main Crates in the NVIDIA Switchyard Project: Complete Architecture Guide

The NVIDIA Switchyard project comprises eight independent Rust crates that form a layered LLM traffic orchestration stack, ranging from the core switchyard-libsy algorithm engine to HTTP servers (switchyard-server), protocol translation layers (switchyard-translation), and Python bindings (switchyard-py).

The NVIDIA-NeMo/Switchyard repository implements a provider-neutral routing system for Large Language Model traffic through a modular crate architecture. Understanding the main crates in the Switchyard project and their distinct responsibilities is essential for developers embedding the routing core, extending algorithms, or deploying the standalone server. Each crate owns a specific layer of the stack—from pure Rust orchestration logic to HTTP-compatible endpoints and Python facades.

Core Orchestration Crates

switchyard-libsy: The Routing Engine

The switchyard-libsy crate contains the provider-neutral orchestration library that implements the core routing logic. It defines the Algorithm trait, which determines which model targets to invoke, in what order, and how to combine their results. This crate contains pure Rust logic with no network I/O, making it suitable for embedding in custom applications.

Key source files include crates/libsy/src/lib.rs, src/algorithms.rs, and src/core/algorithm.rs.

switchyard-protocol: Shared Data Contracts

The switchyard-protocol crate provides the shared data model for requests, responses, usage statistics, and streaming envelopes. It defines protocol-neutral contracts such as Request, Response, and LlmResponse used by all downstream crates. Located in crates/protocol/src/lib.rs, this crate ensures type safety across the asynchronous boundaries between the routing engine and client implementations.

switchyard-llm-client: Terminal HTTP Execution

The switchyard-llm-client crate serves as the ready-made consumer for switchyard-libsy streams. It drives the algorithm's step stream, performs the terminal HTTP call to upstream LLM providers, and handles retries, fallbacks, and usage aggregation. Found in crates/libsy-llm-client/src/lib.rs, this crate bridges the decision-making logic with actual network operations.

Interface and Translation Crates

switchyard-translation: Provider Format Bridging

The switchyard-translation crate acts as a codec layer translating between three supported provider formats—openai_chat, openai_responses, and anthropic_messages—and the internal protocol types. Located in crates/switchyard-translation/src/lib.rs, it enables the server and Python bindings to accept any of these three API formats while internally using the unified protocol.

switchyard-server: HTTP API Exposure

The switchyard-server crate provides a stand-alone HTTP server exposing routing algorithms as OpenAI-compatible endpoints (/v1/chat/completions, /v1/messages, etc.). It parses a TOML configuration declaring LLM clients, targets, and algorithms, then routes each request through switchyard-libsy. The entry point resides in crates/switchyard-server/src/main.rs, with configuration handling in src/config.rs.

switchyard-py: Python FFI Bindings

The switchyard-py crate uses PyO3 bindings to expose the Rust core to Python applications. It provides the switchyard package (switchyard.libsy, switchyard.server, etc.), allowing Python applications to embed the same routing logic without launching the native server. The bindings are implemented in crates/switchyard-py/src/lib.rs and src/server_bindings.rs.

Specialized and Testing Crates

switchyard-skill-distillation: Workflow-Specific Algorithms

The optional switchyard-skill-distillation crate implements a skill distillation algorithm usable as a routing target. It defines specific ports, model identifiers, and error handling logic for distillation workflows. Source files include crates/switchyard-skill-distillation/src/lib.rs and src/model.rs.

switchyard-soak: Load Testing Utilities

The switchyard-soak crate provides a mock server and integration-test harness for load-testing (soaking) the routing stack. While not part of the production runtime, it offers essential performance validation tools via examples/mock_server.rs and supporting test utilities.

Integration Patterns and Code Examples

Understanding how these crates interact requires examining practical usage patterns. The following examples demonstrate the three primary deployment modes.

Running the Standalone Server

To deploy the full HTTP interface using switchyard-server:

export OPENAI_API_KEY="sk-..."
cargo install --locked switchyard-server
switchyard-server --config routes.toml

The server reads the TOML configuration, instantiates switchyard-libsy algorithms, translates incoming OpenAI-style payloads via switchyard-translation, and delegates terminal calls to switchyard-llm-client.

Embedding in Rust Applications

For custom binaries using the core crates directly:

use switchyard_libsy::{Algorithm, run};
use switchyard_protocol::Request;
use switchyard_llm_client::client::HttpClient;

let alg = switchyard_libsy::algorithms::Random::new(vec!["model_a", "model_b"]);
let req: Request = /* build a protocol request */;
let http = HttpClient::new();               // from switchyard-llm-client
let outcome = run(alg, req, http).await?;
println!("Selected model: {}", outcome.selected_model);

This pattern imports the Algorithm trait from switchyard-libsy, constructs protocol requests via switchyard-protocol, and executes them through the HttpClient in switchyard-llm-client.

Calling from Python

To use the routing logic without the HTTP server:

from switchyard import server, libsy

# Load a TOML configuration file

svc = server.Server(config_path="routes.toml")

# Make an OpenAI‑style request

response = svc.chat_completion(
    model="switchyard/general",
    messages=[{"role": "user", "content": "Hello"}]
)
print(response["choices"][0]["message"]["content"])

The switchyard-py bindings expose the same TOML-based configuration and routing capabilities available in the native server.

Summary

  • switchyard-libsy provides the core routing algorithms and decision logic without network dependencies.
  • switchyard-protocol defines the shared data contracts ensuring type safety across crate boundaries.
  • switchyard-llm-client executes terminal HTTP calls to upstream providers with retry and fallback logic.
  • switchyard-translation converts between OpenAI and Anthropic API formats and internal types.
  • switchyard-server exposes the routing stack as OpenAI-compatible HTTP endpoints.
  • switchyard-py offers Python bindings for embedding the engine in Python applications.
  • switchyard-skill-distillation implements specialized algorithms for distillation workflows.
  • switchyard-soak provides load-testing utilities for performance validation.

Frequently Asked Questions

What is the relationship between switchyard-libsy and switchyard-server?

The switchyard-libsy crate contains the pure Rust algorithm implementation and routing logic, while switchyard-server provides the HTTP interface and configuration management. The server crate depends on switchyard-libsy to execute routing decisions but adds network I/O, TOML parsing, and OpenAI-compatible endpoints. You can use switchyard-libsy independently in embedded applications without the server overhead.

Can I use Switchyard from Python without running the HTTP server?

Yes. The switchyard-py crate provides native Python bindings via PyO3 that expose the underlying Rust functionality directly. You can instantiate a Server object from Python, load TOML configurations, and call chat_completion() methods without spawning a separate HTTP process. This approach reduces latency for Python applications that want to embed the routing logic internally.

How does switchyard-translation handle different LLM provider formats?

The switchyard-translation crate implements a codec pattern that maps between three external formats—openai_chat, openai_responses, and anthropic_messages—and the internal switchyard-protocol types. When the server receives a request, it uses the translation layer to normalize the payload into protocol-native structs before passing them to the switchyard-libsy algorithm. Responses undergo the reverse transformation before returning to the client.

Where are the routing algorithms defined in the codebase?

Routing algorithms are defined in the switchyard-libsy crate, specifically within crates/libsy/src/algorithms.rs and src/core/algorithm.rs. The Algorithm trait in these files specifies how to select model targets, sequence calls, and combine results. Concrete implementations like the random selection algorithm reside alongside the trait definition, with additional specialized algorithms available in the optional switchyard-skill-distillation crate.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →