Switchyard Components: The Complete Architecture of NVIDIA’s LLM Router
Switchyard is a modular LLM router written in Rust with Python bindings, built around seven core crates—libsy, llm-client, runner, server, Python bindings, protocol, and translation—that together enable intelligent routing across multiple LLM providers while maintaining provider-agnostic data structures.
Developed under the NVIDIA-NeMo organization, Switchyard separates routing logic from transport concerns through a strict crate-based architecture. Each component in the crates/ directory handles a single responsibility, from core algorithmic decision-making to HTTP proxying and language interoperability.
Core Routing Engine: switchyard-libsy
The switchyard-libsy crate forms the algorithmic heart of the system. Located at crates/libsy/src/lib.rs, this crate exposes the Algorithm trait and its primary entry point Algorithm::run_stream, which drives the routing decision process.
According to the Switchyard source code, the core algorithm loop is implemented in crates/libsy/src/core/algorithm.rs. This crate operates entirely on normalized data structures, remaining agnostic to specific LLM providers or transport protocols.
Provider Communication: switchyard-llm-client
The switchyard-llm-client crate handles HTTP communication with upstream LLM providers. Its main entry point is the run function defined in crates/libsy-llm-client/src/run.rs.
This component feeds normalized responses back into the routing algorithms. It abstracts away provider-specific networking, retries, and credential management, allowing the core libsy algorithms to focus purely on routing logic.
Configuration and Orchestration: switchyard-runner
The switchyard-runner crate serves as the orchestration layer, parsing TOML routing configurations and wiring the client to the algorithm. The primary implementation resides in crates/switchyard-runner/src/runner.rs.
The runner drives the routing process until a model is selected, managing the lifecycle of algorithm execution and coordinating between the configuration file and runtime components.
HTTP Interface: switchyard-server
The switchyard-server crate provides a thin HTTP proxy that exposes OpenAI and Anthropic-compatible endpoints. The server implementation is located in crates/switchyard-server/src/lib.rs.
This component delegates incoming requests to the runner, effectively wrapping the entire routing pipeline behind a standard HTTP interface. Any OpenAI-compatible client can connect without modification.
Python Interoperability: switchyard-py
The switchyard-py crate provides Python bindings that expose the same routing primitives available in Rust. The Foreign Function Interface (FFI) layer is implemented in crates/switchyard-py/src/lib.rs.
These bindings allow Python applications to embed Switchyard directly, utilizing the stage_router algorithm and other components without leaving the Python runtime environment.
Shared Infrastructure
Protocol Crate
The protocol crate defines provider-neutral data structures including requests, responses, and streaming events. Defined in crates/protocol/src/lib.rs, these structures serve as the internal intermediate representation (IR) used across all Switchyard components.
Translation Layer
The switchyard-translation crate handles bidirectional conversion between vendor-specific JSON formats (OpenAI, Anthropic) and the internal protocol IR. The translation logic resides in crates/switchyard-translation/src/lib.rs.
This isolation ensures that vendor API changes affect only the translation layer, leaving core algorithms and routing logic unchanged.
Integration Ecosystem
Beyond the core seven crates, Switchyard provides optional integration plugins:
- NeMo Relay Plugin: Embeds Switchyard directly into NeMo Relay deployments (
crates/switchyard-nemo-relay-plugin/) - LiteLLM Routing Plugin: Allows LiteLLM to use Switchyard as its decision engine (
examples/litellm/)
Component Interaction Flow
The Switchyard components operate in a specific pipeline:
-
Request Normalization:
switchyard-pyor direct Rust code builds aprotocol::Requestusing shared data structures from theprotocolcrate. -
Configuration Loading: The
switchyard-runnerparses the TOML configuration and instantiates the appropriate algorithm fromlibsy. -
Algorithm Execution: The algorithm in
switchyard-libsymay invoke the LLM client (libsy-llm-client) multiple times, applying judges, stages, or escalation logic. -
Response Translation:
switchyard-translationconverts provider-specific JSON into the internal IR before returning it to the algorithm. -
HTTP Serving: When using
switchyard-server, the entire pipeline is exposed behind OpenAI-compatible endpoints.
Implementation Examples
Embedding Switchyard in Python
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router
# Build an algorithm (stage router)
algorithm = stage_router(
capable="capable",
efficient="efficient",
picker="efficient_first",
confidence_threshold=0.5,
)
# Helper to call a real model
async def call_with_fallback(request: dict, models: list[str], clients: dict) -> LlmResponse.Agg:
for model in models:
try:
return LlmResponse.Agg(await clients[model].call({**request, "model": model}))
except Exception:
continue
raise RuntimeError("All candidates failed")
# Drive the algorithm
async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
async for step in algorithm.run_stream(request):
match step:
case Step.CallModel(call):
try:
call.respond(await call_with_fallback(call.request, call.models, clients))
except Exception as e:
call.fail(e)
case Step.Done(outcome):
if outcome.response:
return outcome.response
return await call_with_fallback(
outcome.request,
outcome.selected_model_ids,
clients
)
raise RuntimeError("Algorithm ended without a decision")
Running the Standalone Proxy
# Install the server
cargo install --locked switchyard-server
# Create a TOML configuration
cat > routes.toml <<'TOML'
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"
[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"
[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
TOML
# Start the proxy
export OPENROUTER_API_KEY="your_key_here"
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
Summary
- switchyard-libsy (
crates/libsy/src/lib.rs) provides the coreAlgorithm::run_streamlogic for routing decisions. - switchyard-llm-client (
crates/libsy-llm-client/src/run.rs) manages HTTP communication with upstream providers via therunfunction. - switchyard-runner (
crates/switchyard-runner/src/runner.rs) parses TOML configurations and orchestrates the routing pipeline. - switchyard-server (
crates/switchyard-server/src/lib.rs) exposes OpenAI/Anthropic-compatible HTTP endpoints. - switchyard-py (
crates/switchyard-py/src/lib.rs) enables Python embedding through FFI bindings. - protocol (
crates/protocol/src/lib.rs) defines provider-neutral request/response structures. - switchyard-translation (
crates/switchyard-translation/src/lib.rs) normalizes vendor-specific JSON formats to internal IR.
Frequently Asked Questions
What is the primary purpose of Switchyard?
Switchyard serves as a modular LLM router that intelligently directs requests across multiple AI providers based on configurable algorithms. It separates routing logic from transport concerns, allowing teams to implement complex fallback strategies, cost optimization, and capability-based routing without modifying application code.
How does Switchyard handle different LLM provider APIs?
The switchyard-translation crate handles all provider-specific format conversions. It translates vendor-specific JSON from OpenAI, Anthropic, and other providers into a provider-neutral intermediate representation defined in the protocol crate. This ensures that core routing algorithms remain agnostic to specific API formats.
Can Switchyard be used without the HTTP server component?
Yes. While switchyard-server provides a convenient OpenAI-compatible HTTP interface, you can embed Switchyard directly in Rust or Python applications using switchyard-libsy and switchyard-py. The switchyard-runner can be invoked programmatically without the HTTP layer, making it suitable for internal service integration.
Which files contain the most critical routing logic?
The algorithmic core resides in crates/libsy/src/core/algorithm.rs, while the public API is exposed through crates/libsy/src/lib.rs. Request processing and provider communication are handled in crates/libsy-llm-client/src/run.rs, and the orchestration layer that binds these together is located in crates/switchyard-runner/src/runner.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →