What Are the Core Components of NVIDIA Switchyard? A Complete Architecture Breakdown
NVIDIA Switchyard is built on six tightly-integrated Rust crates and a thin Python façade that together form a provider-neutral LLM routing layer: switchyard-server, switchyard-libsy, switchyard-protocol, switchyard-translation, switchyard-llm-client, and the switchyard_rust Python package.
NVIDIA Switchyard, available in the NVIDIA-NeMo/Switchyard repository, is an open-source routing framework designed to direct LLM traffic across multiple backends without vendor lock-in. Its architecture cleanly separates routing logic, protocol contracts, format conversion, and network I/O into independently swappable Rust crates. This article breaks down each core component of NVIDIA Switchyard, explains how they interact, and provides runnable code examples from the source.
The Six Core Components of NVIDIA Switchyard
Each crate in the Switchyard workspace has a single responsibility, which makes the system modular, testable, and easy to extend. The table below summarizes the components before we dive into each one.
| Component | Role | Key Source |
|---|---|---|
switchyard-server |
Stand-alone HTTP proxy translating OpenAI/Anthropic requests to the neutral IR, routing, and forwarding. | [crates/switchyard-server/README.md](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/README.md) |
switchyard-libsy |
Embeddable routing core (Algorithm trait) with no network calls. |
[crates/libsy/README.md](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/README.md) |
switchyard-protocol |
Provider-neutral type definitions for requests, responses, streaming, and metadata. | [crates/protocol/README.md](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/README.md) |
switchyard-translation |
Codec converting between the neutral IR and OpenAI/Anthropic wire formats. | [crates/switchyard-translation/README.md](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/README.md) |
switchyard-llm-client |
Optional HTTP client that consumes the neutral IR, selects a backend, and performs the call. | [crates/libsy-llm-client/README.md](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/README.md) |
Python façade (switchyard_rust) |
Exposes Rust crates to Python for server or embedded algorithm usage. | [switchyard_rust/__init__.py](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/__init__.py) |
switchyard-server: The Stand-Alone HTTP Proxy
The switchyard-server crate is the deployment-ready binary. It accepts OpenAI-compatible or Anthropic-compatible HTTP requests, translates them into the neutral intermediate representation (IR), runs the configured routing algorithm, forwards the call to an upstream model, and translates the response back to the original provider format.
Key capabilities include TOML-based route configuration, dry-run validation, and metric collection. You can install and run it directly from the crate as follows:
# Install the binary (requires Rust)
cargo install --locked switchyard-server
# Validate a TOML deployment (dry-run)
switchyard-server --config routes.toml --dry-run
# Start the proxy
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
The server is the recommended entry point for production deployments that want a standalone network process.
switchyard-libsy: The Embeddable Routing Core
switchyard-libsy is the intelligence center of Switchyard. It implements the Algorithm trait and performs no network operations — the host application is responsible for making actual model calls. All routing algorithms, including random, llm_classifier, and stage_router, live here.
Because libsy is a pure library, you can embed it inside any Rust or Python process without starting a separate server. This makes it the best choice for integrations that need programmatic routing without HTTP overhead.
switchyard-protocol: The Shared Neutral IR
The switchyard-protocol crate defines the provider-neutral type layer — requests, responses, streaming events, metadata, and routing I/O. Every algorithm in libsy and every translation in switchyard-translation operates on these types rather than on vendor-specific formats. This is what makes Switchyard extensible: adding a new provider only requires a new translation codec, not changes to the routing core.
switchyard-translation: The Wire-Format Codec
switchyard-translation is the codec component that converts between the neutral IR and the concrete wire formats of OpenAI Chat, OpenAI Responses, and Anthropic Messages. It contains the conversion logic needed to receive a request in one format, convert it to the IR, run the algorithm, and then re-encode the response into the original format for the client. The code belongs under crates/switchyard-translation/ and is the only component that needs to change when a new provider format is supported.
switchyard-llm-client: The Network I/O Layer
The switchyard-llm-client crate is an optional HTTP client that consumes the neutral IR, selects a backend from its configuration, encodes the request, performs the HTTP call, and decodes the reply. It pairs with switchyard-libsy to drive an algorithm end-to-end. With this crate, you can run routing and model calls from a single Rust process without deploying the standalone server.
Python Façade (switchyard_rust)
For Python users, the switchyard_rust package exposes the Rust crates through the switchyard_rust/__init__.py fileio. It enables two workflows:
- Running the server (
switchyard-server) from Python. - Embedding
switchyard-libsyalgorithms via theswitchyard.libsymodule.
This is the primary surface for Python-based agent orchestration frameworks and integrations.
How the Core Components Fit Together
Understanding how the components form a request pipeline reveals the architecture's beauty. The NVIDIA Switchyard request lifecycle follows five steps:
- Client request — A user or agent sends an OpenAI/Anthropic-compatible HTTP request to the
switchyard-server(or directly to a Python-level façade). - Translation —
switchyard-translationconverts the core payload into the neutralswitchyard-protocol::Request. - Routing —
switchyard-libsyruns the selectedAlgorithm(e.g.,random,llm_classifier,stage_router) on the neutral request, producing a stream ofSteps indicating which target(s) to call. - Model call — If a
CallModelstep is emitted, either the built-inswitchyard-llm-clientor a custom host implementation forwards the request to the upstream backend using that backend's wire format. - Response translation — The backend response is decoded back into the neutral IR, then
switchyard-translationre-encodes it into the original provider format for the client.
This separation lets you swap the routing algorithm, add a new LLM provider, or replace the HTTP client without touching the other layers. All routing logic lives in libsy, all protocol contracts live in protocol, all conversions between formats live in translation, and all network I/O lives in llm-client.
Routing Requests in Python with libsy
Here is a minimal Python example from examples/libsy.py that embeds a routing algorithm directly:
#!/usr/bin/env python3
import asyncio
from switchyard.libsy import algorithms, Step, LlmResponse
class EchoClient:
async def call(self, request, model):
# Return a fixed aggregation response (no streaming)
return LlmResponse.Agg({
"model": model,
"outputs": [{"role": "assistant", "content": [{"type": "text", "text": "Hello"}]}],
})
async def main() -> None:
request = {
"model": "auto",
"stream": True,
"messages": [{"role": "user", "content": [{"type": "text", "text": "Hello"}]}],
}
client = EchoClient()
# Random routing among two targets
algo = algorithms.random(["fast", "quality"], weights=[1, 3], seed=42)
async for step in algo.run_stream(request):
match step:
case Step.CallModel(call):
call.respond(await client.call(call.request, call.models[0]))
case Step.Done(outcome):
print("Chosen model:", outcome.selected_model_id)
if __name__ == "__main__":
asyncio.run(main())
Driving an Algorithm from Rust with the run Helper
The Rust API offers a run convenience function that combines an algorithm with an HTTP client router:
use std::sync::Arc;
use switchyard_libsy::Algorithm;
use switchyard_llm_client::{ClientRouter, TranslatingLlmClient};
use switchyard_protocol::Request;
// Build a translating client (OpenAI Chat backend)
let client = TranslatingLlmClient::new(&[ModelConfig::new(
"gpt-4o-mini",
Backend::OpenAiChat(HttpBackendConfig {
base_url: "https://api.openai.com/v1".into(),
api_key: std::env::var("OPENAI_API_KEY").ok(),
..Default::default()
}),
None,
)])?;
// Choose any libsy algorithm (e.g., random)
let algorithm: Arc<dyn Algorithm> = Arc::new(algorithms::random(
vec!["fast".into(), "quality".into()], vec![1, 3], None,
));
// Route a request
let router = ClientRouter::single(Arc::new(client));
let request = Request::default(); // fill with desired LLMRequest …
let (selected_model, _response) = switchyard_llm_client::run(algorithm, router, request, None).await?;
println!("Model selected by algorithm: {}", selected_model);
Key Files and Documentation Guide
To deepen your familiarity with the NVIDIA Switchyard codebase, focus on the following paths:
| Path | Description |
|---|---|
README.md |
High-level overview, feature list, quick-start instructions. |
crates/switchyard-server/README.md |
Server binary usage, TOML schema, metric collection. |
crates/libsy/README.md |
Core routing algorithms and their interaction with the neutral protocol. |
crates/protocol/README.md |
Provider-neutral type definitions. |
crates/switchyard-translation/README.md |
Codec for OpenAI Chat, OpenAI Responses, Anthropic Messages. |
crates/libsy-llm-client/README.md |
HTTP client consuming the neutral IR. |
switchyard_rust/__init__.py |
Python entry point exposing the Rust crates. |
examples/libsy.py |
Minimal Python example driving a libsy algorithm. |
docs/core_concepts.md |
Runtime surfaces, request flow, and routing algorithms. |
docs/architecture.md |
Mermaid diagram of the end-to-end lifecycle. |
Summary
- NVIDIA Switchyard core components are
switchyard-server,switchyard-libsy,switchyard-protocol,switchyard-translation,switchyard-llm-client, and theswitchyard_rustPython package. - The design separates routing logic (
libsy), protocol contracts (protocol), format conversion (translation), and network I/O (llm-client). - You can deploy Switchyard as a standalone proxy or embed it directly in a Python or Rust process.
- Translation layers support OpenAI Chat, OpenAI Responses, and Anthropic Messages out of the box.
- The neutral IR implementation ensures adding a new LLM provider doesn't require changes to the routing core.
Frequently Asked Questions
What is the difference between switchyard-server and switchyard-libsy?
switchyard-server is a network-standing HTTP proxy that accepts requests, translates them. It uses switchyard-libsy for routing logic but also manages HTTP serving, metrics, and deployments. switchyard-libsy is an embeddable Rust library that implements the Algorithm trait and performs no network I/O — you can compile it into any process and drive it directly from code.
Which routing algorithms are included in Switchyard?
The switchyard-libsy crate ships random, llm_classifier, and stage_router algorithms. The random algorithm routes based on weights, llm_classifier uses a separate LLM to classify requests, and stage_router supports multi-stage routing pipelines. You can implement the Algorithm trait to add custom strategies.
How does Switchyard handle different LLM providers' wire formats?
All providers are handled by the translation layer (switchyard-translation). This crate converts between the neutral IR and the concrete wire formats of OpenAI Chat, OpenAI Responses, and Anthropic Messages. Protocols like Anthropic use a single codec, while OpenAI's two product APIs map to two codecs.
Can I use Switchyard without the HTTP server?
Yes. You can embed switchyard-libsy directly in a Python or Rust process and implement your own model-call callbacks. The Python package switchyard_rust exposes the libsy module, and the Rust run helper pairs an algorithm with an HTTP client to drive routing end-to-end without the server binary.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →