Switchyard Development Best Practices: A Guide to Modular LLM Routing
Switchyard development best practices center on implementing pure, asynchronous routing algorithms via the Algorithm trait, using the shared protocol crate for type safety, and maintaining comprehensive unit tests with mock requests to ensure reliable LLM call routing across Rust and Python environments.
Switchyard is a Rust-based routing framework developed by NVIDIA-NeMo for orchestrating LLM calls across multiple models. Adhering to Switchyard development best practices ensures your custom routers remain composable, testable, and compatible with both the Rust core library and Python bindings.
Architecture Overview
The Switchyard codebase is organized into focused crates that separate concerns between algorithm logic, protocol definitions, and client implementations. Understanding this separation is fundamental to Switchyard development.
libsy: The core routing library containing theAlgorithmtrait and concrete implementations. Located incrates/libsy/src/lib.rs, this crate defines the central entry pointAlgorithm::run_streamwhich yields a stream ofStepitems.protocol: Defines provider-neutral types such asRequest,Step, andLlmResponseincrates/protocol/src/lib.rs. All crates depend on these shared structures to ensure type consistency across the ecosystem.libsy-llm-client: Drives algorithms by performing actual LLM calls and handling observability, implemented incrates/libsy-llm-client/src/run.rs.switchyard-py: Python bindings exposing the core library to Python environments, defined incrates/switchyard-py/src/lib.rs.
Core Development Guidelines
Implement Stateless, Asynchronous Algorithms
When creating custom routing logic, implement the Algorithm trait defined in crates/libsy/src/lib.rs. The trait's run_stream method receives a normalized Request and returns a Pin<Box<dyn Stream<Item = Step> + Send>>.
Implementations must remain pure logic without performing I/O directly. For example, the StageRouter in crates/libsy/src/algorithms/stage.rs demonstrates this pattern by yielding Step::CallModel instructions rather than executing HTTP requests itself. This separation allows the libsy-llm-client to handle all network operations while your algorithm focuses on routing decisions.
Leverage the Protocol Crate for Type Safety
All request and response structures must originate from the protocol crate. The Request type represents normalized LLM queries, while Step enum variants dictate actions like CallModel or Done. Referencing crates/protocol/src/lib.rs ensures your algorithm remains compatible with the Python bindings and server proxy.
Avoid creating parallel type definitions. Instead, re-export and use switchyard_protocol::{Request, Step, LlmResponse} as shown in crates/libsy/src/lib.rs.
Write Deterministic Unit Tests with Mock Requests
Place unit tests alongside implementations in crates/*/tests/unit/ directories. Use the mock helpers provided in the protocol crate to construct test fixtures.
The test suite in crates/prefill-router/tests/unit/algorithm.rs demonstrates the standard pattern: initialize your algorithm with test parameters, invoke run_stream with Request::mock(), and collect the resulting steps using futures::executor::block_on or #[tokio::test]. Assert that the algorithm yields expected Step variants in the correct order.
If your algorithm requires randomness, inject a seeded PRNG to maintain deterministic test outcomes.
Maintain Cross-Language Compatibility
When adding algorithms to libsy, ensure they can be exposed through switchyard-py. This typically requires adding a pyclass wrapper in crates/switchyard-py/src/lib.rs that delegates to the underlying Rust implementation.
Keep the Rust API surface stable; breaking changes require version bumps in the crate's Cargo.toml and clear documentation in the changelog. The Python bindings rely on the stability of the underlying Algorithm trait signatures.
Practical Code Examples
Creating a Custom Routing Algorithm
The following example implements a latency-aware router that selects between a capable and efficient model based on a latency threshold. This follows the pattern established in crates/libsy/src/algorithms/stage.rs.
// crates/libsy/src/algorithms/latency_aware.rs
use switchyard_protocol::{Request, Step, CallModel};
use crate::Algorithm;
use futures::stream::{self, Stream};
use std::pin::Pin;
pub struct LatencyAware {
pub capable: String,
pub efficient: String,
pub max_latency_ms: u64,
}
impl LatencyAware {
pub fn new(capable: &str, efficient: &str, max_latency_ms: u64) -> Self {
Self {
capable: capable.to_string(),
efficient: efficient.to_string(),
max_latency_ms,
}
}
}
impl Algorithm for LatencyAware {
fn run_stream(&self, request: Request) -> Pin<Box<dyn Stream<Item = Step> + Send>> {
let models = vec![self.efficient.clone(), self.capable.clone()];
Box::pin(stream::once(async move {
Step::CallModel(CallModel {
request,
models,
..Default::default()
})
}))
}
}
Python Integration Pattern
When integrating Switchyard into a Python application, use the bindings to instantiate algorithms and handle the yielded steps asynchronously.
from switchyard.libsy import AlgorithmWrapper, Step, LlmResponse
async def route_request(request: dict, algorithm, clients: dict):
"""Process a request through a Switchyard algorithm."""
async for step in algorithm.run_stream(request):
match step:
case Step.CallModel(call):
try:
response = await clients[call.models[0]].execute(call.request)
call.respond(response)
except Exception as err:
call.fail(err)
case Step.Done(outcome):
if outcome.response:
return outcome.response
raise RuntimeError("Algorithm completed without resolution")
Testing with Mock Requests
Validate your algorithm using the mock request builders from the protocol crate, following the pattern in crates/prefill-router/tests/unit/algorithm.rs.
#[tokio::test]
async fn test_latency_aware_selection() {
let algo = LatencyAware::new("capable-model", "efficient-model", 100);
let request = Request::mock();
let steps: Vec<_> = algo.run_stream(request).collect().await;
assert_eq!(steps.len(), 1);
if let Step::CallModel(call) = &steps[0] {
assert!(call.models.contains(&"efficient-model".to_string()));
} else {
panic!("Expected CallModel step");
}
}
Configuration and Observability
All runtime configuration follows the TOML schema defined in the documentation. When adding configurable parameters to your algorithm, ensure they serialize correctly according to the schema expected by the server and runner components.
For error handling, use the thiserror crate as demonstrated in crates/libsy/src/error.rs. This provides structured error types that propagate correctly through both Rust and Python interfaces.
Emit metrics via the observability module in crates/libsy/src/observability.rs. The standalone server exposes these via a Prometheus endpoint at /metrics, allowing monitoring of routing latency and algorithm decision patterns.
Summary
- Implement the
Algorithmtrait incrates/libsy/src/lib.rsto create stateless, async routing logic that yieldsStepitems rather than performing I/O. - Strictly use types from the
protocolcrate (Request,Step,LlmResponse) to maintain compatibility across Rust, Python, and server components. - Write deterministic unit tests using
Request::mock()incrates/*/tests/unit/directories, injecting seeded PRNGs when randomness is required. - Expose new algorithms through
pyclasswrappers incrates/switchyard-py/src/lib.rsto maintain Python binding parity. - Handle errors using structured types in
crates/libsy/src/error.rsand emit observability data via the dedicated module. - Pin exact commit versions when deploying pre-1.0 API code to production environments.
Frequently Asked Questions
How do I implement a new routing algorithm in Switchyard?
Create a struct in crates/libsy/src/algorithms/ that implements the Algorithm trait from crates/libsy/src/lib.rs. Your implementation must define run_stream to return a stream of Step items based on the input Request. Reference existing implementations like StageRouter in crates/libsy/src/algorithms/stage.rs for the correct structure, and ensure your algorithm remains pure by delegating actual LLM calls to the yielded Step::CallModel instructions.
What testing utilities are available for Switchyard development?
The protocol crate provides Request::mock() for creating test fixtures without constructing full HTTP payloads. The test suite in crates/prefill-router/tests/unit/algorithm.rs demonstrates the standard pattern for collecting and asserting on algorithm output streams. Use futures::executor::block_on for synchronous test contexts or #[tokio::test] for async tests.
How do I expose a Rust algorithm to Python?
Add a pyclass wrapper in crates/switchyard-py/src/lib.rs that wraps your Rust algorithm struct. The wrapper should implement methods that convert between Python types and the Rust Request/Step types defined in the protocol crate. Ensure the underlying Rust API remains stable, as breaking changes will affect Python consumers who import the switchyard module.
Where should observability metrics be emitted in a custom algorithm?
Use the observability module exported from crates/libsy/src/observability.rs rather than printing to stdout or creating custom metric collectors. This integrates with the server's Prometheus endpoint at /metrics, providing consistent monitoring of routing decisions and latency across all Switchyard deployment modes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →