# Switchyard Development Best Practices: A Guide to Modular LLM Routing

> Discover Switchyard development best practices for modular LLM routing. Learn about async algorithms, type safety, and testing for reliable Rust and Python integrations.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: best-practices
- Published: 2026-09-11

---

**Switchyard development best practices center on implementing pure, asynchronous routing algorithms via the `Algorithm` trait, using the shared `protocol` crate for type safety, and maintaining comprehensive unit tests with mock requests to ensure reliable LLM call routing across Rust and Python environments.**

Switchyard is a Rust-based routing framework developed by NVIDIA-NeMo for orchestrating LLM calls across multiple models. Adhering to Switchyard development best practices ensures your custom routers remain composable, testable, and compatible with both the Rust core library and Python bindings.

## Architecture Overview

The Switchyard codebase is organized into focused crates that separate concerns between algorithm logic, protocol definitions, and client implementations. Understanding this separation is fundamental to Switchyard development.

- **`libsy`**: The core routing library containing the `Algorithm` trait and concrete implementations. Located in [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs), this crate defines the central entry point `Algorithm::run_stream` which yields a stream of `Step` items.
- **`protocol`**: Defines provider-neutral types such as `Request`, `Step`, and `LlmResponse` in [`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs). All crates depend on these shared structures to ensure type consistency across the ecosystem.
- **`libsy-llm-client`**: Drives algorithms by performing actual LLM calls and handling observability, implemented in [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs).
- **`switchyard-py`**: Python bindings exposing the core library to Python environments, defined in [`crates/switchyard-py/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/lib.rs).

## Core Development Guidelines

### Implement Stateless, Asynchronous Algorithms

When creating custom routing logic, implement the `Algorithm` trait defined in [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs). The trait's `run_stream` method receives a normalized `Request` and returns a `Pin<Box<dyn Stream<Item = Step> + Send>>`.

Implementations must remain pure logic without performing I/O directly. For example, the `StageRouter` in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs) demonstrates this pattern by yielding `Step::CallModel` instructions rather than executing HTTP requests itself. This separation allows the `libsy-llm-client` to handle all network operations while your algorithm focuses on routing decisions.

### Leverage the Protocol Crate for Type Safety

All request and response structures must originate from the `protocol` crate. The `Request` type represents normalized LLM queries, while `Step` enum variants dictate actions like `CallModel` or `Done`. Referencing [`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs) ensures your algorithm remains compatible with the Python bindings and server proxy.

Avoid creating parallel type definitions. Instead, re-export and use `switchyard_protocol::{Request, Step, LlmResponse}` as shown in [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs).

### Write Deterministic Unit Tests with Mock Requests

Place unit tests alongside implementations in `crates/*/tests/unit/` directories. Use the mock helpers provided in the protocol crate to construct test fixtures.

The test suite in [`crates/prefill-router/tests/unit/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/prefill-router/tests/unit/algorithm.rs) demonstrates the standard pattern: initialize your algorithm with test parameters, invoke `run_stream` with `Request::mock()`, and collect the resulting steps using `futures::executor::block_on` or `#[tokio::test]`. Assert that the algorithm yields expected `Step` variants in the correct order.

If your algorithm requires randomness, inject a seeded PRNG to maintain deterministic test outcomes.

### Maintain Cross-Language Compatibility

When adding algorithms to `libsy`, ensure they can be exposed through `switchyard-py`. This typically requires adding a `pyclass` wrapper in [`crates/switchyard-py/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/lib.rs) that delegates to the underlying Rust implementation.

Keep the Rust API surface stable; breaking changes require version bumps in the crate's [`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml) and clear documentation in the changelog. The Python bindings rely on the stability of the underlying `Algorithm` trait signatures.

## Practical Code Examples

### Creating a Custom Routing Algorithm

The following example implements a latency-aware router that selects between a capable and efficient model based on a latency threshold. This follows the pattern established in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs).

```rust
// crates/libsy/src/algorithms/latency_aware.rs
use switchyard_protocol::{Request, Step, CallModel};
use crate::Algorithm;
use futures::stream::{self, Stream};
use std::pin::Pin;

pub struct LatencyAware {
    pub capable: String,
    pub efficient: String,
    pub max_latency_ms: u64,
}

impl LatencyAware {
    pub fn new(capable: &str, efficient: &str, max_latency_ms: u64) -> Self {
        Self {
            capable: capable.to_string(),
            efficient: efficient.to_string(),
            max_latency_ms,
        }
    }
}

impl Algorithm for LatencyAware {
    fn run_stream(&self, request: Request) -> Pin<Box<dyn Stream<Item = Step> + Send>> {
        let models = vec![self.efficient.clone(), self.capable.clone()];
        
        Box::pin(stream::once(async move {
            Step::CallModel(CallModel {
                request,
                models,
                ..Default::default()
            })
        }))
    }
}

```

### Python Integration Pattern

When integrating Switchyard into a Python application, use the bindings to instantiate algorithms and handle the yielded steps asynchronously.

```python
from switchyard.libsy import AlgorithmWrapper, Step, LlmResponse

async def route_request(request: dict, algorithm, clients: dict):
    """Process a request through a Switchyard algorithm."""
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                try:
                    response = await clients[call.models[0]].execute(call.request)
                    call.respond(response)
                except Exception as err:
                    call.fail(err)
            case Step.Done(outcome):
                if outcome.response:
                    return outcome.response
    raise RuntimeError("Algorithm completed without resolution")

```

### Testing with Mock Requests

Validate your algorithm using the mock request builders from the protocol crate, following the pattern in [`crates/prefill-router/tests/unit/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/prefill-router/tests/unit/algorithm.rs).

```rust
#[tokio::test]
async fn test_latency_aware_selection() {
    let algo = LatencyAware::new("capable-model", "efficient-model", 100);
    let request = Request::mock();
    
    let steps: Vec<_> = algo.run_stream(request).collect().await;
    
    assert_eq!(steps.len(), 1);
    if let Step::CallModel(call) = &steps[0] {
        assert!(call.models.contains(&"efficient-model".to_string()));
    } else {
        panic!("Expected CallModel step");
    }
}

```

## Configuration and Observability

All runtime configuration follows the TOML schema defined in the documentation. When adding configurable parameters to your algorithm, ensure they serialize correctly according to the schema expected by the server and runner components.

For error handling, use the `thiserror` crate as demonstrated in [`crates/libsy/src/error.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/error.rs). This provides structured error types that propagate correctly through both Rust and Python interfaces.

Emit metrics via the `observability` module in [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs). The standalone server exposes these via a Prometheus endpoint at `/metrics`, allowing monitoring of routing latency and algorithm decision patterns.

## Summary

- Implement the `Algorithm` trait in [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs) to create stateless, async routing logic that yields `Step` items rather than performing I/O.
- Strictly use types from the `protocol` crate (`Request`, `Step`, `LlmResponse`) to maintain compatibility across Rust, Python, and server components.
- Write deterministic unit tests using `Request::mock()` in `crates/*/tests/unit/` directories, injecting seeded PRNGs when randomness is required.
- Expose new algorithms through `pyclass` wrappers in [`crates/switchyard-py/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/lib.rs) to maintain Python binding parity.
- Handle errors using structured types in [`crates/libsy/src/error.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/error.rs) and emit observability data via the dedicated module.
- Pin exact commit versions when deploying pre-1.0 API code to production environments.

## Frequently Asked Questions

### How do I implement a new routing algorithm in Switchyard?

Create a struct in `crates/libsy/src/algorithms/` that implements the `Algorithm` trait from [`crates/libsy/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/lib.rs). Your implementation must define `run_stream` to return a stream of `Step` items based on the input `Request`. Reference existing implementations like `StageRouter` in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs) for the correct structure, and ensure your algorithm remains pure by delegating actual LLM calls to the yielded `Step::CallModel` instructions.

### What testing utilities are available for Switchyard development?

The `protocol` crate provides `Request::mock()` for creating test fixtures without constructing full HTTP payloads. The test suite in [`crates/prefill-router/tests/unit/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/prefill-router/tests/unit/algorithm.rs) demonstrates the standard pattern for collecting and asserting on algorithm output streams. Use `futures::executor::block_on` for synchronous test contexts or `#[tokio::test]` for async tests.

### How do I expose a Rust algorithm to Python?

Add a `pyclass` wrapper in [`crates/switchyard-py/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/lib.rs) that wraps your Rust algorithm struct. The wrapper should implement methods that convert between Python types and the Rust `Request`/`Step` types defined in the protocol crate. Ensure the underlying Rust API remains stable, as breaking changes will affect Python consumers who import the `switchyard` module.

### Where should observability metrics be emitted in a custom algorithm?

Use the `observability` module exported from [`crates/libsy/src/observability.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/observability.rs) rather than printing to stdout or creating custom metric collectors. This integrates with the server's Prometheus endpoint at `/metrics`, providing consistent monitoring of routing decisions and latency across all Switchyard deployment modes.