# How to Configure Switchyard's Routing Algorithms: TOML and Rust Integration Guide

> Learn to configure Switchyard routing algorithms using TOML definitions and Rust integration. Explore stage_router and llm_classifier strategies with the switchyard.libsy Python package.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-21

---

**Switchyard routes LLM requests through configurable algorithms defined in TOML route definitions, exposing Rust-implemented strategies like `stage_router` and `llm_classifier` via the `switchyard.libsy` Python package.**

Switchyard, NVIDIA's open-source LLM serving framework, determines how incoming requests reach target models through pluggable routing algorithms. You configure these algorithms in **TOML route definitions** that bind specific decision-making logic to your model deployments. Whether you are running simple A/B tests or complex multi-stage inference pipelines, understanding how to configure Switchyard's routing algorithms ensures optimal traffic distribution across your model fleet.

## Understanding Switchyard's Routing Architecture

Switchyard implements a three-tier architecture that separates route configuration from algorithm execution. The routing core lives in the Rust `libsy` crate and is exposed to Python via `switchyard.libsy`, allowing both declarative TOML configuration and programmatic Rust embedding.

### Route Definition Layer (TOML)

Every route in Switchyard is declared in a TOML configuration file that specifies both the target model and the algorithm used to reach it. According to the TOML schema documented in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md), each route entry must include an `algorithm` field that selects the decision-making strategy. This declarative approach lets operators change routing behavior without modifying application code.

### Algorithm Factory Bridge

The Python module [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) serves as the bridge between configuration and execution. This file re-exports Rust-implemented algorithms—including `llm_classifier`, `llm_task_classifier`, `noop`, `random`, and `stage_router`—making them available to the Switchyard server. When the server loads your TOML configuration, it instantiates the corresponding Rust algorithm objects through this Python façade.

### Routing Execution Engine

When a request arrives, the server normalizes it and invokes the selected algorithm. The algorithm yields a stream of execution steps—such as `CallModel` or `ClassifierCall`—until a final response is generated. This process is documented in [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md), which illustrates how algorithms dynamically select targets based on request content or predefined rules.

## Available Routing Algorithms

Switchyard provides several built-in algorithms optimized for different deployment scenarios. Choose an algorithm based on whether you need simple traffic splitting, content-aware routing, or cost-optimized escalation strategies.

| Algorithm | TOML Value | Use Case |
|-----------|------------|----------|
| **Passthrough** | `passthrough` | Simple one-target deployments with no routing logic |
| **Random** | `random` | Fixed traffic splitting for A/B testing or baseline measurements |
| **LLM Classifier** | `llm_classifier` | Content-aware decisions routing requests to weak vs. strong model tiers |
| **Stage Router** | `stage_router` | Multi-stage inference using tool results and progress signals to select efficient targets |
| **Escalation Router** | `llm_classifier` (with escalation mode) | Cost-efficient routing that starts with cheap models and escalates when difficulty is detected |
| **Advisor Gate** | `advisor` | Single-model execution with a stronger reviewer gating "done" claims |

Detailed configuration options for each algorithm are available in the individual documentation files under `docs/routing_algorithms/`, such as [`stage_router_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/stage_router_routing.md) for the stage router implementation.

## Configuring Algorithms in TOML

To configure Switchyard's routing algorithms, define your routes in a TOML file using the `algorithm` field to select the strategy and the optional `algorithm_config` table for algorithm-specific parameters.

```toml
[[routes]]
name = "efficient_inference"
target = "primary-model"
algorithm = "stage_router"

[routes.algorithm_config]
efficient_first = true
fallback_target = "fallback-model"

```

In this configuration:
- The `algorithm` field selects the Rust implementation instantiated from [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py)
- The `algorithm_config` table passes parameters directly to the Rust side; for `stage_router`, options include `efficient_first` and `fallback_target`
- The `target` specifies which model deployment receives the request when the algorithm selects it

The complete schema for route definitions, including all valid `algorithm` values and their configuration parameters, is documented in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md).

## Embedding Algorithms in Rust

For applications requiring tighter integration, you can configure routing algorithms directly in Rust using the `switchyard-libsy` crate. This bypasses the TOML configuration layer and allows dynamic algorithm construction at compile time.

```rust
use switchyard_libsy::{stage_router::StageRouter, routing::Algorithm};

let algorithm = Algorithm::StageRouter(StageRouter::new(
    efficient_first: true,
    fallback_target: Some("fallback-model".into()),
));

```

The constructed `Algorithm` enum variant can then be injected into the `switchyard_server` runtime as part of a programmatic route definition. The Rust source implementations for each algorithm reside in `crates/libsy/src/` (e.g., [`stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/stage_router.rs), [`random.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/random.rs)), while the public API is documented in [`docs/reference/rust_api.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/rust_api.md).

## Summary

- Switchyard routing algorithms are configured in TOML route definitions under the `algorithm` field, with algorithm-specific options nested in `algorithm_config`
- The [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) module exposes Rust-implemented strategies including `random`, `llm_classifier`, and `stage_router` to the Python runtime
- Algorithms execute as step generators (yielding `CallModel`, `ClassifierCall`, etc.) until producing a final response
- For embedded deployments, instantiate algorithms directly via the `switchyard-libsy` Rust crate using the documented constructors
- Reference the TOML schema in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md) and algorithm-specific docs in `docs/routing_algorithms/` for detailed parameter specifications

## Frequently Asked Questions

### What file contains the built-in algorithm definitions in Switchyard?

The Python façade for all built-in algorithms is located in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py), which imports and re-exports the Rust implementations from the `libsy` crate. The actual algorithm logic resides in the Rust source files under `crates/libsy/src/` (such as [`stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/stage_router.rs) and [`random.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/random.rs)).

### How do I pass custom parameters to a routing algorithm?

Use the `algorithm_config` table within your TOML route definition. This nested configuration object accepts algorithm-specific keys—such as `efficient_first` or `fallback_target` for the stage router—that are serialized and passed directly to the Rust implementation. Consult [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md) for the complete list of valid parameters for each algorithm.

### Can I implement a custom routing algorithm in Python?

While Switchyard exposes algorithms to Python via `switchyard.libsy`, the core routing implementations are written in Rust for performance. The current architecture requires custom algorithms to be implemented in Rust within the `libsy` crate and exposed through the algorithm factory pattern, though they can then be instantiated from Python or TOML configurations.

### Which algorithm should I use for cost-efficient inference?

Use the **Stage Router** (`stage_router`) for general cost optimization, as it leverages tool results and progress signals to select the most efficient target for each request. For workloads with highly variable complexity, consider the **Escalation Router** mode of `llm_classifier`, which starts with cheaper models and only escalates to expensive ones when a judge detects high difficulty.