# What Is Stage Routing in Switchyard? A Signal-Driven LLM Routing Algorithm

> Discover Stage routing in NVIDIA Switchyard, a signal-driven LLM algorithm that optimizes model selection for quality and cost based on real-time context and history. Learn how it works.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-08-21

---

**Stage routing is a signal-driven routing algorithm in NVIDIA Switchyard that dynamically selects between high-quality "capable" models and cost-effective "efficient" models based on real-time analysis of tool-result history and conversation context.**

Stage routing in Switchyard optimizes inference costs for coding-agent workloads by routing each LLM turn to the most appropriate model tier. Developed in the `NVIDIA-NeMo/Switchyard` repository, this algorithm examines signals from recent tool executions to determine whether a request requires expensive, high-quality reasoning or can be handled by a cheaper, faster model.

## Why Switchyard Uses Stage Routing for Coding Agents

Coding agent workloads exhibit distinct behavioral patterns across conversation turns. **Early turns** typically involve exploration, error recovery, and complex reasoning that benefit from capable models with strong reasoning capabilities. **Later turns** often become mechanical, involving simple file writes, edits, or repetitive tasks that do not require advanced reasoning.

By analyzing these patterns, Stage routing reduces operational costs without sacrificing output quality. The system routes exploratory, high-cognitive-load requests to expensive capable models while delegating routine, mechanical operations to efficient models.

## Core Concepts of the Stage Router

The Stage router evaluates four primary components to make routing decisions:

**Tool-Result Signals**  
The algorithm extracts behavioral indicators from the conversation's tool-result history, including:
- Severity of recent errors
- Spinning states (repeated attempts without progress)  
- Exploration patterns (branching or trial-and-error behaviors)
- Recent production intensity (volume of successful outputs)

**Confidence Scoring**  
These signals combine into a signed score that undergoes a `tanh` squashing function, producing a final confidence value bounded between `[0, 1]`.

**Confidence Threshold**  
When confidence exceeds the configured `confidence_threshold`, the router flips the turn to the opposite tier from the default. Values below the threshold trigger fallback behaviors.

**Picker Modes**  
- `efficient_first` (default): Makes the efficient tier the default, escalating to capable only when signals warrant it
- `capable_first` (experimental): Makes the capable tier the default, demoting to efficient only when confidence is low

**Optional LLM Classifier**  
When signal confidence falls below the threshold, the system can consult an `LlmTaskClassifier` to break ties and determine the appropriate tier.

## How Stage Routing Works Internally

According to the Switchyard source code in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs), the implementation assembles a processing cascade:

1. **`ToolSignalProcessor`** extracts and parses recent tool results from the conversation history
2. **`StageClassifier`** scores these signals against both model tiers, applying the `tanh` normalization to generate confidence values
3. **`LlmTaskClassifier`** (optional) evaluates ambiguous cases when signal confidence is insufficient
4. **`FallThrough`** driver executes the cascade, ultimately defaulting to the picker's configured tier if no classifier claims the turn

The signal scoring logic resides in [`crates/libsy/src/algorithms/util/tool_signals.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/tool_signals.rs), which defines how raw tool outputs translate into the severity and behavioral metrics used by the classifier.

## Configuring Stage Routing in TOML

To enable Stage routing, define a route in your [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration file:

```toml
[schema]
version = 1

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "strong"
efficient_target = "weak"
picker = "efficient_first"
confidence_threshold = 0.5
recent_turn_window = 3

```

Start the server with your configuration:

```bash
switchyard-server --config routes.toml --port 4000

```

The `recent_turn_window` parameter controls how many previous turns the `ToolSignalProcessor` examines when extracting signals.

## Using the Stage Router in Python and Rust

### Python Implementation

The Python API allows direct instantiation of the stage router for custom orchestration:

```python
from switchyard.libsy import algorithms, Step, LlmResponse

async def run_one_turn():
    router = algorithms.stage_router(
        capable="strong",          # capable model id from routes.toml

        efficient="weak",          # efficient model id from routes.toml

        picker="efficient_first",
        confidence_threshold=0.5,
    )
    async for step in router.run_stream(your_normalized_turn):
        match step:
            case Step.CallModel(call):
                response = await your_llm_client.call(call.request)
                call.respond(LlmResponse.Agg(response))
            case Step.Done(outcome):
                print("Selected model:", outcome.selected_model_id)

```

This example appears in the repository at [`examples/experimental/litellm/example.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/experimental/litellm/example.py).

### Rust Implementation

For native performance, implement the router directly in Rust:

```rust
use switchyard_libsy::{StageRouter, StageRouterConfig, PickerMode, ModelId};

let capable = ModelId::from("strong");
let efficient = ModelId::from("weak");
let config = StageRouterConfig::new(PickerMode::EfficientFirst, 0.5);
let router = StageRouter::new(capable, efficient, config)?;

```

The `StageRouter` struct and `StageRouterConfig` are defined in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs), with `PickerMode` supporting both `EfficientFirst` and `CapableFirst` variants.

## Summary

- Stage routing in Switchyard intelligently routes LLM requests between **capable** (expensive, high-quality) and **efficient** (cheap, fast) model tiers based on real-time conversation analysis.
- The algorithm extracts **tool-result signals** (severity, spinning, exploring, production intensity) and applies a `tanh` squashing function to generate a confidence score in `[0, 1]`.
- Configuration occurs in TOML via `type = "stage_router"` with parameters for `picker` mode, `confidence_threshold`, and `recent_turn_window`.
- The implementation cascade in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs) processes signals through `ToolSignalProcessor` and `StageClassifier`, with optional `LlmTaskClassifier` fallback.
- Both Python and Rust APIs support programmatic router instantiation for integration into custom agent frameworks.

## Frequently Asked Questions

### How does Stage routing determine which model tier to use?

Stage routing analyzes the **tool-result history** of recent conversation turns (controlled by `recent_turn_window`). It extracts signals like error severity, exploration patterns, and production intensity, then combines these into a confidence score using a `tanh` function. If confidence exceeds the configured threshold, the router switches to the non-default tier; otherwise, it uses the `picker` default (`efficient_first` or `capable_first`).

### Can Stage routing operate without an LLM classifier?

Yes. The `LlmTaskClassifier` is optional. When signal confidence falls below `confidence_threshold` and no classifier is configured, the `FallThrough` driver automatically defaults to the picker’s configured tier. This setup reduces latency and cost by avoiding additional LLM calls for borderline cases.

### What is the difference between `efficient_first` and `capable_first` picker modes?

`efficient_first` (the default) routes requests to the efficient model unless signals indicate high complexity or errors, optimizing for cost. `capable_first` (experimental) defaults to the capable model unless the conversation shows purely mechanical patterns, prioritizing quality over cost savings.

### Where can I find the Stage router implementation in the Switchyard codebase?

The primary implementation resides in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs), which defines the `StageRouter` struct, configuration, and processing cascade. Signal extraction logic is located in [`crates/libsy/src/algorithms/util/tool_signals.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/tool_signals.rs). Documentation and configuration examples appear in [`docs/routing_algorithms/stage_router_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/stage_router_routing.md).