What Is Stage Routing in Switchyard? A Signal-Driven LLM Routing Algorithm

Stage routing is a signal-driven routing algorithm in NVIDIA Switchyard that dynamically selects between high-quality "capable" models and cost-effective "efficient" models based on real-time analysis of tool-result history and conversation context.

Stage routing in Switchyard optimizes inference costs for coding-agent workloads by routing each LLM turn to the most appropriate model tier. Developed in the NVIDIA-NeMo/Switchyard repository, this algorithm examines signals from recent tool executions to determine whether a request requires expensive, high-quality reasoning or can be handled by a cheaper, faster model.

Why Switchyard Uses Stage Routing for Coding Agents

Coding agent workloads exhibit distinct behavioral patterns across conversation turns. Early turns typically involve exploration, error recovery, and complex reasoning that benefit from capable models with strong reasoning capabilities. Later turns often become mechanical, involving simple file writes, edits, or repetitive tasks that do not require advanced reasoning.

By analyzing these patterns, Stage routing reduces operational costs without sacrificing output quality. The system routes exploratory, high-cognitive-load requests to expensive capable models while delegating routine, mechanical operations to efficient models.

Core Concepts of the Stage Router

The Stage router evaluates four primary components to make routing decisions:

Tool-Result Signals
The algorithm extracts behavioral indicators from the conversation's tool-result history, including:

  • Severity of recent errors
  • Spinning states (repeated attempts without progress)
  • Exploration patterns (branching or trial-and-error behaviors)
  • Recent production intensity (volume of successful outputs)

Confidence Scoring
These signals combine into a signed score that undergoes a tanh squashing function, producing a final confidence value bounded between [0, 1].

Confidence Threshold
When confidence exceeds the configured confidence_threshold, the router flips the turn to the opposite tier from the default. Values below the threshold trigger fallback behaviors.

Picker Modes

  • efficient_first (default): Makes the efficient tier the default, escalating to capable only when signals warrant it
  • capable_first (experimental): Makes the capable tier the default, demoting to efficient only when confidence is low

Optional LLM Classifier
When signal confidence falls below the threshold, the system can consult an LlmTaskClassifier to break ties and determine the appropriate tier.

How Stage Routing Works Internally

According to the Switchyard source code in crates/libsy/src/algorithms/stage.rs, the implementation assembles a processing cascade:

  1. ToolSignalProcessor extracts and parses recent tool results from the conversation history
  2. StageClassifier scores these signals against both model tiers, applying the tanh normalization to generate confidence values
  3. LlmTaskClassifier (optional) evaluates ambiguous cases when signal confidence is insufficient
  4. FallThrough driver executes the cascade, ultimately defaulting to the picker's configured tier if no classifier claims the turn

The signal scoring logic resides in crates/libsy/src/algorithms/util/tool_signals.rs, which defines how raw tool outputs translate into the severity and behavioral metrics used by the classifier.

Configuring Stage Routing in TOML

To enable Stage routing, define a route in your routes.toml configuration file:

[schema]
version = 1

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "strong"
efficient_target = "weak"
picker = "efficient_first"
confidence_threshold = 0.5
recent_turn_window = 3

Start the server with your configuration:

switchyard-server --config routes.toml --port 4000

The recent_turn_window parameter controls how many previous turns the ToolSignalProcessor examines when extracting signals.

Using the Stage Router in Python and Rust

Python Implementation

The Python API allows direct instantiation of the stage router for custom orchestration:

from switchyard.libsy import algorithms, Step, LlmResponse

async def run_one_turn():
    router = algorithms.stage_router(
        capable="strong",          # capable model id from routes.toml

        efficient="weak",          # efficient model id from routes.toml

        picker="efficient_first",
        confidence_threshold=0.5,
    )
    async for step in router.run_stream(your_normalized_turn):
        match step:
            case Step.CallModel(call):
                response = await your_llm_client.call(call.request)
                call.respond(LlmResponse.Agg(response))
            case Step.Done(outcome):
                print("Selected model:", outcome.selected_model_id)

This example appears in the repository at examples/experimental/litellm/example.py.

Rust Implementation

For native performance, implement the router directly in Rust:

use switchyard_libsy::{StageRouter, StageRouterConfig, PickerMode, ModelId};

let capable = ModelId::from("strong");
let efficient = ModelId::from("weak");
let config = StageRouterConfig::new(PickerMode::EfficientFirst, 0.5);
let router = StageRouter::new(capable, efficient, config)?;

The StageRouter struct and StageRouterConfig are defined in crates/libsy/src/algorithms/stage.rs, with PickerMode supporting both EfficientFirst and CapableFirst variants.

Summary

  • Stage routing in Switchyard intelligently routes LLM requests between capable (expensive, high-quality) and efficient (cheap, fast) model tiers based on real-time conversation analysis.
  • The algorithm extracts tool-result signals (severity, spinning, exploring, production intensity) and applies a tanh squashing function to generate a confidence score in [0, 1].
  • Configuration occurs in TOML via type = "stage_router" with parameters for picker mode, confidence_threshold, and recent_turn_window.
  • The implementation cascade in crates/libsy/src/algorithms/stage.rs processes signals through ToolSignalProcessor and StageClassifier, with optional LlmTaskClassifier fallback.
  • Both Python and Rust APIs support programmatic router instantiation for integration into custom agent frameworks.

Frequently Asked Questions

How does Stage routing determine which model tier to use?

Stage routing analyzes the tool-result history of recent conversation turns (controlled by recent_turn_window). It extracts signals like error severity, exploration patterns, and production intensity, then combines these into a confidence score using a tanh function. If confidence exceeds the configured threshold, the router switches to the non-default tier; otherwise, it uses the picker default (efficient_first or capable_first).

Can Stage routing operate without an LLM classifier?

Yes. The LlmTaskClassifier is optional. When signal confidence falls below confidence_threshold and no classifier is configured, the FallThrough driver automatically defaults to the picker’s configured tier. This setup reduces latency and cost by avoiding additional LLM calls for borderline cases.

What is the difference between efficient_first and capable_first picker modes?

efficient_first (the default) routes requests to the efficient model unless signals indicate high complexity or errors, optimizing for cost. capable_first (experimental) defaults to the capable model unless the conversation shows purely mechanical patterns, prioritizing quality over cost savings.

Where can I find the Stage router implementation in the Switchyard codebase?

The primary implementation resides in crates/libsy/src/algorithms/stage.rs, which defines the StageRouter struct, configuration, and processing cascade. Signal extraction logic is located in crates/libsy/src/algorithms/util/tool_signals.rs. Documentation and configuration examples appear in docs/routing_algorithms/stage_router_routing.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →