How the Switchyard Stage Router Chooses Between Tool-Response Pattern Matching and an LLM Judge
The Switchyard stage router evaluates six built-in tool-response signals to calculate a confidence score; if that score exceeds a configurable threshold, it routes to the capable tier immediately, otherwise it either defaults to the efficient tier or consults an optional LLM classifier for a final verdict.
The NVIDIA-NeMo/Switchyard framework provides an intelligent routing layer that dynamically selects between high-capability and high-efficiency language models on a per-turn basis. Understanding how the stage router balances deterministic signal processing against learned LLM judgment is essential for optimizing both cost and accuracy in production agent systems.
Understanding the Two-Tier Routing Architecture
The stage router operates on a capable versus efficient model dichotomy. For every conversational turn, the system must decide whether complex reasoning (capable) or fast, cheap inference (efficient) is appropriate. Rather than relying on a single heuristic, the router implements a two-step decision mechanism that first analyzes tool execution patterns, then optionally defers to a neural judge when confidence is low.
Signal-Based Routing with Tool-Response Pattern Matching
At the core of the stage router is a compact history analysis of recent tool results, including file reads, writes, edits, test runs, and shell commands. The system extracts six distinct signals from this telemetry to compute a routing confidence value.
The Six Built-in Signals
| Signal | Description | Routing Effect |
|---|---|---|
| severity | Windowed error severity from tool executions | Pushes toward capable |
| spinning | Repeated attempts without substantive reads/writes | Pushes toward capable |
| exploring | Planning or reading activity without output production | Pushes toward capable |
| recent_production_intensity | Frequency of writes/edits in the recent window | Pushes toward efficient |
| confidence | Derived tanh-squashed aggregate of signal weights |
Determines threshold crossing |
| override | Hard-override flag for critical/fatal errors | Forces capable immediately |
According to the source code in crates/libsy/src/algorithms/util/tool_signals.rs, these signals are extracted from the conversation state and weighted to produce a composite score.
Confidence Score Calculation
The router applies a tanh squashing function to the aggregated signal weights, yielding a normalized confidence value in the range [0, 1]. In crates/libsy/src/algorithms/util/stage.rs, this value is exposed as the CONFIDENCE_METRIC.
If the confidence score exceeds the user-defined confidence_threshold (defaulting to 0.5), the router overrides the picker’s default tier and selects the opposite of the current default (e.g., flipping from efficient to capable). If the score falls below the threshold, the system respects the picker’s default strategy (efficient_first or capable_first) or proceeds to the optional LLM judge.
When the LLM Judge Steps In
When tool-response patterns provide ambiguous evidence, the stage router can defer to an LLM classifier that acts as a learned judge to break ties.
The Classifier Configuration
Enabling the LLM judge requires adding a classifier block to the route configuration in crates/libsy/src/algorithms/stage.rs. When present, the router checks the signal-derived confidence score against the threshold. If the score is below the threshold, the turn is forwarded to the specified classifier route rather than to the default model tier.
[classifier]
id = "switchyard/llm_classifier"
type = "llm_classifier"
model = "openai/gpt-4o"
The Decision Flow
The complete decision chain implemented in stage.rs follows this precedence:
- Extract tool signals and compute confidence score
- If confidence ≥ threshold: route based on signal override (capable or efficient)
- If confidence < threshold and classifier configured: send to LLM judge
- If confidence < threshold and no classifier: use picker default
This design minimizes expensive classifier calls by only invoking the LLM judge when historical tool signals are inconclusive.
Practical Configuration Examples
Example 1: Signal-Only Routing
To use the stage router with pure pattern matching and no LLM judge, omit the classifier block:
[schema]
schema_version = 1
[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"
[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"
[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "strong"
efficient_target = "weak"
picker = "efficient_first"
confidence_threshold = 0.5
With this configuration, the router evaluates tool signals exclusively. High-severity errors or spinning behavior will trigger the capable model, while routine edits remain on the efficient tier.
Example 2: Hybrid Routing with LLM Fallback
To enable the two-stage fallback where ambiguous signals trigger an LLM classifier:
[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "strong"
efficient_target = "weak"
picker = "efficient_first"
confidence_threshold = 0.5
[classifier]
id = "switchyard/llm_classifier"
type = "llm_classifier"
model = "openai/gpt-4o"
Here, any turn with a confidence score below 0.5 is passed to the classifier for a final capable/efficient determination before execution.
Example 3: Python Invocation
The Python bindings exposed in crates/switchyard-py/src/libsy_bindings.rs allow direct invocation of the stage router:
import switchyard
from switchyard import Runner
runner = Runner.from_toml("config.toml")
response = await runner.run(
messages=[{"role": "user", "content": "Please refactor the database layer"}],
route="stage"
)
The Runner automatically applies the signal scoring and classifier logic defined in the TOML configuration.
Summary
- The stage router in NVIDIA-NeMo/Switchyard routes each conversational turn between capable and efficient model tiers.
- Six tool-response signals (severity, spinning, exploring, production intensity, derived confidence, and critical override) are aggregated using a tanh function to produce a [0, 1] confidence score in
tool_signals.rs. - A configurable confidence threshold (default
0.5) determines whether the router overrides the default picker or proceeds to an optional LLM judge. - The LLM classifier is only invoked when the signal confidence is below the threshold, reducing costs while maintaining accuracy on ambiguous inputs.
- Core implementation resides in
crates/libsy/src/algorithms/stage.rswith metrics exported viacrates/libsy/src/algorithms/util/stage.rs.
Frequently Asked Questions
How does the Switchyard stage router calculate confidence scores?
The router extracts six signals from recent tool execution history defined in crates/libsy/src/algorithms/util/tool_signals.rs. These signals are weighted and passed through a tanh squashing function to produce a normalized confidence value between 0 and 1, as implemented in the scoring logic and exposed via the CONFIDENCE_METRIC in the utility module.
When should I enable the LLM classifier versus relying on signals alone?
Enable the LLM classifier when your application encounters ambiguous tool-response patterns that deterministic heuristics cannot reliably classify, such as novel error types or nuanced planning phases. For well-defined workflows with clear success/failure signals (e.g., test suites or file writes), the signal-only approach reduces latency and cost by avoiding unnecessary classifier calls.
What happens when the confidence score exactly equals the threshold?
When the confidence score equals the confidence_threshold exactly, the router treats this as meeting the threshold condition. According to the logic in crates/libsy/src/algorithms/stage.rs, a score greater than or equal to the threshold triggers the signal-based override, routing to the tier indicated by the aggregated signals rather than falling back to the classifier or default picker.
Can I customize the signals used by the stage router?
The current implementation provides a fixed vocabulary of six signals (severity, spinning, exploring, recent_production_intensity, confidence, and override) defined in tool_signals.rs. While you cannot add custom signals without modifying the Rust source, you can tune the routing behavior by adjusting the confidence_threshold or by adding a custom LLM classifier that incorporates domain-specific logic for the ambiguous cases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →