Switchyard Advisor Gate vs Stage Routing: Functional Differences in Plan Evaluation

Advisor Gate provides quality-control checkpoints that can rewrite plans via APPROVE/REDO verdicts, while Stage Routing dynamically selects between efficient and capable models for each turn without modifying the underlying plan.

Plan evaluation in NVIDIA-NeMo/Switchyard relies on two distinct routing strategies to optimize performance and accuracy. Understanding the functional difference between Switchyard's advisor gate and stage routing in plan evaluation is essential for designing cost-effective, reliable agentic LLM applications. While both mechanisms determine how requests flow through the system, they operate at different granularities—the Advisor Gate intercepts and validates terminal turns, while Stage Router handles per-turn model selection.

Architectural Goals and Design Philosophy

Advisor Gate as a Quality Checkpoint

The Advisor Gate serves as a protective barrier for the executor model. According to the implementation in crates/libsy/src/algorithms/advisor_gate.rs, this strategy assumes a single executor model handles every client-visible turn, but intercepts the first terminal turn—defined as a turn with no tool calls or a matching pattern—to request validation from a stronger advisor model. The advisor evaluates the buffered turn and returns either APPROVE or REDO.

Stage Router as a Dynamic Model Selector

In contrast, Stage Routing—implemented in crates/libsy/src/algorithms/util/stage.rs—focuses on economic efficiency and capability matching. Rather than validating content quality, it chooses which model tier (efficient vs. capable) should generate each individual turn. The router analyzes tool-result signals including severity, spinning, exploring, and production_intensity to compute a confidence score that determines routing.

Execution Flow and Trigger Conditions

When the Advisor Gate Fires

The Advisor Gate does not evaluate every turn. Instead, it triggers based on specific criteria defined in docs/routing_algorithms/advisor_gate_routing.md:

  • Gate Trigger: Defaults to no_tool_call but can match specific patterns
  • Checkpoint Stalling: Optionally activates after gate_stall_turns mid-task checkpoints
  • Minimum Tool Results: Respects gate_min_tool_results before allowing evaluation
  • Review Budget: Enforces max_reviews limit per session (typically one review per session)

When triggered, the system sends the entire transcript—including task context, tool results, and the buffered terminal turn—to the advisor endpoint.

Per-Turn Evaluation in Stage Routing

The Stage Router evaluates every LLM call without exception. As detailed in docs/routing_algorithms/stage_router_routing.md, the algorithm:

  1. Scores the current turn using signal extraction from tool results
  2. Compares the aggregated confidence against confidence_threshold
  3. Routes to the capable target if confidence is high (indicating complex needs) or the efficient target if the turn appears routine
  4. Optionally falls back to an LLM classifier if confidence scores are ambiguous

Impact on Plan Execution

Plan Modification Capabilities

The functional distinction becomes clearest when examining how each strategy affects the execution plan:

  • Advisor Gate: Can rewrite the plan. If the advisor returns REDO, the executor's premature "done" claim is discarded, and the advisor's revised plan is injected as user feedback. The executor then re-invokes with corrected instructions.
  • Stage Router: Preserves the original plan structure but may change the underlying model generating subsequent turns. This affects latency and cost but does not alter the plan's logical flow.

Budget and Failure Handling

Advisor Gate implements explicit budget controls through max_reviews and supports fail_open semantics—if the advisor service fails, the gate defaults to APPROVE to maintain availability. Stage Router operates without review limits; every turn incurs a routing decision, constrained only by the confidence_threshold parameter.

Configuration and Implementation

Configuring the Advisor Gate

The Advisor Gate requires defining distinct executor and advisor targets in TOML configuration:

[targets.executor]
id = "small/model"
llm_client = "provider"

[targets.advisor]
id = "frontier/model"
llm_client = "provider"

[routes.gated]
id = "switchyard/gated"
type = "advisor"
executor_target = "executor"
advisor_target = "advisor"
max_reviews = 3
gate_stall_turns = 30
gate_min_tool_results = 3

Python clients access this route using the switchyard/gated model identifier:

import switchyard
from switchyard import SwitchyardClient

client = SwitchyardClient(
    url="http://localhost:8000/v1/chat/completions",
    model="switchyard/gated",   # route type = "advisor"

)

response = client.chat(
    messages=[
        {"role": "system", "content": "You are a coding assistant."},
        {"role": "user", "content": "Write a function to add two numbers."}
    ]
)
print(response.choices[0].message["content"])

Setting Up Stage Routing

Stage Routing requires defining capability tiers and selection logic:

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "strong"
efficient_target = "weak"
picker = "efficient_first"
confidence_threshold = 0.5

Access via the Python SDK uses the switchyard/stage route:

import switchyard
from switchyard import SwitchyardClient

client = SwitchyardClient(
    url="http://localhost:8000/v1/chat/completions",
    model="switchyard/stage",   # route type = "stage_router"

)

response = client.chat(
    messages=[
        {"role": "system", "content": "You are a coding assistant."},
        {"role": "user", "content": "Write a function to add two numbers."}
    ]
)
print(response.choices[0].message["content"])

Observability and Telemetry

Both routing strategies expose distinct metrics through the /v1/stats endpoint. The Advisor Gate reports to advisor_gate statistics, tracking verdict distributions, consult failures, and discarded tokens per session. Stage Router metrics appear under stage_router, capturing routing decisions, confidence score distributions, and signal frequencies (severity counts, spinning detection).

These telemetry endpoints—implemented in crates/switchyard-server/src/stats/algorithms/advisor_gate.rs and crates/switchyard-server/src/stats/algorithms/stage_router.rs respectively—enable operational monitoring of review overhead versus routing efficiency.

Summary

  • Advisor Gate intervenes at terminal turns to validate plan completion, potentially rewriting the plan via advisor feedback while operating under a strict max_reviews budget.
  • Stage Router evaluates every turn to minimize costs by selecting appropriate model tiers, using signal-based confidence scoring without altering plan structure.
  • The Advisor Gate sends complete transcripts to a dedicated advisor endpoint; Stage Router uses lightweight signal extraction for local decision-making.
  • Configuration differs fundamentally: Advisor Gate pairs an executor with an advisor, while Stage Router balances efficient and capable targets against confidence thresholds.

Frequently Asked Questions

Can Advisor Gate and Stage Routing be used together in the same deployment?

Yes, these mechanisms address orthogonal concerns and can be composed. You might use Stage Routing to select the initial executor model tier, then apply an Advisor Gate to validate the final output before client delivery. The configurations operate on different routing identifiers (type = "advisor" vs type = "stage_router") and can coexist within the same Switchyard instance.

What happens when the Advisor Gate returns REDO?

When the advisor returns REDO, the buffered terminal turn is discarded rather than sent to the client. The advisor's revised plan is injected into the conversation history as user feedback, and the executor model re-invokes against this corrected context. This effectively rewinds the conversation by one turn and substitutes the advisor's reasoning for the executor's premature completion claim, as implemented in crates/libsy/src/algorithms/advisor_gate.rs.

How does Stage Routing handle uncertainty when confidence scores are ambiguous?

If the signal-based confidence falls below confidence_threshold, the Stage Router can escalate to an optional LLM classifier for additional evaluation. If the classifier also returns uncertain results, the system typically defaults to the capable target to ensure quality, though this behavior is configurable via the picker strategy parameter in the route configuration.

Which routing strategy reduces API costs more effectively?

Stage Routing typically delivers greater cost reduction for multi-turn conversations by routing routine steps to cheaper efficient models and reserving expensive capable models for complex turns. The Advisor Gate incurs overhead only at terminal turns but involves sending full transcripts to a powerful advisor model, making it more cost-effective for preventing expensive downstream errors rather than optimizing per-turn expenses.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →