# Switchyard Advisor Gate vs Stage Routing: Functional Differences in Plan Evaluation

> Understand the functional differences between Switchyard Advisor Gate and Stage Routing. Advisor Gate offers quality control checkpoints, while Stage Routing dynamically selects models for plan evaluation.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-12

---

**Advisor Gate provides quality-control checkpoints that can rewrite plans via APPROVE/REDO verdicts, while Stage Routing dynamically selects between efficient and capable models for each turn without modifying the underlying plan.**

Plan evaluation in NVIDIA-NeMo/Switchyard relies on two distinct routing strategies to optimize performance and accuracy. Understanding the functional difference between Switchyard's advisor gate and stage routing in plan evaluation is essential for designing cost-effective, reliable agentic LLM applications. While both mechanisms determine how requests flow through the system, they operate at different granularities—the Advisor Gate intercepts and validates terminal turns, while Stage Router handles per-turn model selection.

## Architectural Goals and Design Philosophy

### Advisor Gate as a Quality Checkpoint

The **Advisor Gate** serves as a protective barrier for the executor model. According to the implementation in [`crates/libsy/src/algorithms/advisor_gate.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/advisor_gate.rs), this strategy assumes a single executor model handles every client-visible turn, but intercepts the *first terminal turn*—defined as a turn with no tool calls or a matching pattern—to request validation from a stronger advisor model. The advisor evaluates the buffered turn and returns either **APPROVE** or **REDO**.

### Stage Router as a Dynamic Model Selector

In contrast, **Stage Routing**—implemented in [`crates/libsy/src/algorithms/util/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/stage.rs)—focuses on economic efficiency and capability matching. Rather than validating content quality, it chooses which model tier (efficient vs. capable) should generate each individual turn. The router analyzes tool-result signals including `severity`, `spinning`, `exploring`, and `production_intensity` to compute a confidence score that determines routing.

## Execution Flow and Trigger Conditions

### When the Advisor Gate Fires

The Advisor Gate does not evaluate every turn. Instead, it triggers based on specific criteria defined in [`docs/routing_algorithms/advisor_gate_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/advisor_gate_routing.md):

- **Gate Trigger**: Defaults to `no_tool_call` but can match specific patterns
- **Checkpoint Stalling**: Optionally activates after `gate_stall_turns` mid-task checkpoints
- **Minimum Tool Results**: Respects `gate_min_tool_results` before allowing evaluation
- **Review Budget**: Enforces `max_reviews` limit per session (typically one review per session)

When triggered, the system sends the *entire transcript*—including task context, tool results, and the buffered terminal turn—to the advisor endpoint.

### Per-Turn Evaluation in Stage Routing

The Stage Router evaluates **every LLM call** without exception. As detailed in [`docs/routing_algorithms/stage_router_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/stage_router_routing.md), the algorithm:

1. Scores the current turn using signal extraction from tool results
2. Compares the aggregated confidence against `confidence_threshold`
3. Routes to the **capable** target if confidence is high (indicating complex needs) or the **efficient** target if the turn appears routine
4. Optionally falls back to an LLM classifier if confidence scores are ambiguous

## Impact on Plan Execution

### Plan Modification Capabilities

The functional distinction becomes clearest when examining how each strategy affects the execution plan:

- **Advisor Gate**: Can *rewrite* the plan. If the advisor returns **REDO**, the executor's premature "done" claim is discarded, and the advisor's revised plan is injected as user feedback. The executor then re-invokes with corrected instructions.
- **Stage Router**: Preserves the original plan structure but may *change the underlying model* generating subsequent turns. This affects latency and cost but does not alter the plan's logical flow.

### Budget and Failure Handling

Advisor Gate implements explicit budget controls through `max_reviews` and supports `fail_open` semantics—if the advisor service fails, the gate defaults to **APPROVE** to maintain availability. Stage Router operates without review limits; every turn incurs a routing decision, constrained only by the `confidence_threshold` parameter.

## Configuration and Implementation

### Configuring the Advisor Gate

The Advisor Gate requires defining distinct executor and advisor targets in TOML configuration:

```toml
[targets.executor]
id = "small/model"
llm_client = "provider"

[targets.advisor]
id = "frontier/model"
llm_client = "provider"

[routes.gated]
id = "switchyard/gated"
type = "advisor"
executor_target = "executor"
advisor_target = "advisor"
max_reviews = 3
gate_stall_turns = 30
gate_min_tool_results = 3

```

Python clients access this route using the `switchyard/gated` model identifier:

```python
import switchyard
from switchyard import SwitchyardClient

client = SwitchyardClient(
    url="http://localhost:8000/v1/chat/completions",
    model="switchyard/gated",   # route type = "advisor"

)

response = client.chat(
    messages=[
        {"role": "system", "content": "You are a coding assistant."},
        {"role": "user", "content": "Write a function to add two numbers."}
    ]
)
print(response.choices[0].message["content"])

```

### Setting Up Stage Routing

Stage Routing requires defining capability tiers and selection logic:

```toml
[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "strong"
efficient_target = "weak"
picker = "efficient_first"
confidence_threshold = 0.5

```

Access via the Python SDK uses the `switchyard/stage` route:

```python
import switchyard
from switchyard import SwitchyardClient

client = SwitchyardClient(
    url="http://localhost:8000/v1/chat/completions",
    model="switchyard/stage",   # route type = "stage_router"

)

response = client.chat(
    messages=[
        {"role": "system", "content": "You are a coding assistant."},
        {"role": "user", "content": "Write a function to add two numbers."}
    ]
)
print(response.choices[0].message["content"])

```

## Observability and Telemetry

Both routing strategies expose distinct metrics through the `/v1/stats` endpoint. The Advisor Gate reports to `advisor_gate` statistics, tracking verdict distributions, consult failures, and discarded tokens per session. Stage Router metrics appear under `stage_router`, capturing routing decisions, confidence score distributions, and signal frequencies (severity counts, spinning detection).

These telemetry endpoints—implemented in [`crates/switchyard-server/src/stats/algorithms/advisor_gate.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/advisor_gate.rs) and [`crates/switchyard-server/src/stats/algorithms/stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/stage_router.rs) respectively—enable operational monitoring of review overhead versus routing efficiency.

## Summary

- **Advisor Gate** intervenes at terminal turns to validate plan completion, potentially rewriting the plan via advisor feedback while operating under a strict `max_reviews` budget.
- **Stage Router** evaluates every turn to minimize costs by selecting appropriate model tiers, using signal-based confidence scoring without altering plan structure.
- The Advisor Gate sends complete transcripts to a dedicated advisor endpoint; Stage Router uses lightweight signal extraction for local decision-making.
- Configuration differs fundamentally: Advisor Gate pairs an executor with an advisor, while Stage Router balances efficient and capable targets against confidence thresholds.

## Frequently Asked Questions

### Can Advisor Gate and Stage Routing be used together in the same deployment?

Yes, these mechanisms address orthogonal concerns and can be composed. You might use Stage Routing to select the initial executor model tier, then apply an Advisor Gate to validate the final output before client delivery. The configurations operate on different routing identifiers (`type = "advisor"` vs `type = "stage_router"`) and can coexist within the same Switchyard instance.

### What happens when the Advisor Gate returns REDO?

When the advisor returns **REDO**, the buffered terminal turn is discarded rather than sent to the client. The advisor's revised plan is injected into the conversation history as user feedback, and the executor model re-invokes against this corrected context. This effectively rewinds the conversation by one turn and substitutes the advisor's reasoning for the executor's premature completion claim, as implemented in [`crates/libsy/src/algorithms/advisor_gate.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/advisor_gate.rs).

### How does Stage Routing handle uncertainty when confidence scores are ambiguous?

If the signal-based confidence falls below `confidence_threshold`, the Stage Router can escalate to an optional LLM classifier for additional evaluation. If the classifier also returns uncertain results, the system typically defaults to the **capable** target to ensure quality, though this behavior is configurable via the `picker` strategy parameter in the route configuration.

### Which routing strategy reduces API costs more effectively?

**Stage Routing** typically delivers greater cost reduction for multi-turn conversations by routing routine steps to cheaper efficient models and reserving expensive capable models for complex turns. The **Advisor Gate** incurs overhead only at terminal turns but involves sending full transcripts to a powerful advisor model, making it more cost-effective for preventing expensive downstream errors rather than optimizing per-turn expenses.