Switchyard Advisor Gate Approval Loop: Mechanism and Signals
The advisor gate in NVIDIA NeMo Switchyard implements an approval loop that intercepts executor turns at terminal states or stall thresholds, consults a stronger advisor model for APPROVE or REDO verdicts, and signals the executor to continue through feedback injection and Prometheus metrics when work needs revision.
The advisor gate routing strategy in NVIDIA-NeMo/Switchyard pairs a fast executor model with a quality-checking advisor to balance latency and accuracy. When the executor produces a turn that might be final, the approval loop mechanism determines whether to deliver the response to the client or send the executor back for additional work. This quality gate protects client-facing outputs while allowing the executor to iterate on complex tasks under explicit guidance.
How the Advisor Gate Approval Loop Works
Trigger Detection
The gate monitors every executor turn for two primary conditions defined in docs/routing_algorithms/advisor_gate_routing.md (lines 23-30). First, no_tool_call fires when the executor produces a turn containing no tool calls—typically indicating a plan or completion claim—provided gate_min_tool_results tool results have been observed to filter early chatter. Second, pattern matches user-defined regex via gate_trigger_pattern against visible turn text.
Mid-Task Checkpointing
Independent of content-based triggers, gate_stall_turns forces a review after a configurable number of assistant turns (lines 32-35). This catches executors that "grind" through iterations without declaring completion, ensuring the advisor can intervene during long-running tasks before token budgets exhaust.
Advisor Consultation
When triggered, the gate serializes the full session transcript—including task description, tool results, and the gated turn—and transmits it to the advisor under a reviewer contract (lines 37-42). The advisor must prefix its response with either APPROVE or REDO, creating a discrete binary verdict that the gate evaluates deterministically.
Verdict Handling
The gate implements two distinct paths based on the advisor's verdict (lines 39-42, 50-52).
- APPROVE: The buffered executor turn replays verbatim to the client, preserving all provider events.
- REDO: The gated turn is discarded entirely (invisible to the client). The advisor's plan injects as user-feedback prefixed by
redo_feedback_prefix, and the executor re-invokes with this guidance.
Signals That Send the Executor Back
When the advisor returns REDO, the system emits multiple signals to track the rejection and guide the executor's continuation.
Prometheus Metrics
The implementation in crates/libsy/src/algorithms/advisor_gate.rs exposes operational telemetry:
switchyard_advisor_gate_discarded_turns_total(lines 13-15, 64-65): Counts rejected turns.switchyard_advisor_gate_discarded_tokens_totalwithkindlabel (lines 14-15, 84-88): Tracks wasted input/output tokens.switchyard_advisor_gate_reviews_totalwithverdict="redo"andtriggerlabels (lines 11-12, 59-61): Records review outcomes by trigger type.
Feedback Injection
The primary control signal occurs at line 351 of advisor_gate.rs, where the gate constructs the feedback message:
let feedback = format!("{}{}", self.config.redo_feedback_prefix, plan);
This string prepends to the next user message, effectively rewinding the conversation state with explicit reviewer instructions. The configuration value redo_feedback_prefix is defined in crates/switchyard-runner/src/config.rs (line 1765) and wired through crates/switchyard-runner/src/algorithm.rs. The executor receives no indication that its previous turn was discarded; it simply observes new user feedback requesting changes.
Configuration and Implementation
Enabling the Approval Loop
Configure the gate via TOML in your Switchyard deployment:
[targets.executor]
id = "small/model"
llm_client = "provider"
[targets.advisor]
id = "frontier/model"
llm_client = "provider"
[routes.gated]
id = "switchyard/gated"
type = "advisor"
executor_target = "executor"
advisor_target = "advisor"
max_reviews = 3 # allow up to three reviews per session
gate_stall_turns = 30 # trigger a mid-task review after 30 assistant turns
gate_min_tool_results = 3 # ignore early no-tool-call turns until 3 tool results appear
redo_feedback_prefix = "REVIEWER SAYS: " # prefix injected before the advisor's plan
Source: lines 81-98 of docs/routing_algorithms/advisor_gate_routing.md.
Internal REDO Handling
The Rust implementation in crates/libsy/src/algorithms/advisor_gate.rs processes verdicts:
if verdict == "REDO" {
// Discard the gated turn
self.stats.increment_discarded_turn(); // updates discarded_turns metric
self.stats.record_discarded_tokens(&turn); // updates discarded_tokens metric
// Inject the advisor's plan as user feedback
let feedback = format!("{}{}", self.config.redo_feedback_prefix, plan);
self.session.append_user_message(feedback);
// Re-invoke the executor with the new plan
self.invoke_executor().await?;
}
Source: line 351 of advisor_gate.rs.
Monitoring Review Metrics
Query the /v1/stats endpoint to observe approval loop behavior:
curl http://localhost:8000/v1/stats | jq '.advisor_gate'
Example response:
{
"reviews": {
"approve": { "total": 5, "by_trigger": { "no_tool_call": 5 } },
"redo": { "total": 2, "by_trigger": { "stall": 2 } }
},
"consult_failures": { "client_error": 1 },
"discarded": {
"turns": 2,
"tokens": { "input": 240, "output": 60 }
}
}
Source: lines 74-77 of documentation and crates/switchyard-server/src/stats/algorithms/advisor_gate.rs.
Summary
- The approval loop intercepts executor turns at terminal states (
no_tool_call,pattern) or stall thresholds (gate_stall_turns). - The advisor evaluates transcripts under a reviewer contract requiring
APPROVEorREDOprefixes. - REDO verdicts discard the executor turn entirely, inject advisor feedback via
redo_feedback_prefix, and re-invoke the executor without client visibility. - Prometheus metrics track discarded turns, token costs, and review triggers under the
advisor_gatenamespace at/v1/stats. - Budget constraints (
max_reviews, default 1) limit reviews per session, with failed consultations refunding the budget and logging toconsult_failures.
Frequently Asked Questions
What triggers the advisor gate to review an executor turn?
The gate fires on three conditions: (1) no_tool_call when the executor produces a turn without tool calls after gate_min_tool_results tool results; (2) pattern when turn text matches the gate_trigger_pattern regex; or (3) gate_stall_turns when assistant turn count exceeds the configured threshold.
How does the executor know it needs to redo work?
The executor receives no explicit "redo" command. Instead, the gate discards the rejected turn and injects the advisor's plan as a new user message prefixed by redo_feedback_prefix. The executor perceives this as additional user feedback requiring response, effectively looping it back into the task without awareness of the rejection.
What happens if the advisor fails to respond or returns malformed output?
Failed advisor consultations increment consult_failures metrics and refund the review budget (lines 58-66 of advisor_gate_routing.md). The original executor turn proceeds to the client as a fallback, ensuring that advisor unavailability does not block task completion.
Can I configure multiple reviews for a single session?
Yes. Set max_reviews in the route configuration (default 1, maximum configurable as needed). Each REDO verdict consumes one review from the budget until exhausted, at which point subsequent turns pass directly to the client without advisor consultation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →