How Advisor Gate Routing Works in Switchyard: Executor-Reviewer Architecture Explained

Advisor Gate Routing pairs an executor LLM with a reviewer (advisor) LLM, buffering every turn so the advisor can approve or reject responses before they reach the client, ensuring quality through automated quality gates.

Advisor Gate Routing is a sophisticated routing algorithm implemented in the NVIDIA-NeMo/Switchyard framework that orchestrates two distinct language models to enforce quality controls on AI-generated outputs. Unlike simple routing mechanisms, this approach introduces a quality gate where every client-visible turn is inspected by a secondary reviewer model before delivery. The architecture ensures that only approved responses reach the end user, while rejected turns trigger automated regeneration cycles with corrective feedback.

The Dual-Model Architecture: Executor and Advisor

The algorithm distinguishes between two specialized roles that operate in tandem to process user requests.

The Executor Role (executor_target)

The executor serves as the primary workhorse model that handles every client-visible turn. According to the implementation in crates/libsy/src/algorithms/advisor_gate.rs, this model executes the task, makes tool calls, and generates what the system designates as the "terminal" turn. All initial processing and tool interactions occur through this model, making it the client-facing engine of the routing pair.

The Advisor (Reviewer) Role (advisor_target)

The advisor functions as a non-streaming quality gate that never transmits directly to the client. Its sole responsibility is reviewing the executor's buffered terminal turn and returning a binary verdict: either APPROVE to release the turn to the client, or REDO to discard the response and trigger a regeneration cycle. This model operates asynchronously behind the scenes, inspecting conversation quality through structured analysis.

The Advisor Gate Routing Flow

The routing sequence follows a strict multi-stage pipeline that ensures comprehensive review before client delivery:

  1. Initial Execution and Buffering – All turns are first routed to the executor. The response is immediately buffered so the gate can inspect it before any client sees the output.

  2. Trigger Evaluation – The system monitors for review triggers using GateTrigger conditions. By default, this fires on the first turn without tool calls that also satisfies gate_min_tool_results, though custom regex patterns (GateTrigger::Pattern) or mid-task stall checkpoints can also initiate review.

  3. Consultation Request – When triggered, the buffered turn and full session transcript are sent to the advisor via AdvisorGate::build_consult_request. The transcript concatenates request instructions and messages, truncated at transcript_max_chars to manage context window constraints.

  4. Verdict Parsing – The advisor receives a specialized REVIEWER_SYSTEM_PROMPT that mandates responses begin with either APPROVE or REDO. The parse_verdict function processes this reply to determine the routing outcome.

  5. Output Handling –

    • APPROVE releases the buffered turn verbatim to the client
    • REDO discards the buffered turn, prefixes the advisor's plan with REDO_FEEDBACK_PREFIX, injects it as a user message, and re-invokes the executor

Trigger Configuration and Review Logic

The system supports multiple trigger mechanisms defined in the GateTrigger enum and compiled into CompiledTrigger instances for efficient evaluation.

The default NoToolCall trigger activates when the executor produces a turn without tool invocations after accumulating the minimum required tool results specified by gate_min_tool_results. Alternatively, Pattern triggers allow regex-based matching for custom review conditions.

For long-running tasks, the implementation supports stall detection keyed by a hash of the first user message (stall_key), enabling mid-task quality checkpoints without waiting for terminal turns.

State Management and Budget Controls

Per-session state is maintained through a Mutex<GateState> that tracks ScopeState for each conversation identified by proxy_x_session_id, the host's session ID, or a global instance scope. All mutations occur within short critical sections and are never held across async calls, preventing contention during high-throughput operations.

The gate enforces strict budgets to prevent infinite review loops:

  • max_reviews limits the total number of advisory consultations per session
  • MAX_FAILED_CONSULTS tracks consecutive advisor failures
  • fail_open configuration automatically approves turns when the advisor becomes unavailable, ensuring system resilience

Expired scopes are evicted automatically to maintain bounded memory usage in the ledger.

REDO Operations and Request Reconstruction

When the advisor returns REDO, the AdvisorGate::redo method reconstructs the conversation context by:

  • Appending the discarded turn's content (or a placeholder) as an assistant message
  • Injecting the advisor's corrective plan as a user message prefixed with REDO_FEEDBACK_PREFIX
  • Dropping the exact-replay cache to ensure the executor processes the new feedback rather than repeating the rejected output

This cycle repeats until the advisor approves the turn or the review budget is exhausted.

Observability and Performance Metrics

The algorithm exposes comprehensive Prometheus metrics through crates/switchyard-server/src/stats/algorithms/advisor_gate.rs, including:

  • switchyard_advisor_gate_reviews_total – total review operations
  • consult_failures_total – failed advisory consultations
  • discarded_turns_total – turns rejected and regenerated

The AdvisorGateStatsSnapshot provides human-readable projections of these counters for operational monitoring. Each consultation through AdvisorGate::consult records latency metrics and emits audit events via emit_review_audit after invoking the advisor through driver.call_model.

Implementation Examples

The following Python client demonstrates invoking a gated route:

import switchyard

client = switchyard.SwitchyardClient()
response = client.chat(
    model="switchyard/gated",
    messages=[{"role": "user", "content": "Write a Python function to add two numbers"}],
)
print(response["content"])

For direct Rust integration using the libsy crate:

use switchyard_libsy::algorithms::advisor_gate::{
    AdvisorGate, AdvisorGateConfig, GateTrigger,
};

let config = AdvisorGateConfig {
    executor_target: "small/model".into(),
    advisor_target: "frontier/model".into(),
    gate_trigger: GateTrigger::NoToolCall,
    max_reviews: 3,
    ..Default::default()
};

let gate = AdvisorGate::new(
    "small/model".into(),
    "frontier/model".into(),
    config,
).unwrap();

Summary

  • Advisor Gate Routing implements a dual-model architecture where an executor generates responses and an advisor reviews them before client delivery
  • Quality gates enforce standards through binary verdicts (APPROVE/REDO) parsed from structured advisor outputs
  • Trigger mechanisms include tool-call absence patterns, regex matching, and stall checkpoints compiled via CompiledTrigger
  • Budget controls prevent infinite loops through per-session max_reviews limits and fail_open resilience options
  • State management uses Mutex<GateState> with short critical sections to track ScopeState across sessions identified by proxy_x_session_id
  • Observability is built-in through Prometheus metrics like switchyard_advisor_gate_reviews_total and latency-tracking audits in AdvisorGateStatsSnapshot
  • Source implementation resides in crates/libsy/src/algorithms/advisor_gate.rs with documentation in docs/routing_algorithms/advisor_gate_routing.md

Frequently Asked Questions

What happens when the advisor fails to respond during a consultation?

If the advisor fails to return a valid verdict and fail_open is set to true in the configuration, the gate automatically approves the buffered turn to prevent service interruption. The system tracks failed consultations via consult_failures_total and evicts problematic scopes to maintain system stability.

How does the system prevent infinite REDO loops?

Each session maintains a review budget specified by max_reviews in AdvisorGateConfig. Once this limit is reached, the gate stops consulting the advisor and releases subsequent turns directly to the client. This ensures that quality control does not result in unbounded recursion or excessive latency.

Can developers customize when reviews are triggered?

Yes, the GateTrigger enum supports multiple activation methods including NoToolCall (default), regex-based Pattern matching, and stall checkpoint detection. Developers can configure these triggers during AdvisorGate initialization to suit specific quality assurance requirements for different task types.

What is the difference between the executor and advisor models?

The executor (executor_target) is the client-facing model that performs work, makes tool calls, and generates terminal turns visible to users. The advisor (advisor_target) is a non-streaming reviewer that operates behind the scenes, exclusively outputting APPROVE or REDO verdicts based on REVIEWER_SYSTEM_PROMPT instructions without ever transmitting directly to the client.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →