# How Advisor Gate Routing Works in Switchyard: Executor-Reviewer Architecture Explained

> Discover how Advisor Gate Routing in Switchyard uses an executor and advisor LLM to ensure response quality through automated approval gates.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-08-22

---

**Advisor Gate Routing pairs an executor LLM with a reviewer (advisor) LLM, buffering every turn so the advisor can approve or reject responses before they reach the client, ensuring quality through automated quality gates.**

Advisor Gate Routing is a sophisticated routing algorithm implemented in the NVIDIA-NeMo/Switchyard framework that orchestrates two distinct language models to enforce quality controls on AI-generated outputs. Unlike simple routing mechanisms, this approach introduces a quality gate where every client-visible turn is inspected by a secondary reviewer model before delivery. The architecture ensures that only approved responses reach the end user, while rejected turns trigger automated regeneration cycles with corrective feedback.

## The Dual-Model Architecture: Executor and Advisor

The algorithm distinguishes between two specialized roles that operate in tandem to process user requests.

### The Executor Role (`executor_target`)

The **executor** serves as the primary workhorse model that handles every client-visible turn. According to the implementation in [`crates/libsy/src/algorithms/advisor_gate.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/advisor_gate.rs), this model executes the task, makes tool calls, and generates what the system designates as the "terminal" turn. All initial processing and tool interactions occur through this model, making it the client-facing engine of the routing pair.

### The Advisor (Reviewer) Role (`advisor_target`)

The **advisor** functions as a non-streaming quality gate that never transmits directly to the client. Its sole responsibility is reviewing the executor's buffered terminal turn and returning a binary verdict: either **APPROVE** to release the turn to the client, or **REDO** to discard the response and trigger a regeneration cycle. This model operates asynchronously behind the scenes, inspecting conversation quality through structured analysis.

## The Advisor Gate Routing Flow

The routing sequence follows a strict multi-stage pipeline that ensures comprehensive review before client delivery:

1. **Initial Execution and Buffering** – All turns are first routed to the executor. The response is immediately buffered so the gate can inspect it before any client sees the output.

2. **Trigger Evaluation** – The system monitors for review triggers using `GateTrigger` conditions. By default, this fires on the first turn **without tool calls** that also satisfies `gate_min_tool_results`, though custom regex patterns (`GateTrigger::Pattern`) or mid-task stall checkpoints can also initiate review.

3. **Consultation Request** – When triggered, the buffered turn and full session transcript are sent to the advisor via `AdvisorGate::build_consult_request`. The transcript concatenates request instructions and messages, truncated at `transcript_max_chars` to manage context window constraints.

4. **Verdict Parsing** – The advisor receives a specialized `REVIEWER_SYSTEM_PROMPT` that mandates responses begin with either `APPROVE` or `REDO`. The `parse_verdict` function processes this reply to determine the routing outcome.

5. **Output Handling** – 
   - **APPROVE** releases the buffered turn verbatim to the client
   - **REDO** discards the buffered turn, prefixes the advisor's plan with `REDO_FEEDBACK_PREFIX`, injects it as a user message, and re-invokes the executor

## Trigger Configuration and Review Logic

The system supports multiple trigger mechanisms defined in the `GateTrigger` enum and compiled into `CompiledTrigger` instances for efficient evaluation.

The default `NoToolCall` trigger activates when the executor produces a turn without tool invocations after accumulating the minimum required tool results specified by `gate_min_tool_results`. Alternatively, `Pattern` triggers allow regex-based matching for custom review conditions.

For long-running tasks, the implementation supports stall detection keyed by a hash of the first user message (`stall_key`), enabling mid-task quality checkpoints without waiting for terminal turns.

## State Management and Budget Controls

Per-session state is maintained through a `Mutex<GateState>` that tracks `ScopeState` for each conversation identified by `proxy_x_session_id`, the host's session ID, or a global instance scope. All mutations occur within short critical sections and are never held across async calls, preventing contention during high-throughput operations.

The gate enforces strict budgets to prevent infinite review loops:
- **`max_reviews`** limits the total number of advisory consultations per session
- **`MAX_FAILED_CONSULTS`** tracks consecutive advisor failures
- **`fail_open`** configuration automatically approves turns when the advisor becomes unavailable, ensuring system resilience

Expired scopes are evicted automatically to maintain bounded memory usage in the ledger.

## REDO Operations and Request Reconstruction

When the advisor returns `REDO`, the `AdvisorGate::redo` method reconstructs the conversation context by:
- Appending the discarded turn's content (or a placeholder) as an assistant message
- Injecting the advisor's corrective plan as a user message prefixed with `REDO_FEEDBACK_PREFIX`
- Dropping the exact-replay cache to ensure the executor processes the new feedback rather than repeating the rejected output

This cycle repeats until the advisor approves the turn or the review budget is exhausted.

## Observability and Performance Metrics

The algorithm exposes comprehensive Prometheus metrics through [`crates/switchyard-server/src/stats/algorithms/advisor_gate.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/advisor_gate.rs), including:
- `switchyard_advisor_gate_reviews_total` – total review operations
- `consult_failures_total` – failed advisory consultations
- `discarded_turns_total` – turns rejected and regenerated

The `AdvisorGateStatsSnapshot` provides human-readable projections of these counters for operational monitoring. Each consultation through `AdvisorGate::consult` records latency metrics and emits audit events via `emit_review_audit` after invoking the advisor through `driver.call_model`.

## Implementation Examples

The following Python client demonstrates invoking a gated route:

```python
import switchyard

client = switchyard.SwitchyardClient()
response = client.chat(
    model="switchyard/gated",
    messages=[{"role": "user", "content": "Write a Python function to add two numbers"}],
)
print(response["content"])

```

For direct Rust integration using the `libsy` crate:

```rust
use switchyard_libsy::algorithms::advisor_gate::{
    AdvisorGate, AdvisorGateConfig, GateTrigger,
};

let config = AdvisorGateConfig {
    executor_target: "small/model".into(),
    advisor_target: "frontier/model".into(),
    gate_trigger: GateTrigger::NoToolCall,
    max_reviews: 3,
    ..Default::default()
};

let gate = AdvisorGate::new(
    "small/model".into(),
    "frontier/model".into(),
    config,
).unwrap();

```

## Summary

- **Advisor Gate Routing** implements a dual-model architecture where an executor generates responses and an advisor reviews them before client delivery
- **Quality gates** enforce standards through binary verdicts (APPROVE/REDO) parsed from structured advisor outputs
- **Trigger mechanisms** include tool-call absence patterns, regex matching, and stall checkpoints compiled via `CompiledTrigger`
- **Budget controls** prevent infinite loops through per-session `max_reviews` limits and `fail_open` resilience options
- **State management** uses `Mutex<GateState>` with short critical sections to track `ScopeState` across sessions identified by `proxy_x_session_id`
- **Observability** is built-in through Prometheus metrics like `switchyard_advisor_gate_reviews_total` and latency-tracking audits in `AdvisorGateStatsSnapshot`
- **Source implementation** resides in [`crates/libsy/src/algorithms/advisor_gate.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/advisor_gate.rs) with documentation in [`docs/routing_algorithms/advisor_gate_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/advisor_gate_routing.md)

## Frequently Asked Questions

### What happens when the advisor fails to respond during a consultation?

If the advisor fails to return a valid verdict and `fail_open` is set to `true` in the configuration, the gate automatically approves the buffered turn to prevent service interruption. The system tracks failed consultations via `consult_failures_total` and evicts problematic scopes to maintain system stability.

### How does the system prevent infinite REDO loops?

Each session maintains a review budget specified by `max_reviews` in `AdvisorGateConfig`. Once this limit is reached, the gate stops consulting the advisor and releases subsequent turns directly to the client. This ensures that quality control does not result in unbounded recursion or excessive latency.

### Can developers customize when reviews are triggered?

Yes, the `GateTrigger` enum supports multiple activation methods including `NoToolCall` (default), regex-based `Pattern` matching, and stall checkpoint detection. Developers can configure these triggers during `AdvisorGate` initialization to suit specific quality assurance requirements for different task types.

### What is the difference between the executor and advisor models?

The **executor** (`executor_target`) is the client-facing model that performs work, makes tool calls, and generates terminal turns visible to users. The **advisor** (`advisor_target`) is a non-streaming reviewer that operates behind the scenes, exclusively outputting `APPROVE` or `REDO` verdicts based on `REVIEWER_SYSTEM_PROMPT` instructions without ever transmitting directly to the client.