# Switchyard Escalation Router: Trigger Conditions and Mode Differences Explained

> Understand Switchyard escalation router trigger conditions and mode differences. Learn how escalation differs from standard LLM classifiers for improved conversational AI.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-13

---

**Switchyard’s escalation router promotes conversations to a strong model only after a judge returns consecutive `escalate` verdicts matching a configurable confirmation threshold, while standard LLM classifiers predict difficulty before any model execution.**

The NVIDIA-NeMo/Switchyard framework provides intelligent request routing for multi-turn LLM workloads. Its **escalation router** dynamically monitors conversation health at runtime and triggers tier promotion only when the weak model demonstrably fails. This contrasts sharply with the standard `llm_classifier` mode, which predicts query difficulty upfront and commits to a tier before execution begins.

## How the Escalation Mechanism Works

The escalation router in [`crates/libsy/src/algorithms/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/escalation.rs) implements a **runtime monitoring** pattern. Every conversation begins on a cost-effective weak model. After each turn, the router invokes a **judge** model to evaluate whether the weak model is stuck.

### Judge Verdict Processing

The judge returns a structured verdict of either `escalate` or `decline`. The router maintains a **per-session streak counter** that increments on each `escalate` verdict and resets to zero on any `decline`. Only when this streak reaches the `confirmations` threshold (typically configured to `2`) does the router **latch** the session to the strong model, bypassing further judging for subsequent turns.

If the judge times out, errors, or returns an unparsable verdict, the router **fails open**: it serves the buffered weak-model reply without incrementing the streak. This guarantees that a faulty judge never forces an unintended strong-model latch.

### Session State Requirements

Unlike standard routing modes, escalation requires a persistent session identifier passed via the `x-switchyard-session-id` header. The router stores streak counters per session ID, enabling multi-turn tracking across distributed requests.

## Conditions That Trigger Escalation

The judge evaluates conversation history—controlled by `recent_turn_window` (default 28 turns) and `window_message_chars` (default 500 characters)—to identify specific failure patterns that warrant escalation:

- **Repeated errors**: The weak model generates nonsensical outputs or failing code repeatedly within the monitored window.
- **Loops or drift**: The conversation enters a repetitive cycle or drifts away from the intended task objective.
- **Sustained trouble**: Any pattern of problematic behavior identified by the judge's prompt logic as requiring rescue intervention.

The escalation trigger is **not** based on single-turn failures but on sustained dysfunction across multiple consecutive evaluations.

## Escalation Mode vs. Standard LLM Classifier

The `mode = "escalation"` setting fundamentally changes the routing decision timeline compared to the default classifier behavior:

| Aspect | `mode = "escalation"` | Standard `llm_classifier` |
|--------|----------------------|---------------------------|
| **Decision timing** | After weak model generates output and judge evaluates it | Before any model call, based on request text analysis |
| **Trigger mechanism** | Consecutive `escalate` verdicts matching `confirmations` threshold | Single prediction (e.g., `high`, `low`) from classifier model |
| **Cost structure** | One weak call + one judge call per unlatched turn; strong call only post-escalation | One classifier call upfront, then only the selected tier |
| **State management** | Requires session persistence (`x-switchyard-session-id`) | Stateless; each request independent |
| **Optimal use case** | Multi-turn agent workloads where failure emerges during execution | One-shot requests where difficulty is predictable beforehand |

## Configuration and Implementation

### Escalation Router Configuration

Define an escalation route in your TOML configuration by setting `type = "llm_classifier"` with `mode = "escalation"`:

```toml
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.judge]
id = "google/gemini-3.5-flash"
llm_client = "openrouter"

[targets.strong]
id = "anthropic/claude-opus-4.7"
llm_client = "openrouter"

[targets.weak]
id = "moonshotai/kimi-k2.6"
llm_client = "openrouter"

[routes.agent]
id = "agent"
type = "llm_classifier"
mode = "escalation"
classifier_target = "judge"
strong_target = "strong"
weak_target = "weak"
prompt = "Judge whether the weak model is stuck. Return the required structured verdict."
escalation = { confirmations = 2, recent_turn_window = 28, window_message_chars = 500 }

```

The `escalation` table configures the streak threshold and history window sizes. See [`docs/routing_algorithms/escalation_router_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/escalation_router_routing.md) for complete schema documentation.

### Making Escalation-Enabled Requests

Include the session header to enable streak tracking across turns:

```bash
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "x-switchyard-session-id: demo-session" \
  -d '{
        "model":"agent",
        "messages":[{"role":"user","content":"Debug this recursive Python function"}]
      }'

```

The first turn executes on the weak model. If the judge returns `escalate` on two consecutive turns (per `confirmations = 2`), the third turn routes directly to the strong model without judge overhead.

### Standard Classifier Comparison

A standard configuration omits the escalation mode, making tier decisions before execution:

```toml
[routes.agent]
id = "agent"
type = "llm_classifier"

# Defaults to capability prediction mode

classifier_target = "classifier"
strong_target = "strong"
weak_target = "weak"

```

Here, the router predicts query difficulty using the `classifier` target before invoking either tier, eliminating runtime judging overhead but losing the ability to rescue conversations that degrade mid-stream.

## Summary

- Switchyard's escalation router monitors conversation health post-execution using a judge model that returns `escalate` or `decline` verdicts.
- Escalation triggers only after consecutive `escalate` verdicts match the configured `confirmations` threshold, requiring persistent session IDs.
- Trigger conditions include repeated errors, conversational loops, and task drift detected within the configured history window.
- Unlike standard LLM classifiers that predict difficulty before execution, escalation mode acts as a runtime rescue mechanism for multi-turn agent workloads.
- The implementation in [`crates/libsy/src/algorithms/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/escalation.rs) handles streak counting, session latching, and fail-open safety for judge failures.

## Frequently Asked Questions

### What happens if the judge model fails or returns an invalid verdict?

The router implements **fail-open** safety. If the judge times out, errors, or returns unparsable structured output, the router serves the weak model's buffered response and does not increment the escalation streak. This prevents judge malfunctions from forcing unnecessary strong-model usage.

### Can I adjust how much conversation history the judge evaluates?

Yes. The `recent_turn_window` parameter (default 28) controls how many prior turns the judge examines, while `window_message_chars` (default 500) limits the total character count of history provided to the judge. Configure these in the `escalation` table of your route definition to balance context richness against judge latency and cost.

### Why does escalation mode require session headers while standard classification does not?

Escalation mode maintains a **streak counter** per conversation to track consecutive `escalate` verdicts. The router stores this state keyed by the `x-switchyard-session-id` header value. Standard classifiers make independent predictions per request without tracking conversational state, eliminating the need for session persistence.

### How does the cost profile differ between escalation and standard classification?

Escalation mode incurs **one weak model call plus one judge call** for every turn until the escalation threshold is reached and the session latches to the strong model. Standard classification incurs **one classifier call upfront**, followed by only the selected tier's calls. Escalation optimizes for scenarios where most conversations succeed on the weak model, while standard classification optimizes for predictable per-request routing without runtime judging overhead.