# Escalation Router Routing in NVIDIA-NeMo Switchyard: Automatic Tier Escalation Strategy

> Explore Escalation Router Routing in NVIDIA-NeMo Switchyard. Automatically elevate conversations to strong models using LLM judges for efficient cost management.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-08-22

---

**Escalation Router Routing is a trajectory-judge based strategy that starts every conversation on a low-cost weak model and automatically promotes the session to a high-capability strong model when a dedicated LLM judge returns consecutive escalation verdicts.**

Escalation Router Routing enables cost-efficient LLM deployments by deferring expensive inference calls until a cheaper model proves insufficient. Implemented in the NVIDIA-NeMo Switchyard repository, this routing algorithm uses a tiered approach where sessions begin on economical weak models and escalate to premium strong models only when trajectory analysis indicates persistent failure. The strategy balances operational costs against response quality through a configurable streak-based latching mechanism defined in [`crates/libsy/src/algorithms/util/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs).

## How Escalation Router Routing Works

The routing flow operates through three distinct logical phases: weak-tier execution, trajectory judgment, and streak-based latching. Each phase is designed to minimize expensive model calls while maintaining conversation quality through continuous monitoring.

### Weak-Tier Execution and Reply Buffering

When a request enters an escalation route, Switchyard immediately forwards it to the **weak target** configured in `weak_target`. The response from this economical model is buffered in memory but not immediately returned to the client. This buffering allows the system to evaluate the quality of the weak response before committing to serving it, effectively creating a "look-before-you-leap" mechanism that prevents serving low-quality outputs when the model struggles.

### Trajectory Judging with EscalationJudgeConfig

The buffered reply is appended to the session transcript and sent to the **judge model** specified by `classifier_target`. Defined in [`EscalationJudgeConfig`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs#L48-L58), this configuration structures how the judge analyzes the conversation trajectory. The judge examines the full turn context—including the weak model's reply—and returns a boolean `escalate` verdict defined in [`EscalationVerdict`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs#L90-L96).

The verdict feeds into the **EscalationPolicy** (lines 24-40 in [`escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/escalation.rs)), which maps the boolean result to a `Classification` that the routing engine consumes. When the verdict is `true`, the policy selects the **strong** model (labeled `capable`); otherwise, it selects the **weak** model (labeled `efficient`).

### Streak-Based Latching Mechanism

Escalation Router Routing implements a **per-session counter** that tracks consecutive `escalate` verdicts. When this counter reaches the `confirmations` threshold—defaulting to `2` as specified in the configuration—the session becomes **latched**. Once latched, the buffered weak reply is discarded, and the current turn is re-routed to the strong model (`strong_target`). All subsequent turns bypass the judge entirely and flow directly to the strong model until the session terminates.

This latching behavior prevents thrashing between models and eliminates judge inference costs for the remainder of the conversation. The `recent_turn_window` (default 28) and `window_message_chars` (default 500) parameters in `EscalationJudgeConfig` control how much historical context the judge evaluates when making escalation decisions.

## Cost Analysis and Call Patterns

Understanding the token economics of Escalation Router Routing requires analyzing the call patterns at each decision point:

| Step | Action | Cost Implication |
|------|--------|------------------|
| 1 | Call weak model and buffer reply | 1 weak model call |
| 2 | Call judge on buffered reply | +1 judge model call |
| 3a | Streak < `confirmations`: Serve weak reply | No strong call (total: 2 calls) |
| 3b | Streak ≥ `confirmations`: Discard weak, call strong | +1 strong call (total: 3 calls for escalation turn) |

An escalated turn therefore incurs three sequential LLM calls: weak inference, judge evaluation, and strong inference. However, once latched, the session incurs only single strong-model calls for all remaining turns, eliminating the overhead of the judgment phase.

## Error Handling and Fail-Open Safety

Switchyard implements defensive programming for the escalation pathway. If the judge times out, errors, or returns an unparsable payload, the route **fails open**—the buffered weak reply is immediately served to the client, and the escalation streak remains unchanged. This prevents transient judge failures from causing unintended latches to expensive models or dropping user requests. The fail-open logic ensures that service degradation defaults to the cost-efficient weak tier rather than erroring or forcing expensive fallbacks.

## Configuring Escalation Router Routing

Escalation routes are declared in Switchyard's TOML configuration using `type = "llm_classifier"` with `mode = "escalation"`. The configuration requires three distinct targets: the judge (`classifier_target`), the capable model (`strong_target`), and the efficient model (`weak_target`).

### Minimal TOML Configuration

```toml
[schema]
version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.judge]
id = "google/gemini-3.5-flash"
llm_client = "openrouter"

[targets.strong]
id = "anthropic/claude-opus-4.7"
llm_client = "openrouter"

[targets.weak]
id = "moonshotai/kimi-k2.6"
llm_client = "openrouter"

[routes.agent]
id = "agent"
type = "llm_classifier"
mode = "escalation"
classifier_target = "judge"
strong_target = "strong"
weak_target = "weak"
prompt = "Judge whether the weak model is stuck. Return the required structured verdict."
escalation = { confirmations = 2, recent_turn_window = 28, window_message_chars = 500 }

```

### HTTP Client Usage

Clients must include the `x-switchyard-session-id` header to maintain session state for streak tracking and latching:

```bash
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "x-switchyard-session-id: demo-session" \
  -d '{"model":"agent","messages":[{"role":"user","content":"Write a short poem about clouds"}]}'

```

### Python SDK Integration

```python
import switchyard

client = switchyard.Client(base_url="http://localhost:4000")
response = client.chat_completions(
    model="agent",
    messages=[{"role": "user", "content": "Explain quantum tunneling"}],
    headers={"x-switchyard-session-id": "demo-session"},
)
print(response["choices"][0]["message"]["content"])

```

## Key Source Files

The Escalation Router Routing implementation spans multiple crates in the Switchyard codebase:

- **[[`crates/libsy/src/algorithms/util/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs)** — Core implementation containing `EscalationJudgeConfig` (lines 48-58), `EscalationVerdict` (lines 90-96), `EscalationPolicy` (lines 24-40), and the `build_judge` factory function.

- **[[`crates/libsy/src/algorithms/llm_class.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/llm_class.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/llm_class.rs)** — Registration point for the escalation classifier that wires the judge logic into the routing engine.

- **[[`docs/routing_algorithms/escalation_router_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/escalation_router_routing.md)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/escalation_router_routing.md)** — User-facing documentation covering configuration schema and examples.

- **[[`crates/switchyard-py/src/libsy_bindings.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/libsy_bindings.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/libsy_bindings.rs)** — Python bindings exposing `EscalationClassifierConfig` to the Switchyard Python package.

## Summary

- **Escalation Router Routing** starts all sessions on economical weak models and promotes them to strong models only when a judge detects persistent failure.
- The **streak-based latching** mechanism requires `confirmations` consecutive escalation verdicts (default: 2) before permanently upgrading a session to the strong tier.
- **Fail-open error handling** ensures judge timeouts or parsing errors default to serving the weak model reply rather than dropping requests or forcing expensive fallbacks.
- Configuration requires three targets—`classifier_target`, `strong_target`, and `weak_target`—specified via TOML with `mode = "escalation"`.
- Escalated turns incur three LLM calls (weak, judge, strong), but latched sessions eliminate judge overhead for all subsequent turns.

## Frequently Asked Questions

### What triggers an escalation in Escalation Router Routing?

An escalation triggers when the judge model returns `escalate = true` for a number of consecutive turns equal to the `confirmations` threshold, which defaults to 2 but is configurable per route. The judge evaluates the conversation trajectory within the `recent_turn_window` (default 28 turns) to determine if the weak model is struggling.

### How does the fail-open mechanism protect against judge failures?

If the judge times out, returns malformed JSON, or encounters a runtime error, the route **fails open** by immediately serving the buffered weak model reply and preserving the current escalation streak count. This prevents transient judge unavailability from causing service disruptions or accidentally latching sessions to expensive models.

### What is the cost implication when a session escalates?

The escalation turn incurs three sequential LLM calls: one to the weak model, one to the judge, and one to the strong model. However, once latched, the session bypasses the judge entirely, sending all subsequent turns directly to the strong model at standard single-call pricing, effectively amortizing the judgment overhead over the session lifetime.

### How do you configure the confirmation threshold for escalation?

Set the `confirmations` value in the `escalation` table of your route configuration, as implemented in [`EscalationJudgeConfig`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs#L48-L58). For example, `escalation = { confirmations = 3 }` requires three consecutive positive verdicts before latching to the strong model, reducing false positives at the cost of potentially delayed escalation.