# How the Escalation Router Optimizes Costs in Switchyard

> Discover how the Escalation router in NVIDIA Switchyard cuts LLM serving costs by intelligently routing conversations to affordable models, saving money while maintaining quality.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: performance
- Published: 2026-08-21

---

**The Escalation router reduces LLM serving expenses by starting every conversation on a cheap "weak" model and only promoting the session to an expensive "strong" model when a built-in trajectory judge determines that the weak model is stuck or looping.**

NVIDIA-NeMo/Switchyard is an open-source LLM routing framework that implements intelligent traffic distribution across multiple model tiers. The Escalation router specifically targets cost optimization through a two-stage processing pipeline that minimizes expensive API calls while maintaining output quality.

## Two-Stage Architecture for Cost Reduction

The Escalation router implements a **cold-start optimization pattern** that defers expensive computation until it is provably necessary. This architecture is hardcoded into the routing logic found in [`crates/libsy/src/algorithms/util/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs).

### Initial Cheap Processing on Weak Models

For each turn of an *unlatched* session, Switchyard first invokes the configured **weak target** (for example, `moonshotai/kimi-k2.6`) and buffers its reply. This incurs only the token cost of the weak model. The router holds this response in memory while it awaits the judge's verdict, ensuring that users receive a response even if the escalation check fails.

The weak model handles the initial reasoning, code generation, or chat completion, while the system maintains a **transcript summary** that respects bounds like `MAX_REQUEST_CHARS` and `SYSTEM_CHARS` to prevent token bloat during the judging phase.

### Judge-Driven Escalation Decisions

After the weak reply is appended to the transcript, the router invokes the **trajectory-judge** classifier ( implemented as `EscalationJudge` in the Rust source) to evaluate whether the weak model is making progress. The judge processes a condensed transcript that caps per-message length and total characters to keep its token usage bounded.

The judge returns a structured verdict (`escalate: true/false`) that determines whether the session requires the strong model's capabilities. This evaluation happens transparently between the weak model's response and the final output to the user.

## The Confirmation Threshold as a Cost Dial

The router maintains a **streak** of consecutive "escalate" verdicts. Only when the streak reaches the `confirmations` threshold does the session *latch*, causing Switchyard to route future turns directly to the strong model (such as `anthropic/claude-opus-4.7`).

### Tuning the confirmations Parameter

The `confirmations` setting acts as a direct cost-control lever:

- **Setting `confirmations = 1`** causes immediate latching to the strong tier upon the first positive judge verdict, increasing strong-model usage and overall cost.
- **The default `confirmations = 2`** keeps the session on the weak tier for at least one judged turn, saving costs while still allowing rapid escalation when the weak model demonstrably fails.

You configure this behavior in your route definition:

```toml
[routes.agent]
id = "agent"
type = "llm_classifier"
mode = "escalation"
classifier_target = "judge"
strong_target = "strong"
weak_target = "weak"

# Tune the cost-dial here; the default is 2 confirmations.

escalation = { confirmations = 2, recent_turn_window = 28, window_message_chars = 500 }

```

### Fail-Open Safety Mechanisms

If the judge times out or returns an unparsable verdict, Switchyard **fails open**: it serves the buffered weak reply and does **not** latch to the strong tier. This prevents unnecessary strong-model calls caused by judge errors or infrastructure issues, ensuring that transient failures do not inflate your inference costs.

## Implementation in the Switchyard Codebase

The core escalation logic resides in [`crates/libsy/src/algorithms/util/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs). This file implements the `EscalationJudge` struct, which handles:

- Transcript summarization to fit within token limits
- Verdict conversion from the judge model's structured output
- Streak counting and session state management

The algorithm exposes a clean interface to the Python integration layer via [`switchyard/libsy/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/__init__.py), allowing the Escalation router to function as a first-class routing primitive within the Switchyard ecosystem.

## Configuration and Observability

To utilize the Escalation router, you define three distinct targets in your TOML configuration: the weak model, the strong model, and the judge model.

```toml
[targets.judge]
id = "google/gemini-3.5-flash"
llm_client = "openrouter"

[targets.strong]
id = "anthropic/claude-opus-4.7"
llm_client = "openrouter"

[targets.weak]
id = "moonshotai/kimi-k2.6"
llm_client = "openrouter"

```

Send requests with a persistent session ID to maintain the escalation streak across turns:

```bash
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "x-switchyard-session-id: demo-session" \
  -d '{
        "model":"agent",
        "messages":[{"role":"user","content":"Write a Python script that backs up a folder"}]
      }'

```

Monitor per-tier token usage via the stats endpoint to verify cost savings:

```bash
curl -s http://localhost:4000/v1/stats | jq '.models'

```

This output lists token counts and call totals for `weak`, `strong`, and the judge (`classifier`) models, allowing you to calculate the exact cost differential between the Escalation router and a naive single-model deployment.

## Summary

- The Escalation router implements a **two-stage pipeline** (weak model → judge → conditional strong model) to minimize expensive API calls.
- The `confirmations` parameter acts as a tunable **cost dial**, trading off latency for expense reduction.
- **Fail-open behavior** ensures that judge errors never trigger unnecessary strong-model invocations.
- All routing logic is implemented in [`crates/libsy/src/algorithms/util/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs), specifically within the `EscalationJudge` struct.
- The `/v1/stats` endpoint provides observability into per-tier usage, enabling operators to verify cost optimization in production.

## Frequently Asked Questions

### What models should I use for the weak and strong targets?

Select a **weak model** that handles 70-80% of routine queries adequately (such as `moonshotai/kimi-k2.6` or similar cost-efficient options) and a **strong model** that excels at complex reasoning or code generation (such as `anthropic/claude-opus-4.7`). The **judge model** should be fast and inexpensive (like `google/gemini-3.5-flash`) since it runs on every turn of an unlatched session.

### How does the trajectory judge determine if a model is stuck?

The `EscalationJudge` in [`crates/libsy/src/algorithms/util/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs) analyzes a condensed transcript of recent turns, checking for repetitive patterns, lack of progress toward the stated goal, or circular reasoning. It returns a boolean `escalate` verdict based on whether the weak model's trajectory indicates it is unlikely to resolve the query successfully.

### What happens if the judge fails or times out?

Switchyard implements **fail-open semantics**. If the judge request times out or returns malformed JSON, the router serves the buffered weak model response and does not increment the escalation streak or latch the session to the strong tier. This prevents infrastructure glitches from triggering expensive strong-model calls.

### How can I monitor the cost savings from escalation routing?

Query the `/v1/stats` endpoint to retrieve token counts and call volumes for each model tier (weak, strong, and classifier). By comparing the ratio of weak-model calls to strong-model calls and calculating the respective API costs, you can derive the exact cost reduction achieved versus running all traffic through the expensive strong model exclusively.