How the Escalation Router Optimizes Costs in Switchyard

The Escalation router reduces LLM serving expenses by starting every conversation on a cheap "weak" model and only promoting the session to an expensive "strong" model when a built-in trajectory judge determines that the weak model is stuck or looping.

NVIDIA-NeMo/Switchyard is an open-source LLM routing framework that implements intelligent traffic distribution across multiple model tiers. The Escalation router specifically targets cost optimization through a two-stage processing pipeline that minimizes expensive API calls while maintaining output quality.

Two-Stage Architecture for Cost Reduction

The Escalation router implements a cold-start optimization pattern that defers expensive computation until it is provably necessary. This architecture is hardcoded into the routing logic found in crates/libsy/src/algorithms/util/escalation.rs.

Initial Cheap Processing on Weak Models

For each turn of an unlatched session, Switchyard first invokes the configured weak target (for example, moonshotai/kimi-k2.6) and buffers its reply. This incurs only the token cost of the weak model. The router holds this response in memory while it awaits the judge's verdict, ensuring that users receive a response even if the escalation check fails.

The weak model handles the initial reasoning, code generation, or chat completion, while the system maintains a transcript summary that respects bounds like MAX_REQUEST_CHARS and SYSTEM_CHARS to prevent token bloat during the judging phase.

Judge-Driven Escalation Decisions

After the weak reply is appended to the transcript, the router invokes the trajectory-judge classifier ( implemented as EscalationJudge in the Rust source) to evaluate whether the weak model is making progress. The judge processes a condensed transcript that caps per-message length and total characters to keep its token usage bounded.

The judge returns a structured verdict (escalate: true/false) that determines whether the session requires the strong model's capabilities. This evaluation happens transparently between the weak model's response and the final output to the user.

The Confirmation Threshold as a Cost Dial

The router maintains a streak of consecutive "escalate" verdicts. Only when the streak reaches the confirmations threshold does the session latch, causing Switchyard to route future turns directly to the strong model (such as anthropic/claude-opus-4.7).

Tuning the confirmations Parameter

The confirmations setting acts as a direct cost-control lever:

  • Setting confirmations = 1 causes immediate latching to the strong tier upon the first positive judge verdict, increasing strong-model usage and overall cost.
  • The default confirmations = 2 keeps the session on the weak tier for at least one judged turn, saving costs while still allowing rapid escalation when the weak model demonstrably fails.

You configure this behavior in your route definition:

[routes.agent]
id = "agent"
type = "llm_classifier"
mode = "escalation"
classifier_target = "judge"
strong_target = "strong"
weak_target = "weak"

# Tune the cost-dial here; the default is 2 confirmations.

escalation = { confirmations = 2, recent_turn_window = 28, window_message_chars = 500 }

Fail-Open Safety Mechanisms

If the judge times out or returns an unparsable verdict, Switchyard fails open: it serves the buffered weak reply and does not latch to the strong tier. This prevents unnecessary strong-model calls caused by judge errors or infrastructure issues, ensuring that transient failures do not inflate your inference costs.

Implementation in the Switchyard Codebase

The core escalation logic resides in crates/libsy/src/algorithms/util/escalation.rs. This file implements the EscalationJudge struct, which handles:

  • Transcript summarization to fit within token limits
  • Verdict conversion from the judge model's structured output
  • Streak counting and session state management

The algorithm exposes a clean interface to the Python integration layer via switchyard/libsy/__init__.py, allowing the Escalation router to function as a first-class routing primitive within the Switchyard ecosystem.

Configuration and Observability

To utilize the Escalation router, you define three distinct targets in your TOML configuration: the weak model, the strong model, and the judge model.

[targets.judge]
id = "google/gemini-3.5-flash"
llm_client = "openrouter"

[targets.strong]
id = "anthropic/claude-opus-4.7"
llm_client = "openrouter"

[targets.weak]
id = "moonshotai/kimi-k2.6"
llm_client = "openrouter"

Send requests with a persistent session ID to maintain the escalation streak across turns:

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "x-switchyard-session-id: demo-session" \
  -d '{
        "model":"agent",
        "messages":[{"role":"user","content":"Write a Python script that backs up a folder"}]
      }'

Monitor per-tier token usage via the stats endpoint to verify cost savings:

curl -s http://localhost:4000/v1/stats | jq '.models'

This output lists token counts and call totals for weak, strong, and the judge (classifier) models, allowing you to calculate the exact cost differential between the Escalation router and a naive single-model deployment.

Summary

  • The Escalation router implements a two-stage pipeline (weak model → judge → conditional strong model) to minimize expensive API calls.
  • The confirmations parameter acts as a tunable cost dial, trading off latency for expense reduction.
  • Fail-open behavior ensures that judge errors never trigger unnecessary strong-model invocations.
  • All routing logic is implemented in crates/libsy/src/algorithms/util/escalation.rs, specifically within the EscalationJudge struct.
  • The /v1/stats endpoint provides observability into per-tier usage, enabling operators to verify cost optimization in production.

Frequently Asked Questions

What models should I use for the weak and strong targets?

Select a weak model that handles 70-80% of routine queries adequately (such as moonshotai/kimi-k2.6 or similar cost-efficient options) and a strong model that excels at complex reasoning or code generation (such as anthropic/claude-opus-4.7). The judge model should be fast and inexpensive (like google/gemini-3.5-flash) since it runs on every turn of an unlatched session.

How does the trajectory judge determine if a model is stuck?

The EscalationJudge in crates/libsy/src/algorithms/util/escalation.rs analyzes a condensed transcript of recent turns, checking for repetitive patterns, lack of progress toward the stated goal, or circular reasoning. It returns a boolean escalate verdict based on whether the weak model's trajectory indicates it is unlikely to resolve the query successfully.

What happens if the judge fails or times out?

Switchyard implements fail-open semantics. If the judge request times out or returns malformed JSON, the router serves the buffered weak model response and does not increment the escalation streak or latch the session to the strong tier. This prevents infrastructure glitches from triggering expensive strong-model calls.

How can I monitor the cost savings from escalation routing?

Query the /v1/stats endpoint to retrieve token counts and call volumes for each model tier (weak, strong, and classifier). By comparing the ratio of weak-model calls to strong-model calls and calculating the respective API costs, you can derive the exact cost reduction achieved versus running all traffic through the expensive strong model exclusively.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →