Switchyard Escalation Router: Trigger Conditions and Mode Differences Explained

Switchyard’s escalation router promotes conversations to a strong model only after a judge returns consecutive escalate verdicts matching a configurable confirmation threshold, while standard LLM classifiers predict difficulty before any model execution.

The NVIDIA-NeMo/Switchyard framework provides intelligent request routing for multi-turn LLM workloads. Its escalation router dynamically monitors conversation health at runtime and triggers tier promotion only when the weak model demonstrably fails. This contrasts sharply with the standard llm_classifier mode, which predicts query difficulty upfront and commits to a tier before execution begins.

How the Escalation Mechanism Works

The escalation router in crates/libsy/src/algorithms/escalation.rs implements a runtime monitoring pattern. Every conversation begins on a cost-effective weak model. After each turn, the router invokes a judge model to evaluate whether the weak model is stuck.

Judge Verdict Processing

The judge returns a structured verdict of either escalate or decline. The router maintains a per-session streak counter that increments on each escalate verdict and resets to zero on any decline. Only when this streak reaches the confirmations threshold (typically configured to 2) does the router latch the session to the strong model, bypassing further judging for subsequent turns.

If the judge times out, errors, or returns an unparsable verdict, the router fails open: it serves the buffered weak-model reply without incrementing the streak. This guarantees that a faulty judge never forces an unintended strong-model latch.

Session State Requirements

Unlike standard routing modes, escalation requires a persistent session identifier passed via the x-switchyard-session-id header. The router stores streak counters per session ID, enabling multi-turn tracking across distributed requests.

Conditions That Trigger Escalation

The judge evaluates conversation history—controlled by recent_turn_window (default 28 turns) and window_message_chars (default 500 characters)—to identify specific failure patterns that warrant escalation:

  • Repeated errors: The weak model generates nonsensical outputs or failing code repeatedly within the monitored window.
  • Loops or drift: The conversation enters a repetitive cycle or drifts away from the intended task objective.
  • Sustained trouble: Any pattern of problematic behavior identified by the judge's prompt logic as requiring rescue intervention.

The escalation trigger is not based on single-turn failures but on sustained dysfunction across multiple consecutive evaluations.

Escalation Mode vs. Standard LLM Classifier

The mode = "escalation" setting fundamentally changes the routing decision timeline compared to the default classifier behavior:

Aspect mode = "escalation" Standard llm_classifier
Decision timing After weak model generates output and judge evaluates it Before any model call, based on request text analysis
Trigger mechanism Consecutive escalate verdicts matching confirmations threshold Single prediction (e.g., high, low) from classifier model
Cost structure One weak call + one judge call per unlatched turn; strong call only post-escalation One classifier call upfront, then only the selected tier
State management Requires session persistence (x-switchyard-session-id) Stateless; each request independent
Optimal use case Multi-turn agent workloads where failure emerges during execution One-shot requests where difficulty is predictable beforehand

Configuration and Implementation

Escalation Router Configuration

Define an escalation route in your TOML configuration by setting type = "llm_classifier" with mode = "escalation":

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.judge]
id = "google/gemini-3.5-flash"
llm_client = "openrouter"

[targets.strong]
id = "anthropic/claude-opus-4.7"
llm_client = "openrouter"

[targets.weak]
id = "moonshotai/kimi-k2.6"
llm_client = "openrouter"

[routes.agent]
id = "agent"
type = "llm_classifier"
mode = "escalation"
classifier_target = "judge"
strong_target = "strong"
weak_target = "weak"
prompt = "Judge whether the weak model is stuck. Return the required structured verdict."
escalation = { confirmations = 2, recent_turn_window = 28, window_message_chars = 500 }

The escalation table configures the streak threshold and history window sizes. See docs/routing_algorithms/escalation_router_routing.md for complete schema documentation.

Making Escalation-Enabled Requests

Include the session header to enable streak tracking across turns:

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "x-switchyard-session-id: demo-session" \
  -d '{
        "model":"agent",
        "messages":[{"role":"user","content":"Debug this recursive Python function"}]
      }'

The first turn executes on the weak model. If the judge returns escalate on two consecutive turns (per confirmations = 2), the third turn routes directly to the strong model without judge overhead.

Standard Classifier Comparison

A standard configuration omits the escalation mode, making tier decisions before execution:

[routes.agent]
id = "agent"
type = "llm_classifier"

# Defaults to capability prediction mode

classifier_target = "classifier"
strong_target = "strong"
weak_target = "weak"

Here, the router predicts query difficulty using the classifier target before invoking either tier, eliminating runtime judging overhead but losing the ability to rescue conversations that degrade mid-stream.

Summary

  • Switchyard's escalation router monitors conversation health post-execution using a judge model that returns escalate or decline verdicts.
  • Escalation triggers only after consecutive escalate verdicts match the configured confirmations threshold, requiring persistent session IDs.
  • Trigger conditions include repeated errors, conversational loops, and task drift detected within the configured history window.
  • Unlike standard LLM classifiers that predict difficulty before execution, escalation mode acts as a runtime rescue mechanism for multi-turn agent workloads.
  • The implementation in crates/libsy/src/algorithms/escalation.rs handles streak counting, session latching, and fail-open safety for judge failures.

Frequently Asked Questions

What happens if the judge model fails or returns an invalid verdict?

The router implements fail-open safety. If the judge times out, errors, or returns unparsable structured output, the router serves the weak model's buffered response and does not increment the escalation streak. This prevents judge malfunctions from forcing unnecessary strong-model usage.

Can I adjust how much conversation history the judge evaluates?

Yes. The recent_turn_window parameter (default 28) controls how many prior turns the judge examines, while window_message_chars (default 500) limits the total character count of history provided to the judge. Configure these in the escalation table of your route definition to balance context richness against judge latency and cost.

Why does escalation mode require session headers while standard classification does not?

Escalation mode maintains a streak counter per conversation to track consecutive escalate verdicts. The router stores this state keyed by the x-switchyard-session-id header value. Standard classifiers make independent predictions per request without tracking conversational state, eliminating the need for session persistence.

How does the cost profile differ between escalation and standard classification?

Escalation mode incurs one weak model call plus one judge call for every turn until the escalation threshold is reached and the session latches to the strong model. Standard classification incurs one classifier call upfront, followed by only the selected tier's calls. Escalation optimizes for scenarios where most conversations succeed on the weak model, while standard classification optimizes for predictable per-request routing without runtime judging overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →