How the LLM Classifier Routing Escalation Mode Works in Switchyard
Switchyard's escalation mode implements a response-based routing workflow that calls an efficient LLM first, uses a judge LLM to evaluate response quality, and escalates to a capable LLM only when the judge verdict meets configurable confirmation criteria.
Switchyard, NVIDIA's open-source LLM routing framework, provides three built-in routing strategies through the LlmClassifierConfig type. The LLM Classifier Routing escalation mode enables intelligent cost optimization by attempting fast, inexpensive inference first while maintaining quality guarantees through automated escalation triggers. This pattern reduces infrastructure costs by routing only complex queries to expensive, high-capability models.
Understanding the Escalation Workflow
The escalation classifier implements a three-stage decision pipeline:
-
Initial inference – The request routes to the
efficient_target, a cheap or fast model that handles straightforward queries. -
Quality judgment – A specialized
judge_targetevaluates the efficient model's response against quality or safety criteria. -
Conditional escalation – If the judge returns consecutive affirmative verdicts meeting the
confirmationsthreshold, the system re-routes the request to thecapable_targetand returns the higher-quality response.
Escalation Annotations
The workflow attaches optional escalation notes and de-escalation notes to requests, providing observability hooks for downstream logging components to track routing decisions.
Configuring the Escalation Classifier
The EscalationClassifierConfig class controls the escalation behavior. According to the Switchyard source code in switchyard_rust/libsy.py (lines 81-98), the configuration accepts these parameters:
confirmations– Number of consecutive "yes" judgments required before triggering escalation (default: 2)recent_turn_window– Conversation history turns examined during judgment (default: 28)window_message_chars– Minimum character threshold per turn for context inclusion (≥ 50)max_output_tokens– Token limit for the escalation target's response (default: 4096)prompt– Optional system prompt prepended to judge requestsresponse_format_type– Expected judge output format (json_schemaorjson_object)
Implementation in the Source Code
The escalation factory method resides in the Rust-backed Python bindings. The static method LlmClassifierConfig.escalation (lines 78-86 of switchyard_rust/libsy.py) constructs a classifier instance binding the three targets together:
judge_target– The lightweight evaluator modelefficient_target– The fast/cheap initial inference modelcapable_target– The high-quality fallback model
Internally, the Rust algorithm:
- Sends the initial request to the efficient endpoint
- Feeds the response to the judge endpoint for binary verdict classification
- Invokes the capable endpoint only when verdicts satisfy the confirmation rule
Practical Implementation Example
Configure escalation routing using the Python API:
from switchyard.libsy import (
EscalationClassifierConfig,
LlmClassifierConfig,
llm_classifier,
stage_router,
)
# Configure escalation criteria
escalation_cfg = EscalationClassifierConfig(
confirmations=3,
recent_turn_window=20,
window_message_chars=200,
max_output_tokens=2048,
prompt="You are a safety judge. Decide if the answer is sufficient.",
response_format_type="json_object",
)
# Create classifier binding three targets
classifier = LlmClassifierConfig.escalation(
judge_target="safety-judge",
efficient_target="fast-model",
capable_target="high-quality",
config=escalation_cfg,
)
# Build routing algorithm with observability hooks
routing_algo = stage_router(
capable_target="high-quality",
efficient_target="fast-model",
picker="confidence",
confidence_threshold=0.8,
recent_window=10,
escalation_note="Escalated after safety check failed",
deescalation_note="Returned fast answer",
only_on_wrong_signal_escalation=True,
classifier=llm_classifier(classifier),
)
This configuration requires three consecutive affirmative judgments before escalating from the fast model to the high-quality model, with full audit trails attached to routing decisions.
Summary
- The escalation mode implements a "try-fast-first" pattern using three distinct LLM targets: efficient, judge, and capable
- EscalationClassifierConfig controls confirmation thresholds, context windows, and output limits through parameters like
confirmationsandrecent_turn_window - The implementation lives in
switchyard_rust/libsy.py, specifically theEscalationClassifierConfigdefinition (lines 81-98) andLlmClassifierConfig.escalationfactory (lines 78-86) - Optional escalation notes provide observability into routing decisions for downstream logging systems
Frequently Asked Questions
What triggers an escalation in Switchyard's classifier?
An escalation triggers when the judge LLM returns consecutive affirmative verdicts meeting the confirmations threshold defined in EscalationClassifierConfig. By default, the system requires two consecutive "yes" judgments, though this is configurable per deployment.
How does the escalation mode differ from other Switchyard routing strategies?
Unlike static routing or confidence-based selection, the escalation mode specifically implements a sequential evaluation pattern that unconditionally calls the efficient model first, then conditionally escalates based on post-hoc quality assessment rather than pre-request classification.
Can I configure different confirmation thresholds for different query types?
Yes, you instantiate separate EscalationClassifierConfig objects with varying confirmations values and create distinct classifiers using LlmClassifierConfig.escalation(). Each classifier instance can then be bound to different route patterns or API endpoints within your stage_router configuration.
Where is the escalation logic actually executed?
The core escalation algorithm executes within the compiled Rust library loaded via switchyard_rust/_native.py, while the configuration and target binding occur in the Python layer through switchyard_rust/libsy.py. The Python bindings marshal configuration data to the Rust runtime where the actual three-stage routing logic executes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →