How Switchyard Escalation Routing Detects and Handles Model Response Failures
Switchyard escalation routing detects model response failures by categorizing errors from the weak model into context-window overflows, transport failures, or API exceptions, while judge model failures trigger a "fail open" mechanism that preserves the buffered weak response and maintains the session streak without forcing an expensive strong-model call.
Switchyard, the open-source LLM routing engine in the NVIDIA-NeMo/Switchyard repository, implements a cost-optimization strategy that queries an efficient "weak" model first, then uses a judge model to determine if the session requires escalation to a stronger model. This architecture must gracefully handle failures from both the efficient model and the judge classifier to prevent service degradation. Understanding how Switchyard escalation routing manages these failure modes ensures reliable, cost-efficient inference deployments.
Detecting Weak Model Failures
The EscalationClassifier implementation in crates/libsy/src/algorithms/escalation.rs intercepts and categorizes failures when calling the efficient model into three distinct error types, each with specific fallback behaviors.
Context-Window Overflow
When the weak model's prompt exceeds its maximum token limit, the system catches LlmClientError::ContextWindowExceeded at lines 99-101. Rather than attempting to truncate or retry, the route immediately invokes decisive(&capable) to fallback to the strong model without awaiting a judge verdict. This ensures that context limitations on the efficient tier never block user requests.
Transport Failures
If the weak model's streaming transport drops while buffering the response, the code matches LlmClientError::Transport at lines 113-119. Like context-window errors, transport failures trigger an immediate fallback to the capable model via decisive(&capable), maintaining service continuity when network instability affects the efficient tier.
Other API Errors
For API-level errors not covered by the specific transport or context categories—such as authentication failures or rate limits—the system propagates the error via LibsyError::client_call at lines 122-124. These errors bubble up to abort the request, allowing upstream callers to implement custom retry or circuit-breaker logic rather than silently falling back to expensive model calls.
Handling Judge Model Failures
After successfully buffering a weak model response, Switchyard constructs a judge request containing the weak reply as an assistant message and dispatches it to the configured judge model. According to the documentation at docs/routing_algorithms/escalation_router_routing.md lines 89-92, the system implements "fail open" semantics for judge failures:
"A judge that times out, errors, or returns an unparseable verdict fails open: the turn serves the buffered weak reply and the existing streak is held rather than cleared. A judge failure never creates a strong-tier latch."
This design produces three critical behaviors when the judge fails:
- The buffered weak response is served to the client immediately, ensuring no latency penalty from judge timeouts.
- The escalation streak is preserved, allowing subsequent turns to accumulate toward the
confirmationsthreshold defined inEscalationJudgeConfig(located incrates/libsy/src/algorithms/util/escalation.rs). - No strong-tier latch is created, guaranteeing that transient judge glitches cannot force expensive model invocations.
The Complete Escalation Flow
The end-to-end failure handling workflow operates as follows:
-
Weak Model Invocation – The system calls the efficient model. If
LlmClientError::ContextWindowExceededorLlmClientError::Transportoccurs, it immediately serves the strong model viadecisive(&capable). Other errors propagate upward viaLibsyError::client_call. -
Response Buffering – On success, the weak reply is buffered and appended to the conversation transcript. The system then constructs and sends the judge request.
-
Judge Evaluation –
- If the judge returns a valid verdict, the escalation streak increments. When the streak reaches the configured
confirmationsvalue, the session latches to the strong model for subsequent turns. - If the judge times out, errors, or returns malformed JSON, the route fails open: the weak reply is served, the streak remains unchanged, and the session stays unlatched.
- If the judge returns a valid verdict, the escalation streak increments. When the streak reaches the configured
This architecture ensures that model-level failures never silently degrade user experience while maintaining strict cost controls.
Configuration and Usage Examples
The following TOML configuration defines an escalation route with a judge classifier:
[routes.agent]
id = "agent"
type = "llm_classifier"
mode = "escalation"
classifier_target = "judge"
strong_target = "strong"
weak_target = "weak"
prompt = "Judge whether the weak model is stuck. Return the required structured verdict."
escalation = { confirmations = 2 }
To trigger this routing logic via HTTP:
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "x-switchyard-session-id: demo-session" \
-d '{"model":"agent","messages":[{"role":"user","content":"Explain quantum tunneling"}]}'
If the weak model encounters a context-window limit, the server logs the fallback (see the trace! macro at line 90 in crates/libsy/src/algorithms/escalation.rs) and returns the strong model's response immediately. If the judge times out, the system returns the weak model's buffered reply and retains the escalation streak for future turns.
Summary
- Weak model failures are categorized into context-window overflows, transport drops, and other API errors, with specific fallbacks to the strong model or error propagation.
- Judge failures trigger "fail open" behavior that serves the weak response and preserves the escalation streak, preventing expensive strong-model calls during judge outages.
- Key implementation files include
crates/libsy/src/algorithms/escalation.rsfor the core classifier logic andcrates/libsy/src/algorithms/util/escalation.rsfor judge configuration. - Testing coverage includes scenarios like
stage_route_escalates_on_a_signal_in_the_conversationto validate fallback behavior under failure conditions.
Frequently Asked Questions
What happens when the weak model exceeds its context window in Switchyard?
When the weak model encounters a context-window overflow, Switchyard catches LlmClientError::ContextWindowExceeded at lines 99-101 in crates/libsy/src/algorithms/escalation.rs and immediately falls back to the strong model via decisive(&capable). This ensures the user receives a valid response without waiting for a judge verdict.
How does Switchyard handle judge model timeouts during escalation routing?
Judge timeouts trigger "fail open" semantics: the system serves the buffered weak model response to the client, preserves the existing escalation streak, and avoids creating a strong-tier latch. This behavior is documented at lines 89-92 in docs/routing_algorithms/escalation_router_routing.md and prevents transient judge failures from forcing expensive model calls.
Can transport errors from the weak model cause request failures in Switchyard?
No. Transport errors such as dropped streaming connections are caught via LlmClientError::Transport at lines 113-119 and trigger an immediate fallback to the capable model. Only uncategorized API errors propagate upward via LibsyError::client_call to allow upstream handling.
Where is the escalation confirmation threshold configured in Switchyard?
The confirmations threshold—which determines how many consecutive successful judge verdicts are required to latch to the strong model—is defined in EscalationJudgeConfig within crates/libsy/src/algorithms/util/escalation.rs. This value is set in the route configuration TOML under the escalation key.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →