# How Switchyard Escalation Routing Detects and Handles Model Response Failures

> Learn how Switchyard escalation routing detects and handles model response failures by categorizing errors and triggering fail open mechanisms to maintain session streaks and preserve weak responses.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-12

---

**Switchyard escalation routing detects model response failures by categorizing errors from the weak model into context-window overflows, transport failures, or API exceptions, while judge model failures trigger a "fail open" mechanism that preserves the buffered weak response and maintains the session streak without forcing an expensive strong-model call.**

Switchyard, the open-source LLM routing engine in the NVIDIA-NeMo/Switchyard repository, implements a cost-optimization strategy that queries an efficient "weak" model first, then uses a judge model to determine if the session requires escalation to a stronger model. This architecture must gracefully handle failures from both the efficient model and the judge classifier to prevent service degradation. Understanding how Switchyard escalation routing manages these failure modes ensures reliable, cost-efficient inference deployments.

## Detecting Weak Model Failures

The `EscalationClassifier` implementation in [`crates/libsy/src/algorithms/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/escalation.rs) intercepts and categorizes failures when calling the efficient model into three distinct error types, each with specific fallback behaviors.

### Context-Window Overflow

When the weak model's prompt exceeds its maximum token limit, the system catches `LlmClientError::ContextWindowExceeded` at lines 99-101. Rather than attempting to truncate or retry, the route immediately invokes `decisive(&capable)` to fallback to the strong model without awaiting a judge verdict. This ensures that context limitations on the efficient tier never block user requests.

### Transport Failures

If the weak model's streaming transport drops while buffering the response, the code matches `LlmClientError::Transport` at lines 113-119. Like context-window errors, transport failures trigger an immediate fallback to the capable model via `decisive(&capable)`, maintaining service continuity when network instability affects the efficient tier.

### Other API Errors

For API-level errors not covered by the specific transport or context categories—such as authentication failures or rate limits—the system propagates the error via `LibsyError::client_call` at lines 122-124. These errors bubble up to abort the request, allowing upstream callers to implement custom retry or circuit-breaker logic rather than silently falling back to expensive model calls.

## Handling Judge Model Failures

After successfully buffering a weak model response, Switchyard constructs a judge request containing the weak reply as an `assistant` message and dispatches it to the configured judge model. According to the documentation at [`docs/routing_algorithms/escalation_router_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/escalation_router_routing.md) lines 89-92, the system implements "fail open" semantics for judge failures:

> "A judge that times out, errors, or returns an unparseable verdict fails open: the turn serves the buffered weak reply and the existing streak is held rather than cleared. A judge failure never creates a strong-tier latch."

This design produces three critical behaviors when the judge fails:

- **The buffered weak response is served** to the client immediately, ensuring no latency penalty from judge timeouts.
- **The escalation streak is preserved**, allowing subsequent turns to accumulate toward the `confirmations` threshold defined in `EscalationJudgeConfig` (located in [`crates/libsy/src/algorithms/util/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs)).
- **No strong-tier latch is created**, guaranteeing that transient judge glitches cannot force expensive model invocations.

## The Complete Escalation Flow

The end-to-end failure handling workflow operates as follows:

1. **Weak Model Invocation** – The system calls the efficient model. If `LlmClientError::ContextWindowExceeded` or `LlmClientError::Transport` occurs, it immediately serves the strong model via `decisive(&capable)`. Other errors propagate upward via `LibsyError::client_call`.

2. **Response Buffering** – On success, the weak reply is buffered and appended to the conversation transcript. The system then constructs and sends the judge request.

3. **Judge Evaluation** –
   - If the judge returns a valid verdict, the escalation streak increments. When the streak reaches the configured `confirmations` value, the session latches to the strong model for subsequent turns.
   - If the judge times out, errors, or returns malformed JSON, the route fails open: the weak reply is served, the streak remains unchanged, and the session stays unlatched.

This architecture ensures that model-level failures never silently degrade user experience while maintaining strict cost controls.

## Configuration and Usage Examples

The following TOML configuration defines an escalation route with a judge classifier:

```toml
[routes.agent]
id = "agent"
type = "llm_classifier"
mode = "escalation"
classifier_target = "judge"
strong_target = "strong"
weak_target = "weak"
prompt = "Judge whether the weak model is stuck. Return the required structured verdict."
escalation = { confirmations = 2 }

```

To trigger this routing logic via HTTP:

```bash
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "x-switchyard-session-id: demo-session" \
  -d '{"model":"agent","messages":[{"role":"user","content":"Explain quantum tunneling"}]}'

```

If the weak model encounters a context-window limit, the server logs the fallback (see the `trace!` macro at line 90 in [`crates/libsy/src/algorithms/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/escalation.rs)) and returns the strong model's response immediately. If the judge times out, the system returns the weak model's buffered reply and retains the escalation streak for future turns.

## Summary

- **Weak model failures** are categorized into context-window overflows, transport drops, and other API errors, with specific fallbacks to the strong model or error propagation.
- **Judge failures** trigger "fail open" behavior that serves the weak response and preserves the escalation streak, preventing expensive strong-model calls during judge outages.
- **Key implementation files** include [`crates/libsy/src/algorithms/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/escalation.rs) for the core classifier logic and [`crates/libsy/src/algorithms/util/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs) for judge configuration.
- **Testing coverage** includes scenarios like `stage_route_escalates_on_a_signal_in_the_conversation` to validate fallback behavior under failure conditions.

## Frequently Asked Questions

### What happens when the weak model exceeds its context window in Switchyard?

When the weak model encounters a context-window overflow, Switchyard catches `LlmClientError::ContextWindowExceeded` at lines 99-101 in [`crates/libsy/src/algorithms/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/escalation.rs) and immediately falls back to the strong model via `decisive(&capable)`. This ensures the user receives a valid response without waiting for a judge verdict.

### How does Switchyard handle judge model timeouts during escalation routing?

Judge timeouts trigger "fail open" semantics: the system serves the buffered weak model response to the client, preserves the existing escalation streak, and avoids creating a strong-tier latch. This behavior is documented at lines 89-92 in [`docs/routing_algorithms/escalation_router_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/escalation_router_routing.md) and prevents transient judge failures from forcing expensive model calls.

### Can transport errors from the weak model cause request failures in Switchyard?

No. Transport errors such as dropped streaming connections are caught via `LlmClientError::Transport` at lines 113-119 and trigger an immediate fallback to the capable model. Only uncategorized API errors propagate upward via `LibsyError::client_call` to allow upstream handling.

### Where is the escalation confirmation threshold configured in Switchyard?

The `confirmations` threshold—which determines how many consecutive successful judge verdicts are required to latch to the strong model—is defined in `EscalationJudgeConfig` within [`crates/libsy/src/algorithms/util/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/escalation.rs). This value is set in the route configuration TOML under the `escalation` key.