# How Switchyard Handles Context Window Overflow: Automatic Fallback to Capable Models

> Discover how Switchyard gracefully handles context window overflow, automatically falling back to capable models when efficient models exceed token limits, ensuring uninterrupted performance.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-12

---

**When an efficient model reports `LlmClientError::ContextWindowExceeded`, Switchyard immediately falls back to the capable model without invoking the judge, treating the overflow as a deterministic capacity failure rather than a quality issue.**

Switchyard's tiered routing architecture in the NVIDIA-NeMo/Switchyard repository optimizes cost and latency by attempting efficient models first, but it must gracefully handle cases where prompts exceed the smaller context windows of efficient tiers. The framework implements a deterministic escalation path that bypasses judgment evaluation to guarantee request completion by routing to higher-capacity capable models.

## The Escalation Algorithm for Context Window Limits

The **EscalationClassifier** drives the routing decision in [`crates/libsy/src/algorithms/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/escalation.rs). It implements a fail-fast strategy for capacity constraints that differs fundamentally from quality-based escalation.

### Efficient Model-First Attempt

By default, Switchyard issues the initial request to the efficient model to minimize costs. The classifier calls `driver.call_model()` with the efficient model ID, awaiting either a successful response or a specific error variant indicating capacity limits.

### Detecting ContextWindowExceeded Errors

When the efficient model cannot process the input due to token limits, it returns `LlmClientError::ContextWindowExceeded`. The classifier captures this specific error in a `match` block (lines 87-107) rather than propagating it as a generic failure. This detection triggers an immediate short-circuit to the capable tier.

## Immediate Fallback Without Judge Evaluation

Unlike content-quality failures that require nuanced evaluation, context window overflows represent hard infrastructure limits. Switchyard optimizes for reliability and latency by eliminating the judge from this specific failure path.

### Deterministic Capacity Handling

Because the overflow is a binary capacity constraint rather than a subjective quality judgment, Switchyard guarantees the fallback to the capable model. This design prevents request starvation when prompts exceed efficient model limits while avoiding unnecessary judge API calls.

### Recording Fallback Metadata

Before returning the capable model classification, the system logs the escalation reason via `driver.set_evidence()`. The method stores a structured JSON payload containing `"reason_code": "context_window"` to maintain audit trails of why the efficient tier was bypassed.

## Implementation in escalation.rs

The core logic resides in the error handling arm of the escalation classifier, where the `ContextWindowExceeded` variant triggers the fallback path:

```rust
// crates/libsy/src/algorithms/escalation.rs
let efficient_response = match driver
    .call_model(request.clone(), vec![efficient.clone()])
    .await
{
    Ok(r) => r,
    Err(LibsyError::ClientCall {
        source: LlmClientError::ContextWindowExceeded { .. },
        ..
    }) => {
        // Record fallback reason and select the capable model.
        driver.set_evidence(serde_json::json!({
            "source": "fallback",
            "reason_code": "context_window",
        }));
        return Ok((decisive(&capable), None));
    }
    Err(e) => return Err(e),
};

```

The `decisive(&capable)` function returns a classification that forces router invocation of the capable model, while the `None` indicates no response was extracted from the failed efficient call.

## Server-Side Error Propagation

The mechanism extends to Switchyard's server layer, where backend errors translate into the same fallback pipeline. The mock server implementation in the soak testing framework demonstrates how overflow errors originate:

```rust
// crates/switchyard-soak/examples/mock_server.rs
if model == "mock/weak" && marker == Some("context_overflow") {
    return (
        StatusCode::BAD_REQUEST,
        Json(json!({
            "error": {
                "code": "context_length_exceeded",
                "message": "the weak target context window is too small"
            }
        })),
    )
    .into_response();
}

```

The router layer translates these `context_length_exceeded` server responses into `LlmClientError::ContextWindowExceeded` errors, ensuring consistent handling across mock and production backends.

## Unit Testing the Overflow Fallback

The `falls_back_to_capable_when_efficient_overflows` test in [`crates/libsy/src/algorithms/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/escalation.rs) verifies the deterministic escalation path. It injects a context window error from the efficient model and asserts successful completion by the capable tier:

```rust
#[tokio::test]
async fn falls_back_to_capable_when_efficient_overflows() -> Result<()> {
    let serve = |target: ModelId, _request: Request| async move {
        match target.as_str() {
            "efficient" => Err(LlmClientError::ContextWindowExceeded {
                model: target,
                message: "prompt is too long".to_string(),
            }),
            _ => Ok(reply("capable answer")),
        }
    };

    let (selected_model, response) = test_drive_with_models(
        escalation_router()?,      // confirmations = 1
        classify_request(),
        runtime_models(),
        serve,
    )
    .await?;

    assert_eq!(selected_model, "capable");
    assert_eq!(
        response.llm_response.as_agg().map(completion_text),
        Some("capable answer".to_string())
    );
    Ok(())
}

```

This test confirms that the escalation classifier returns the capable model ID and its response, bypassing the judge and avoiding request failure.

## Summary

- **Deterministic detection**: Switchyard captures `LlmClientError::ContextWindowExceeded` in [`escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/escalation.rs) (lines 87-107) to identify efficient model capacity limits immediately.
- **Judge bypass**: Context window overflows skip judge evaluation entirely, reducing latency and guaranteeing fallback to capable models.
- **Audit logging**: The system records `"reason_code": "context_window"` via `driver.set_evidence()` to track capacity-based escalations.
- **Verified behavior**: Unit tests confirm the `falls_back_to_capable_when_efficient_overflows` path returns capable model responses when efficient tiers overflow.

## Frequently Asked Questions

### What triggers a fallback from efficient to capable models in Switchyard?

When the efficient model returns the `ContextWindowExceeded` error variant, Switchyard immediately escalates to the capable model. This occurs when input prompts contain more tokens than the efficient model's context window supports, typically encountered with long documents or extensive conversation histories.

### Why does Switchyard bypass the judge for context window overflows?

Switchyard treats context window limits as deterministic capacity constraints rather than quality judgments. Bypassing the judge eliminates an unnecessary API call, reduces end-to-end latency, and guarantees that requests with large contexts always reach capable models that can process them.

### How does Switchyard track context window escalation events?

The `EscalationClassifier` calls `driver.set_evidence()` with a JSON payload containing `"source": "fallback"` and `"reason_code": "context_window"` before returning the capable model classification. This metadata appears in routing logs to help operators analyze capacity constraints and model selection patterns.

### Where is the context window overflow handling implemented?

The core logic resides in [`crates/libsy/src/algorithms/escalation.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/escalation.rs) within the `EscalationClassifier` implementation, specifically in the error handling match block spanning lines 87-107. Additional integration coverage exists in [`crates/switchyard-server/tests/server.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/tests/server.rs) and the mock server at [`crates/switchyard-soak/examples/mock_server.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-soak/examples/mock_server.rs).