# How Switchyard Handles Upstream API Failures and Implements Fallback Logic

> Discover how Switchyard's two-layer resilience system manages upstream API failures. Learn about its client-side retry loop and server-side fallback for robust request routing.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-17

---

**Switchyard handles upstream API failures through a two-layer resilience system: a client-side retry loop with exponential back-off for transient HTTP errors, and a server-side health-aware candidate fallback mechanism that routes requests to alternative backends when retries are exhausted.**

NVIDIA-NeMo/Switchyard implements robust handling of upstream API failures and fallback logic through a clear architectural separation between low-level HTTP resilience and high-level routing decisions. This design ensures that transient errors are automatically recovered while persistent failures trigger seamless failover to healthy backends. The implementation spans the LLM client library, the server-side router, and algorithmic fallback modules that work together to maintain service availability.

## Retry Logic and Exponential Back-Off

The first layer of defense against upstream instability lives in the LLM client, where every request is wrapped in a sophisticated retry loop that respects both exponential back-off and server-provided retry directives.

### The Core Retry Loop Implementation

In [`crates/libsy-llm-client/src/client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs) (lines 31‑78), the `call_upstream` method implements the retry logic using a configurable `max_retries` value per backend. The loop captures rich tracing spans and calculates delays using the `retry_delay` function, which considers both the attempt number and any `Retry-After` header returned by the upstream server.

```rust
let max_retries = u64::from(backend.max_retries());
let mut attempt = 0_u64;
loop {
    let result = self.send_once(&url, backend, &body, metadata, model, streaming).await;
    match result {
        Ok(resp) => return Ok(resp),
        Err(failure) => {
            let will_retry = attempt < max_retries && failure.is_retryable();
            if !will_retry { return Err(failure.error); }

            let delay = retry_delay(attempt, failure.retry_after);
            tokio::time::sleep(delay).await;
            attempt += 1;
        }
    }
}

```

### Classifying Retryable Errors

Whether an error triggers a retry is determined by `AttemptFailure.is_retryable()`, which delegates to `metrics::is_retryable_http_status` defined in [`crates/libsy-llm-client/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/metrics.rs) (lines 10‑18). This function checks the HTTP status code against known transient error codes such as **429** (Too Many Requests), **500** (Internal Server Error), and **504** (Gateway Timeout), while also parsing the `Retry-After` header to respect server-side rate limiting.

### Enforcing Global Retry Budgets

To prevent retry storms across the entire service, Switchyard implements a **retry budget** enforced at the server level in [`crates/switchyard-server/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/config.rs). This budget caps the total number of retries across all concurrent requests, ensuring that transient upstream failures do not cascade into systemic overload.

## Health-Aware Candidate Fallback

When the retry budget is exhausted or the error is classified as non-retryable, Switchyard’s second resilience layer activates to route the request to an alternative backend without surfacing the failure to the caller.

### The StageRouter Decision Engine

The routing logic resides in [`crates/switchyard-server/src/stats/algorithms/stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/stage_router.rs). The **StageRouter** maintains continuous health metrics for each candidate (model/endpoint pair) including latency percentiles and error rates. When a candidate returns a fatal error, the router evaluates whether to fall through to the next viable candidate based on real-time health data.

### The Fall-Through Algorithm

The actual fallback behavior is codified in [`crates/libsy/src/algorithms/fall_through.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/fall_through.rs). When a candidate error is deemed non-recoverable, the algorithm returns `CandidateResult::Fallback`, signaling the StageRouter to select the next healthy candidate from the deployment configuration.

```rust
match candidate.run(request).await {
    Ok(resp) => return Ok(resp),
    Err(err) if err.is_fatal() => {
        // Candidate exhausted – fall through to the next healthy one
        let next = router.next_candidate()?;
        return router.dispatch(next, request).await;
    }
    Err(err) => return Err(err),
}

```

## Complete Request Flow

Understanding how Switchyard handles upstream API failures requires tracing the full lifecycle of a request from construction through potential fallback.

### Request Construction and Initial Dispatch

The flow begins in client launchers such as [`switchyard/cli/launchers/codex_cli_launcher.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launchers/codex_cli_launcher.py), which constructs an `LlmRequest` and forwards it to the native server. The server initializes a **candidate**—a concrete model endpoint selected from the deployment TOML configuration—and invokes `client.call_once` to perform the initial HTTP attempt.

### Retry Evaluation and Back-Off

Upon failure, the `AttemptFailure` struct captures the error, HTTP status code, and `Retry-After` value. The outer loop in `client.call_upstream` evaluates `will_retry` using the criteria `attempt < max_retries && failure.is_retryable()`. If true, the system records the outcome, computes the exponential back-off delay, and sleeps before the next attempt.

### Fallback Activation

If retries are exhausted or the error is non-retryable, the error bubbles up to the StageRouter. The router checks `candidate.health` and invokes the fall-through algorithm to yield a new candidate. The request then re-enters the retry loop against the alternative backend using the same resilience logic.

### Observability and Metrics

Every attempt—including retries and fallbacks—updates Prometheus counters such as `switchyard_upstream_attempts_total` and `switchyard_router_retry_recovered_total` (see [`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs)). Tracing spans labeled `libsy.upstream_attempt` capture attempt numbers, outcome classifications, status codes, and back-off delays, enabling end-to-end debugging of failure scenarios.

## Summary

- **Layered resilience**: Switchyard combines client-side retry loops with server-side fallback routing to handle upstream API failures gracefully.
- **Exponential back-off**: The retry mechanism in [`client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/client.rs) respects both exponential delays and the `Retry-After` header while enforcing a global retry budget.
- **Health-aware routing**: The StageRouter continuously monitors candidate health and uses the fall-through algorithm to select alternative backends when errors are non-retryable.
- **Granular observability**: Built-in Prometheus metrics and tracing spans provide detailed visibility into retry attempts, fallback events, and upstream error classification.

## Frequently Asked Questions

### How does Switchyard determine if an HTTP error is retryable?

Switchyard uses the `AttemptFailure.is_retryable()` method, which checks HTTP status codes against `metrics::is_retryable_http_status` in [`crates/libsy-llm-client/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/metrics.rs). Transient errors like 429, 500, and 504 trigger retries, while the system also respects the `Retry-After` header to comply with upstream rate limiting.

### What happens when the retry budget is exhausted?

When the global retry budget defined in [`crates/switchyard-server/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/config.rs) is depleted, or when `max_retries` is reached for a specific backend, the client stops retrying and returns the error to the StageRouter. The router then initiates the fallback sequence to route the request to an alternative candidate.

### How does the fallback mechanism choose alternative backends?

The StageRouter in [`crates/switchyard-server/src/stats/algorithms/stage_router.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/algorithms/stage_router.rs) evaluates candidate health metrics including latency and error rates. When the current candidate fails with a non-retryable error, the fall-through algorithm in [`crates/libsy/src/algorithms/fall_through.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/fall_through.rs) returns `CandidateResult::Fallback`, prompting the router to select the next healthy candidate from the deployment configuration.

### Where is the retry delay calculated in the source code?

The retry delay calculation occurs in [`crates/libsy-llm-client/src/client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs) within the retry loop. The `retry_delay` function computes the wait time using exponential back-off based on the attempt number, while optionally incorporating the `Retry-After` header value if provided by the upstream server.