# How Bounded Failover Works in a Grok Provider: Mechanism and Implementation

> Discover how bounded failover in Grok provider caps request retries, reroutes to fresh credentials, and prevents infinite loops. Learn the mechanism and implementation.

- Repository: [Chenyme/grok2api](https://github.com/chenyme/grok2api)
- Tags: internals
- Published: 2026-08-09

---

**Bounded failover in the chenyme/grok2api project caps request retries at a configurable maximum (up to 65,535 attempts), automatically rerouting to fresh credentials or alternative providers on retryable errors while preventing infinite loops.**

The **bounded failover** mechanism is a critical reliability feature in Grok2API that prevents resource exhaustion by limiting how many times a single request can cycle through different accounts or providers. By enforcing a strict attempt counter at the gateway level, the system ensures that transient failures trigger automatic recovery without allowing uncontrolled cascading retries that could block clients indefinitely.

## Core Principles of Bounded Failover

### Attempt Limiting with MaxAttempts

Every request processed by the gateway carries a **max‑attempts counter** defined as a `uint16` field, supporting a range of **1 to 65,535** attempts. The initial provider selection counts as the first attempt, and each subsequent failover or egress retry decrements the remaining budget. When the counter reaches zero, the gateway terminates the request with an explicit error rather than continuing to loop. According to the project documentation, this default ceiling prevents runaway resource consumption while accommodating complex multi-provider routing scenarios【[`README.md`](https://github.com/chenyme/grok2api/blob/main/README.md) line 138】.

### Pre-Response Failover Constraints

Failover operations only occur **before** a usable response is emitted to the client. The system strictly prohibits splicing streaming responses across different accounts, meaning once a stream begins, the bound account remains sticky for the duration of that request. As documented in the frontend internationalization strings, "credential expiry, rate limits, and retryable upstream errors can trigger credential refresh or account failover" exclusively during the pre‑response phase【[`frontend/src/shared/i18n/index.ts`](https://github.com/chenyme/grok2api/blob/main/frontend/src/shared/i18n/index.ts) lines 1526‑1527】.

### Sticky Sessions and Bounded Capacity

Once an account binds to a request, subsequent retries respect a **bounded capacity wait** to prevent over‑utilization of the same credential. The selector implementation references a *large‑pool bounded planner policy* that shapes this behavior, ensuring that retries distribute load across the provider pool rather than hammering a single node【[`gateway/selector.go`](https://github.com/chenyme/grok2api/blob/main/gateway/selector.go) line 331】.

### Immediate Recovery Probes

When a fixed proxy fails, the system launches short‑lived recovery probes immediately. Other requests waiting for that specific node are subject to a **bounded waiting window of ≤ 5 seconds**, after which they either retry a healthy node or return a failure. This design isolates failing proxies quickly while maintaining overall throughput【[`FAILURE_RETRY.md`](https://github.com/chenyme/grok2api/blob/main/FAILURE_RETRY.md) lines 1‑2】.

## Implementation Details in the Gateway

### The Routing Attempt Counter

The request struct maintains internal state to track failover exhaustion:

```go
type Request struct {
    MaxAttempts uint16 // configured via client input or default (≤ 65535)
    attempts    uint16 // internal counter incremented on each routing round
}

```

The gateway evaluates `attempts < MaxAttempts` before permitting another failover cycle. This check occurs atomically during the provider selection phase to ensure thread‑safe decrementing across concurrent requests.

### Provider Selection Logic

The core failover logic resides in [`backend/internal/application/gateway/selector.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/gateway/selector.go). The `selectProvider` function implements recursive retry with bounds checking:

```go
func (s *selector) selectProvider(ctx context.Context, req *Request) (provider.Provider, error) {
    // ... standard selection logic ...
    if err := provider.Call(...); isRetryable(err) && req.attempts < req.MaxAttempts {
        req.attempts++
        // refresh credentials or switch to another provider/account
        return s.selectProvider(ctx, req)
    }
    return nil, err
}

```

This implementation ensures that only retryable errors—such as credential expiry, rate limits, or transport failures—trigger the failover path, while hard failures (authentication errors, invalid requests) abort immediately without consuming the attempt budget.

## Configuring Bounded Failover

### Setting Maximum Attempts

Clients can specify the failover limit when constructing requests. The following example demonstrates routing with a maximum of three attempts (initial selection plus two failovers):

```go
import (
    "context"
    "github.com/chenyme/grok2api/backend/internal/application/gateway"
)

func main() {
    // Register providers including the failoverAdapter for testing
    reg := provider.NewRegistry(&failoverAdapter{}, webStoredResponseAdapter{}, statelessConsoleAdapter{})
    
    req := provider.Request{
        Model:       "grok-2",
        MaxAttempts: 3, // strict upper bound on routing rounds
    }
    
    resp, err := reg.Route(context.Background(), req)
    if err != nil {
        log.Fatalf("request failed after bounded failover: %v", err)
    }
    fmt.Println("Response:", resp)
}

```

The `failoverAdapter` mock in the test suite ([`gateway/service_test.go`](https://github.com/chenyme/grok2api/blob/main/gateway/service_test.go)) simulates retryable errors to verify that the gateway respects the configured bound【[`gateway/service_test.go`](https://github.com/chenyme/grok2api/blob/main/gateway/service_test.go) lines 119‑160】.

### Customizing Egress Retry Windows

For infrastructure‑level failover, the egress manager supports bounded waiting intervals:

```go
mgr := egress.NewManager(egress.Config{
    ProbeTimeout: 5 * time.Second, // abort wait after 5s and try next node
})

```

This configuration aligns with the **bounded waiting** design documented in [`FAILURE_RETRY.md`](https://github.com/chenyme/grok2api/blob/main/FAILURE_RETRY.md), ensuring that proxy failures trigger immediate probes while queued requests timeout quickly rather than blocking indefinitely.

## Summary

- **Bounded failover** prevents infinite retry loops by capping attempts at a `uint16` maximum (65,535) configurable per request.
- The mechanism only triggers **pre‑response**; streaming responses remain bound to their initial account to prevent splicing.
- Core logic lives in **[`backend/internal/application/gateway/selector.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/gateway/selector.go)**, utilizing a large‑pool bounded planner policy for load distribution.
- **Immediate recovery probes** and a **5‑second bounded wait** isolate failing proxies without starving the request pool.
- The test suite validates behavior through the **`failoverAdapter`** mocks in [`service_test.go`](https://github.com/chenyme/grok2api/blob/main/service_test.go).

## Frequently Asked Questions

### What is the maximum number of failover attempts allowed in Grok2API?

The system supports up to **65,535 attempts** per request, defined by the `uint16` type of the `MaxAttempts` field. Operators can configure any value between 1 and 65535 via the request configuration or gateway defaults. Once the internal counter reaches this limit, the gateway aborts the request immediately.

### Does bounded failover work with streaming responses?

No. The bounded failover mechanism only operates **before** a usable response is returned. Streaming responses are never spliced across different accounts or providers. If a failure occurs mid‑stream, the connection terminates rather than failing over to preserve data integrity and client expectations.

### Where is the bounded failover logic implemented?

The primary implementation resides in **[`backend/internal/application/gateway/selector.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/gateway/selector.go)**, specifically within the provider selection and routing functions. Supporting documentation and immediate probe logic are located in **[`backend/internal/infra/egress/FAILURE_RETRY.md`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/egress/FAILURE_RETRY.md)**, while the behavior is verified in **[`backend/internal/application/gateway/service_test.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/gateway/service_test.go)**.

### How does the gateway distinguish between retryable and fatal errors?

The `selectProvider` function evaluates errors using an internal `isRetryable` helper. Retryable conditions include credential expiry, rate‑limit responses, and transport‑level failures. Authentication failures, malformed requests, and other client errors bypass the failover counter and return immediately to prevent wasting resources on doomed requests.