# How Cooldown Periods Drive Provider Failover in Grok2API

> Discover how cooldown periods in Grok2API enable provider failover by temporarily removing failing providers, ensuring automatic routing to healthy accounts and preventing retry loops.

- Repository: [Chenyme/grok2api](https://github.com/chenyme/grok2api)
- Tags: internals
- Published: 2026-07-16

---

**Cooldown periods temporarily remove failing providers from the routing pool, forcing automatic failover to healthy accounts while preventing rapid retry loops.**

In the `chenyme/grok2api` project, each provider account tracks a `cooldown_until` timestamp that governs eligibility for request routing. When a request fails, the egress manager calculates a back-off interval and stores it via `AccountRepository.UpdateHealth`, effectively isolating the misbehaving provider until recovery is likely. This mechanism ensures that the selector logic automatically routes traffic to alternative providers while the failed account remains excluded from the pool.

## The Cooldown State in Account Records

Every provider account in Grok2API maintains a `cooldown_until` field that determines its availability for routing. The egress manager writes this timestamp after detecting failures, while the selector reads it before choosing an account for incoming requests. This simple timestamp acts as a circuit breaker that prevents the system from hammering unstable providers.

## Exponential Backoff Calculation

When a failure occurs, the egress manager calculates the cooldown duration using an exponential backoff algorithm capped at ten minutes. The implementation in [`backend/internal/infra/egress/manager.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/egress/manager.go) uses the failure count to determine the interval, ensuring that repeated failures result in progressively longer isolation periods.

```go
// backend/internal/infra/egress/manager.go
cooldown := min(10*time.Minute,
                30*time.Second*time.Duration(1<<min(value.FailureCount-1, 4)))
until := now.Add(cooldown)
repo.UpdateHealth(..., &until, ..., false)

```

The calculation doubles the base interval of 30 seconds for each consecutive failure, up to a maximum of 10 minutes. This prevents temporary glitches from triggering long outages while protecting against persistent failures.

## How the Selector Implements Failover

The selector component in [`backend/internal/application/gateway/selector.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/gateway/selector.go) queries the account repository and filters out any accounts where `cooldown_until` is in the future. This exclusion happens automatically, causing the router to pick the next available healthy provider without requiring explicit failover configuration.

```go
func (m *Manager) handleFailure(ctx context.Context, accID uint64, failureCount int) {
    now := time.Now().UTC()
    // exponential back‑off, capped at 10 min
    cooldown := min(10*time.Minute,
        30*time.Second*time.Duration(1<<min(failureCount-1, 4)))
    until := now.Add(cooldown)

    _ = m.accountRepo.UpdateHealth(ctx, accID, failureCount, &until,
        "request failed", false)
}

```

When the selector encounters an account still within its cooldown window, it treats that provider as unavailable and skips to the next candidate in the pool.

## The Complete Failover Flow

Understanding how cooldown periods affect provider failover requires following the request lifecycle through the system:

1. **Initial Selection** – The selector picks an enabled, authenticated account whose `cooldown_until` is nil or in the past.
2. **Failure Detection** – When a request fails, the egress manager records the failure, increments `failure_count`, and calculates a new `cooldown_until` timestamp.
3. **Pool Exclusion** – Subsequent requests trigger the selector’s query, which excludes any account with `cooldown_until > now`.
4. **Automatic Failover** – Because the failing provider is hidden from the routing pool, the selector automatically falls back to the next healthiest account, which may be a different provider tier or another account in the same tier.
5. **Recovery** – When `cooldown_until` expires, the account becomes eligible again. A successful request clears the failure count and removes the cooldown stamp entirely.

This flow ensures that **cooldown periods directly drive failover behavior** by manipulating the visibility of providers in the routing pool.

## Configuration and Monitoring

You can inspect current cooldown settings through the HTTP settings handler exposed in [`backend/internal/transport/http/settings/handler.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/transport/http/settings/handler.go). The configuration values `cooldownBase` (30 seconds) and `cooldownMax` (10 minutes) are defined in [`backend/internal/infra/config/config.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/config/config.go) and determine the bounds of the exponential backoff calculation. These settings allow operators to tune how aggressively the system isolates failing providers based on their specific reliability requirements.

## Summary

- **Cooldown periods** are stored as `cooldown_until` timestamps in account records and act as temporary exclusion flags for the routing selector.
- **Exponential backoff** calculates cooldown durations starting at 30 seconds and doubling up to a 10-minute maximum, as implemented in [`backend/internal/infra/egress/manager.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/egress/manager.go).
- **Automatic failover** occurs because [`backend/internal/application/gateway/selector.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/gateway/selector.go) filters out accounts with active cooldowns, forcing selection of healthier providers.
- **Recovery** happens automatically when the cooldown timestamp expires, with successful requests clearing the failure history and restoring full eligibility.

## Frequently Asked Questions

### How long does a provider stay in cooldown?

The duration depends on the failure count using an exponential backoff formula: `min(10 minutes, 30 seconds × 2^(failureCount-1))`. First failures incur 30 seconds, second failures 60 seconds, third failures 120 seconds, and so on, up to a hard cap of 10 minutes.

### What happens when multiple providers fail simultaneously?

If multiple accounts enter cooldown simultaneously, the selector excludes all of them from the routing pool and attempts to use any remaining healthy accounts. If no accounts are available outside their cooldown windows, the system returns an error indicating no eligible providers exist.

### How is the cooldown duration calculated?

The egress manager in [`backend/internal/infra/egress/manager.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/egress/manager.go) calculates the duration using bitwise left shift on the failure count: `30*time.Second * time.Duration(1<<min(failureCount-1, 4))`. This value is then capped at 10 minutes using the `min` function before being added to the current timestamp.

### How does a provider exit cooldown status?

According to the selector logic in [`backend/internal/application/gateway/selector.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/gateway/selector.go), a provider automatically exits cooldown when the current time passes the stored `cooldown_until` timestamp. Additionally, when a request eventually succeeds, the egress manager clears the failure count and removes the cooldown stamp, fully restoring the account to the healthy routing pool.