How Cooldown Periods Drive Provider Failover in Grok2API
Cooldown periods temporarily remove failing providers from the routing pool, forcing automatic failover to healthy accounts while preventing rapid retry loops.
In the chenyme/grok2api project, each provider account tracks a cooldown_until timestamp that governs eligibility for request routing. When a request fails, the egress manager calculates a back-off interval and stores it via AccountRepository.UpdateHealth, effectively isolating the misbehaving provider until recovery is likely. This mechanism ensures that the selector logic automatically routes traffic to alternative providers while the failed account remains excluded from the pool.
The Cooldown State in Account Records
Every provider account in Grok2API maintains a cooldown_until field that determines its availability for routing. The egress manager writes this timestamp after detecting failures, while the selector reads it before choosing an account for incoming requests. This simple timestamp acts as a circuit breaker that prevents the system from hammering unstable providers.
Exponential Backoff Calculation
When a failure occurs, the egress manager calculates the cooldown duration using an exponential backoff algorithm capped at ten minutes. The implementation in backend/internal/infra/egress/manager.go uses the failure count to determine the interval, ensuring that repeated failures result in progressively longer isolation periods.
// backend/internal/infra/egress/manager.go
cooldown := min(10*time.Minute,
30*time.Second*time.Duration(1<<min(value.FailureCount-1, 4)))
until := now.Add(cooldown)
repo.UpdateHealth(..., &until, ..., false)
The calculation doubles the base interval of 30 seconds for each consecutive failure, up to a maximum of 10 minutes. This prevents temporary glitches from triggering long outages while protecting against persistent failures.
How the Selector Implements Failover
The selector component in backend/internal/application/gateway/selector.go queries the account repository and filters out any accounts where cooldown_until is in the future. This exclusion happens automatically, causing the router to pick the next available healthy provider without requiring explicit failover configuration.
func (m *Manager) handleFailure(ctx context.Context, accID uint64, failureCount int) {
now := time.Now().UTC()
// exponential back‑off, capped at 10 min
cooldown := min(10*time.Minute,
30*time.Second*time.Duration(1<<min(failureCount-1, 4)))
until := now.Add(cooldown)
_ = m.accountRepo.UpdateHealth(ctx, accID, failureCount, &until,
"request failed", false)
}
When the selector encounters an account still within its cooldown window, it treats that provider as unavailable and skips to the next candidate in the pool.
The Complete Failover Flow
Understanding how cooldown periods affect provider failover requires following the request lifecycle through the system:
- Initial Selection – The selector picks an enabled, authenticated account whose
cooldown_untilis nil or in the past. - Failure Detection – When a request fails, the egress manager records the failure, increments
failure_count, and calculates a newcooldown_untiltimestamp. - Pool Exclusion – Subsequent requests trigger the selector’s query, which excludes any account with
cooldown_until > now. - Automatic Failover – Because the failing provider is hidden from the routing pool, the selector automatically falls back to the next healthiest account, which may be a different provider tier or another account in the same tier.
- Recovery – When
cooldown_untilexpires, the account becomes eligible again. A successful request clears the failure count and removes the cooldown stamp entirely.
This flow ensures that cooldown periods directly drive failover behavior by manipulating the visibility of providers in the routing pool.
Configuration and Monitoring
You can inspect current cooldown settings through the HTTP settings handler exposed in backend/internal/transport/http/settings/handler.go. The configuration values cooldownBase (30 seconds) and cooldownMax (10 minutes) are defined in backend/internal/infra/config/config.go and determine the bounds of the exponential backoff calculation. These settings allow operators to tune how aggressively the system isolates failing providers based on their specific reliability requirements.
Summary
- Cooldown periods are stored as
cooldown_untiltimestamps in account records and act as temporary exclusion flags for the routing selector. - Exponential backoff calculates cooldown durations starting at 30 seconds and doubling up to a 10-minute maximum, as implemented in
backend/internal/infra/egress/manager.go. - Automatic failover occurs because
backend/internal/application/gateway/selector.gofilters out accounts with active cooldowns, forcing selection of healthier providers. - Recovery happens automatically when the cooldown timestamp expires, with successful requests clearing the failure history and restoring full eligibility.
Frequently Asked Questions
How long does a provider stay in cooldown?
The duration depends on the failure count using an exponential backoff formula: min(10 minutes, 30 seconds × 2^(failureCount-1)). First failures incur 30 seconds, second failures 60 seconds, third failures 120 seconds, and so on, up to a hard cap of 10 minutes.
What happens when multiple providers fail simultaneously?
If multiple accounts enter cooldown simultaneously, the selector excludes all of them from the routing pool and attempts to use any remaining healthy accounts. If no accounts are available outside their cooldown windows, the system returns an error indicating no eligible providers exist.
How is the cooldown duration calculated?
The egress manager in backend/internal/infra/egress/manager.go calculates the duration using bitwise left shift on the failure count: 30*time.Second * time.Duration(1<<min(failureCount-1, 4)). This value is then capped at 10 minutes using the min function before being added to the current timestamp.
How does a provider exit cooldown status?
According to the selector logic in backend/internal/application/gateway/selector.go, a provider automatically exits cooldown when the current time passes the stored cooldown_until timestamp. Additionally, when a request eventually succeeds, the egress manager clears the failure count and removes the cooldown stamp, fully restoring the account to the healthy routing pool.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →