How to Handle Provider Quota Limits with Automatic Fallback in AxonHub

AxonHub handles provider quota limits by isolating quota pools per model and automatically failing over to alternative endpoints or pools when quotas are exhausted or errors occur.

AxonHub is an open-source LLM gateway that manages multiple provider backends through a sophisticated quota and routing system. When working with providers like Antigravity and Gemini-CLI, you need robust mechanisms to handle rate limits and quota exhaustion without interrupting service. The codebase implements a three-layer fallback strategy that operates at the model, endpoint, and API-key levels.

Understanding AxonHub's Quota Architecture

Quota Pools and Model Naming

AxonHub maintains separate quota pools for different providers. The system recognizes two primary pools: Antigravity and Gemini-CLI. The routing decision happens in llm/transformer/antigravity/router.go through the DetermineQuotaPreference function, which parses model names to select the appropriate pool.

Model names can explicitly specify their quota pool using:

  • Suffixes: :antigravity or :gemini-cli
  • Prefixes: antigravity-

The Three-Layer Fallback Strategy

The fallback mechanism operates through three distinct layers:

  1. API-Key Quota Enforcement: Middleware checks if your API key has remaining quota for the selected pool before the request reaches the transformer
  2. Endpoint Fallback: Within the Antigravity pool, the system cycles through three sandboxes (Daily, Autopush, Prod) when encountering retryable errors
  3. Pool-Level Fallback: By configuring dual-model entries, you can fall back from Gemini-CLI to Antigravity pools when the former is exhausted

Implementing Model-Level Quota Selection

The DetermineQuotaPreference function in llm/transformer/antigravity/router.go implements sophisticated logic to route requests to the correct quota pool. It checks for explicit suffixes, prefixes, and model-specific rules.

// router.go
pref := DetermineQuotaPreference(modelName) // antigravity, gemini-cli or default
initial := GetInitialEndpoint(pref)       // always starts with Daily

The function evaluates model names in this priority order:

  1. Explicit suffix :antigravity or :gemini-cli
  2. Explicit prefix antigravity-
  3. Model-specific rules (Claude, GPT-OSS, and image models route to Antigravity)
  4. Legacy Gemini-3 names route to Antigravity
  5. Default to Gemini-CLI quota (for Gemini 2.5/3-preview models)

Configuring Automatic Endpoint Fallback

The Sandbox Priority Order

Within the Antigravity quota pool, AxonHub maintains three distinct endpoints that serve as automatic fallback targets. The GetFallbackEndpoints function in llm/transformer/antigravity/executor.go returns these in priority order:

return []string{
    EndpointDaily,    // newest features, usually the most generous quota
    EndpointAutopush, // staging sandbox
    EndpointProd,     // stable production sandbox
}

The getEndpointsInOrder method constructs the final execution sequence by placing the preferred endpoint first while maintaining the fallback chain:

func (e *Executor) getEndpointsInOrder(modelName string) []string {
    if modelName == "" { return GetFallbackEndpoints() }
    pref := DetermineQuotaPreference(modelName)
    init := GetInitialEndpoint(pref)
    fallbacks := GetFallbackEndpoints()
    // put the preferred endpoint first, keep the rest in the original order
}

Health Tracking and Cooldown Logic

The health tracker in llm/transformer/antigravity/health_tracker.go implements intelligent cooldown mechanisms to prevent hammering failing endpoints. It tracks HTTP status codes and implements lazy TTL-based expiration.

tracker := antigravity.NewAntigravityHealthTracker()
tracker.RecordFailure("gemini-2.5-pro", antigravity.EndpointDaily, 429)

// Later, before the next request:
if tracker.ShouldSkip("gemini-2.5-pro", antigravity.EndpointDaily) {
    // Daily is still in cooldown → executor will try Autopush first
}

The tracker considers these status codes as retryable triggers for cooldown: 429 (Rate Limit), 403 (Forbidden), 404 (Not Found), and 5xx (Server Errors). Success responses immediately clear any existing cooldown for that endpoint-model pair.

Enforcing API Key Quota Limits

Before requests reach the transformer layer, the enforceQuota middleware in internal/server/orchestrator/quota.go validates API key quotas against the selected provider pool. This prevents wasted calls to exhausted endpoints.

result, err := quotaService.CheckAPIKeyQuota(ctx, apiKey.ID, profile.Quota)
if !result.Allowed {
    return nil, &llm.ResponseError{
        StatusCode: http.StatusForbidden,
        Detail: llm.ErrorDetail{
            Code: "quota_exceeded",
            Type: "quota_exceeded_error",
            Message: result.Message,
        },
    }
}

If the quota service returns Allowed: false, the request fails fast with a 403 Quota Exceeded error before any endpoint fallback logic executes. This ensures that quota enforcement happens at the API key level, independent of the provider's own rate limiting.

Dual-Quota Pattern for Pool Fallback

To achieve automatic fallback between quota pools (e.g., from Gemini-CLI to Antigravity when the former exhausts), AxonHub supports a dual-model configuration pattern. By declaring the same model twice with different quota specifiers, clients can retry across pools without code changes.

Configuration Example


# config.yml (excerpt)

providers:
  antigravity:
    models:
      - gemini-2.5-pro           # uses Gemini‑CLI quota (default)

      - antigravity-gemini-2.5-pro   # uses Antigravity quota

As documented in docs/en/guides/antigravity.md, adding the antigravity- prefix enables access to both quota pools for the same underlying model.

Client-Side Pool Selection

Clients can force a specific quota pool at request time using suffix notation:

req := axonhub.NewChatCompletionRequest()
req.Model = "gemini-2.5-pro:antigravity" // force Antigravity quota
resp, err := client.ChatCompletion(context.Background(), req)

The :antigravity suffix forces DetermineQuotaPreference to return QuotaAntigravity, bypassing the default routing logic.

Summary

  • AxonHub isolates quota pools per provider (Antigravity vs Gemini-CLI) and selects pools based on model name prefixes, suffixes, or implicit rules defined in router.go.
  • Automatic endpoint fallback cycles through Daily, Autopush, and Prod sandboxes when encountering retryable HTTP errors (429, 403, 404, 5xx), implemented in executor.go.
  • Health tracking maintains per-model cooldown periods for failed endpoints using lazy TTL expiration in health_tracker.go.
  • API key quota enforcement occurs before transformer execution via middleware in quota.go, returning 403 errors immediately when limits are reached.
  • Dual-model configuration allows automatic pool-level fallback by declaring models with both default and antigravity- prefixed names, enabling clients to retry across quota pools without application changes.

Frequently Asked Questions

What happens when all Antigravity endpoints are in cooldown?

If Daily, Autopush, and Prod are all in cooldown for a specific model, the executor in llm/transformer/antigravity/executor.go will still attempt the request against the first endpoint in the list (Daily) if no other options remain. However, if you have configured dual-model entries for both Antigravity and Gemini-CLI pools, you should retry with the Gemini-CLI variant to bypass the exhausted Antigravity infrastructure entirely.

How long do endpoint cooldowns last in AxonHub?

The health tracker in llm/transformer/antigravity/health_tracker.go uses a default TTL of 10 minutes for failure records. After this period, the cooldown entry is lazily expired and the endpoint becomes eligible for requests again. Successful requests to an endpoint immediately clear its cooldown status for that specific model.

Can I force a specific quota pool without modifying configuration?

Yes, you can force quota pool selection at request time using naming conventions. Append :antigravity or :gemini-cli to the model name to override the default routing logic implemented in llm/transformer/antigravity/router.go. Alternatively, use the antigravity- prefix (e.g., antigravity-gemini-2.5-pro) to achieve the same result without changing your YAML configuration files.

What HTTP status codes trigger automatic endpoint fallback?

AxonHub treats the following status codes as retryable errors that trigger endpoint fallback: 429 (Too Many Requests), 403 (Forbidden), 404 (Not Found), and any 5xx server error. When the executor in llm/transformer/antigravity/executor.go encounters these codes, it records the failure in the health tracker and automatically attempts the request against the next available endpoint in the fallback chain (Daily → Autopush → Prod).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →