How AxonHub Implements Load Balancing Across Multiple API Keys

AxonHub distributes outbound LLM requests across multiple API keys using a round-robin rotation mechanism implemented in the authentication provider layer, automatically selecting a different key for each request to balance traffic and mitigate rate limits.

Managing multiple API keys for external LLM providers requires intelligent distribution strategies to avoid rate limiting and ensure high availability. The looplj/axonhub repository solves this through a built-in load balancing system that rotates authentication credentials at the provider level. This article examines how AxonHub implements API key load balancing using round-robin selection and provider abstractions.

The API Key Provider Abstraction

AxonHub decouples authentication logic from transformer implementations through the auth.APIKeyProvider interface. This abstraction allows the system to support both single-key and multi-key configurations without changing downstream code.

Static vs. Rotating Providers

The repository provides two primary implementations:

  • auth.NewStaticKeyProvider(key string) – Returns a single, fixed API key for every request. Used when only one key is configured.
  • auth.NewRotatingKeyProvider(keys []string) – Maintains a slice of keys and rotates through them using round-robin logic. This is the core mechanism enabling load balancing across multiple API keys.

Round-Robin Load Balancing Implementation

The load balancing logic resides entirely within the authentication layer, ensuring transformers remain agnostic to key management strategies.

Key Rotation Logic in rotating_key_provider.go

The round-robin mechanism is implemented in internal/pkg/auth/rotating_key_provider.go. The provider maintains an internal index that advances with each call to Get(ctx):

// internal/pkg/auth/rotating_key_provider.go
type RotatingKeyProvider struct {
    keys  []string
    index uint64 // atomic counter for thread-safe rotation
}

func (p *RotatingKeyProvider) Get(ctx context.Context) (string, error) {
    // Atomically increment and wrap around
    idx := atomic.AddUint64(&p.index, 1) % uint64(len(p.keys))
    return p.keys[idx], nil
}

This implementation ensures thread-safe round-robin distribution across concurrent requests. When the index reaches the end of the slice, the modulo operation wraps it back to zero, creating a continuous rotation cycle.

Per-Request Key Selection

Transformers consume the provider during request construction. In llm/transformer/openai/outbound.go, each outbound HTTP request retrieves a fresh key from the provider:

// llm/transformer/openai/outbound.go
func (t *Transformer) sendRequest(ctx context.Context, payload any) (*http.Response, error) {
    apiKey := t.config.APIKeyProvider.Get(ctx) // ← selects next key in rotation
    
    req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, body)
    if err != nil {
        return nil, err
    }
    
    req.Header.Set("Authorization", "Bearer "+apiKey)
    return t.httpClient.Do(req)
}

Because Get(ctx) is invoked for every request, successive calls to the same LLM endpoint automatically use different API keys, distributing the load evenly across the configured pool.

Configuring Multiple API Keys in AxonHub

Setting up load balancing requires only YAML configuration and standard bootstrapping code.

YAML Configuration

Define multiple keys in the configuration file:


# conf/conf.yaml

llm:
  openai:
    api_keys:
      - "sk-first-key"
      - "sk-second-key"
      - "sk-third-key"
    model: "gpt-4o"
    timeout: 30s

Bootstrap Implementation

The configuration loader in conf/conf.go instantiates the rotating provider when multiple keys are detected:

// conf/conf.go
import "github.com/looplj/axonhub/internal/pkg/auth"

func loadOpenAIConfig(cfg Config) (*openai.Transformer, error) {
    keys := cfg.LLM.OpenAI.APIKeys // []string from YAML
    
    var provider auth.APIKeyProvider
    if len(keys) == 1 {
        provider = auth.NewStaticKeyProvider(keys[0])
    } else {
        provider = auth.NewRotatingKeyProvider(keys) // ← load balancing enabled
    }
    
    return openai.NewTransformer(openai.Config{
        APIKeyProvider: provider,
        Model:          cfg.LLM.OpenAI.Model,
    }), nil
}

Benefits of AxonHub's Load Balancing Approach

The round-robin API key distribution provides several operational advantages:

  • Automatic Rate Limit Mitigation – By distributing requests across multiple keys, the system naturally avoids hitting per-key rate limits on external providers like OpenAI or Anthropic.
  • Zero-Configuration Fallback – If a specific key fails due to rate limiting or revocation, the next request automatically uses the subsequent key in the rotation without requiring circuit breaker logic in the transformer.
  • Even Traffic Distribution – The atomic round-robin counter ensures statistically uniform distribution across the key pool, preventing hot-spotting on any single credential.
  • Provider Agnostic – The APIKeyProvider interface allows the same load balancing logic to work across OpenAI, Anthropic, Gemini, or any other LLM transformer without code changes.

Summary

AxonHub implements load balancing across multiple API keys through a clean separation of concerns:

  • The auth.RotatingKeyProvider in internal/pkg/auth/rotating_key_provider.go maintains a thread-safe round-robin index over a slice of keys.
  • Transformers retrieve keys via the Get(ctx) method before each HTTP request, ensuring every outbound call potentially uses a different credential.
  • Configuration in conf/conf.go automatically selects between static and rotating providers based on the number of keys supplied.

This architecture provides automatic traffic distribution and built-in resilience against rate limits without requiring external load balancers or complex retry logic in application code.

Frequently Asked Questions

What happens if one API key hits a rate limit?

When a specific API key receives a rate-limit response from the provider, the next outbound request automatically uses the following key in the rotation sequence. Because auth.RotatingKeyProvider advances the index on every Get(ctx) call, failed keys are naturally skipped for subsequent requests without requiring explicit error handling or circuit breaker configuration in the transformer code.

Can I mix static and rotating providers for different LLM backends?

Yes. The auth.APIKeyProvider interface allows each transformer to use a different provider implementation. In conf/conf.go, you can configure a static provider for one LLM service (e.g., a single Anthropic key) while using a rotating provider for another (e.g., multiple OpenAI keys). The transformer remains agnostic to the provider type, calling only the standard Get(ctx) method.

Does AxonHub support weighted or priority-based API key selection?

The current implementation in internal/pkg/auth/rotating_key_provider.go uses a simple round-robin algorithm that treats all keys equally. There is no built-in support for weighted distribution or priority tiers in the open-source codebase. To implement priority-based selection, you would need to extend the APIKeyProvider interface with a custom implementation that tracks key priorities or usage quotas.

How does the round-robin mechanism handle concurrent requests?

The RotatingKeyProvider uses an atomic counter (atomic.AddUint64) to manage the current index position. When multiple goroutines call Get(ctx) simultaneously, the atomic operation ensures each request receives a unique key from the pool without race conditions. The modulo operation (% uint64(len(p.keys))) wraps the index back to zero after reaching the end of the key slice, creating a continuous rotation cycle safe for high-concurrency workloads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →