# How AxonHub Implements Load Balancing Across Multiple API Keys

> Learn how AxonHub implements load balancing across multiple API keys using round-robin rotation in the authentication provider to balance traffic and avoid rate limits.

- Repository: [Loop/axonhub](https://github.com/looplj/axonhub)
- Tags: internals
- Published: 2026-03-06

---

**AxonHub distributes outbound LLM requests across multiple API keys using a round-robin rotation mechanism implemented in the authentication provider layer, automatically selecting a different key for each request to balance traffic and mitigate rate limits.**

Managing multiple API keys for external LLM providers requires intelligent distribution strategies to avoid rate limiting and ensure high availability. The `looplj/axonhub` repository solves this through a built-in load balancing system that rotates authentication credentials at the provider level. This article examines how AxonHub implements API key load balancing using round-robin selection and provider abstractions.

## The API Key Provider Abstraction

AxonHub decouples authentication logic from transformer implementations through the `auth.APIKeyProvider` interface. This abstraction allows the system to support both single-key and multi-key configurations without changing downstream code.

### Static vs. Rotating Providers

The repository provides two primary implementations:

- **`auth.NewStaticKeyProvider(key string)`** – Returns a single, fixed API key for every request. Used when only one key is configured.
- **`auth.NewRotatingKeyProvider(keys []string)`** – Maintains a slice of keys and rotates through them using round-robin logic. This is the core mechanism enabling load balancing across multiple API keys.

## Round-Robin Load Balancing Implementation

The load balancing logic resides entirely within the authentication layer, ensuring transformers remain agnostic to key management strategies.

### Key Rotation Logic in rotating_key_provider.go

The round-robin mechanism is implemented in [`internal/pkg/auth/rotating_key_provider.go`](https://github.com/looplj/axonhub/blob/main/internal/pkg/auth/rotating_key_provider.go). The provider maintains an internal index that advances with each call to `Get(ctx)`:

```go
// internal/pkg/auth/rotating_key_provider.go
type RotatingKeyProvider struct {
    keys  []string
    index uint64 // atomic counter for thread-safe rotation
}

func (p *RotatingKeyProvider) Get(ctx context.Context) (string, error) {
    // Atomically increment and wrap around
    idx := atomic.AddUint64(&p.index, 1) % uint64(len(p.keys))
    return p.keys[idx], nil
}

```

This implementation ensures **thread-safe round-robin distribution** across concurrent requests. When the index reaches the end of the slice, the modulo operation wraps it back to zero, creating a continuous rotation cycle.

### Per-Request Key Selection

Transformers consume the provider during request construction. In [`llm/transformer/openai/outbound.go`](https://github.com/looplj/axonhub/blob/main/llm/transformer/openai/outbound.go), each outbound HTTP request retrieves a fresh key from the provider:

```go
// llm/transformer/openai/outbound.go
func (t *Transformer) sendRequest(ctx context.Context, payload any) (*http.Response, error) {
    apiKey := t.config.APIKeyProvider.Get(ctx) // ← selects next key in rotation
    
    req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, body)
    if err != nil {
        return nil, err
    }
    
    req.Header.Set("Authorization", "Bearer "+apiKey)
    return t.httpClient.Do(req)
}

```

Because `Get(ctx)` is invoked for every request, **successive calls to the same LLM endpoint automatically use different API keys**, distributing the load evenly across the configured pool.

## Configuring Multiple API Keys in AxonHub

Setting up load balancing requires only YAML configuration and standard bootstrapping code.

### YAML Configuration

Define multiple keys in the configuration file:

```yaml

# conf/conf.yaml

llm:
  openai:
    api_keys:
      - "sk-first-key"
      - "sk-second-key"
      - "sk-third-key"
    model: "gpt-4o"
    timeout: 30s

```

### Bootstrap Implementation

The configuration loader in [`conf/conf.go`](https://github.com/looplj/axonhub/blob/main/conf/conf.go) instantiates the rotating provider when multiple keys are detected:

```go
// conf/conf.go
import "github.com/looplj/axonhub/internal/pkg/auth"

func loadOpenAIConfig(cfg Config) (*openai.Transformer, error) {
    keys := cfg.LLM.OpenAI.APIKeys // []string from YAML
    
    var provider auth.APIKeyProvider
    if len(keys) == 1 {
        provider = auth.NewStaticKeyProvider(keys[0])
    } else {
        provider = auth.NewRotatingKeyProvider(keys) // ← load balancing enabled
    }
    
    return openai.NewTransformer(openai.Config{
        APIKeyProvider: provider,
        Model:          cfg.LLM.OpenAI.Model,
    }), nil
}

```

## Benefits of AxonHub's Load Balancing Approach

The round-robin API key distribution provides several operational advantages:

- **Automatic Rate Limit Mitigation** – By distributing requests across multiple keys, the system naturally avoids hitting per-key rate limits on external providers like OpenAI or Anthropic.
- **Zero-Configuration Fallback** – If a specific key fails due to rate limiting or revocation, the next request automatically uses the subsequent key in the rotation without requiring circuit breaker logic in the transformer.
- **Even Traffic Distribution** – The atomic round-robin counter ensures statistically uniform distribution across the key pool, preventing hot-spotting on any single credential.
- **Provider Agnostic** – The `APIKeyProvider` interface allows the same load balancing logic to work across OpenAI, Anthropic, Gemini, or any other LLM transformer without code changes.

## Summary

AxonHub implements load balancing across multiple API keys through a clean separation of concerns:

- The **`auth.RotatingKeyProvider`** in [`internal/pkg/auth/rotating_key_provider.go`](https://github.com/looplj/axonhub/blob/main/internal/pkg/auth/rotating_key_provider.go) maintains a thread-safe round-robin index over a slice of keys.
- Transformers retrieve keys via the **`Get(ctx)`** method before each HTTP request, ensuring every outbound call potentially uses a different credential.
- Configuration in [`conf/conf.go`](https://github.com/looplj/axonhub/blob/main/conf/conf.go) automatically selects between static and rotating providers based on the number of keys supplied.

This architecture provides automatic traffic distribution and built-in resilience against rate limits without requiring external load balancers or complex retry logic in application code.

## Frequently Asked Questions

### What happens if one API key hits a rate limit?

When a specific API key receives a rate-limit response from the provider, the next outbound request automatically uses the following key in the rotation sequence. Because `auth.RotatingKeyProvider` advances the index on every `Get(ctx)` call, failed keys are naturally skipped for subsequent requests without requiring explicit error handling or circuit breaker configuration in the transformer code.

### Can I mix static and rotating providers for different LLM backends?

Yes. The `auth.APIKeyProvider` interface allows each transformer to use a different provider implementation. In [`conf/conf.go`](https://github.com/looplj/axonhub/blob/main/conf/conf.go), you can configure a static provider for one LLM service (e.g., a single Anthropic key) while using a rotating provider for another (e.g., multiple OpenAI keys). The transformer remains agnostic to the provider type, calling only the standard `Get(ctx)` method.

### Does AxonHub support weighted or priority-based API key selection?

The current implementation in [`internal/pkg/auth/rotating_key_provider.go`](https://github.com/looplj/axonhub/blob/main/internal/pkg/auth/rotating_key_provider.go) uses a simple round-robin algorithm that treats all keys equally. There is no built-in support for weighted distribution or priority tiers in the open-source codebase. To implement priority-based selection, you would need to extend the `APIKeyProvider` interface with a custom implementation that tracks key priorities or usage quotas.

### How does the round-robin mechanism handle concurrent requests?

The `RotatingKeyProvider` uses an atomic counter (`atomic.AddUint64`) to manage the current index position. When multiple goroutines call `Get(ctx)` simultaneously, the atomic operation ensures each request receives a unique key from the pool without race conditions. The modulo operation (`% uint64(len(p.keys))`) wraps the index back to zero after reaching the end of the key slice, creating a continuous rotation cycle safe for high-concurrency workloads.