How AxonHub Implements Load Balancing Across Multiple API Keys
AxonHub distributes outbound LLM requests across multiple API keys using a round-robin rotation mechanism implemented in the authentication provider layer, automatically selecting a different key for each request to balance traffic and mitigate rate limits.
Managing multiple API keys for external LLM providers requires intelligent distribution strategies to avoid rate limiting and ensure high availability. The looplj/axonhub repository solves this through a built-in load balancing system that rotates authentication credentials at the provider level. This article examines how AxonHub implements API key load balancing using round-robin selection and provider abstractions.
The API Key Provider Abstraction
AxonHub decouples authentication logic from transformer implementations through the auth.APIKeyProvider interface. This abstraction allows the system to support both single-key and multi-key configurations without changing downstream code.
Static vs. Rotating Providers
The repository provides two primary implementations:
auth.NewStaticKeyProvider(key string)– Returns a single, fixed API key for every request. Used when only one key is configured.auth.NewRotatingKeyProvider(keys []string)– Maintains a slice of keys and rotates through them using round-robin logic. This is the core mechanism enabling load balancing across multiple API keys.
Round-Robin Load Balancing Implementation
The load balancing logic resides entirely within the authentication layer, ensuring transformers remain agnostic to key management strategies.
Key Rotation Logic in rotating_key_provider.go
The round-robin mechanism is implemented in internal/pkg/auth/rotating_key_provider.go. The provider maintains an internal index that advances with each call to Get(ctx):
// internal/pkg/auth/rotating_key_provider.go
type RotatingKeyProvider struct {
keys []string
index uint64 // atomic counter for thread-safe rotation
}
func (p *RotatingKeyProvider) Get(ctx context.Context) (string, error) {
// Atomically increment and wrap around
idx := atomic.AddUint64(&p.index, 1) % uint64(len(p.keys))
return p.keys[idx], nil
}
This implementation ensures thread-safe round-robin distribution across concurrent requests. When the index reaches the end of the slice, the modulo operation wraps it back to zero, creating a continuous rotation cycle.
Per-Request Key Selection
Transformers consume the provider during request construction. In llm/transformer/openai/outbound.go, each outbound HTTP request retrieves a fresh key from the provider:
// llm/transformer/openai/outbound.go
func (t *Transformer) sendRequest(ctx context.Context, payload any) (*http.Response, error) {
apiKey := t.config.APIKeyProvider.Get(ctx) // ← selects next key in rotation
req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, body)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
return t.httpClient.Do(req)
}
Because Get(ctx) is invoked for every request, successive calls to the same LLM endpoint automatically use different API keys, distributing the load evenly across the configured pool.
Configuring Multiple API Keys in AxonHub
Setting up load balancing requires only YAML configuration and standard bootstrapping code.
YAML Configuration
Define multiple keys in the configuration file:
# conf/conf.yaml
llm:
openai:
api_keys:
- "sk-first-key"
- "sk-second-key"
- "sk-third-key"
model: "gpt-4o"
timeout: 30s
Bootstrap Implementation
The configuration loader in conf/conf.go instantiates the rotating provider when multiple keys are detected:
// conf/conf.go
import "github.com/looplj/axonhub/internal/pkg/auth"
func loadOpenAIConfig(cfg Config) (*openai.Transformer, error) {
keys := cfg.LLM.OpenAI.APIKeys // []string from YAML
var provider auth.APIKeyProvider
if len(keys) == 1 {
provider = auth.NewStaticKeyProvider(keys[0])
} else {
provider = auth.NewRotatingKeyProvider(keys) // ← load balancing enabled
}
return openai.NewTransformer(openai.Config{
APIKeyProvider: provider,
Model: cfg.LLM.OpenAI.Model,
}), nil
}
Benefits of AxonHub's Load Balancing Approach
The round-robin API key distribution provides several operational advantages:
- Automatic Rate Limit Mitigation – By distributing requests across multiple keys, the system naturally avoids hitting per-key rate limits on external providers like OpenAI or Anthropic.
- Zero-Configuration Fallback – If a specific key fails due to rate limiting or revocation, the next request automatically uses the subsequent key in the rotation without requiring circuit breaker logic in the transformer.
- Even Traffic Distribution – The atomic round-robin counter ensures statistically uniform distribution across the key pool, preventing hot-spotting on any single credential.
- Provider Agnostic – The
APIKeyProviderinterface allows the same load balancing logic to work across OpenAI, Anthropic, Gemini, or any other LLM transformer without code changes.
Summary
AxonHub implements load balancing across multiple API keys through a clean separation of concerns:
- The
auth.RotatingKeyProviderininternal/pkg/auth/rotating_key_provider.gomaintains a thread-safe round-robin index over a slice of keys. - Transformers retrieve keys via the
Get(ctx)method before each HTTP request, ensuring every outbound call potentially uses a different credential. - Configuration in
conf/conf.goautomatically selects between static and rotating providers based on the number of keys supplied.
This architecture provides automatic traffic distribution and built-in resilience against rate limits without requiring external load balancers or complex retry logic in application code.
Frequently Asked Questions
What happens if one API key hits a rate limit?
When a specific API key receives a rate-limit response from the provider, the next outbound request automatically uses the following key in the rotation sequence. Because auth.RotatingKeyProvider advances the index on every Get(ctx) call, failed keys are naturally skipped for subsequent requests without requiring explicit error handling or circuit breaker configuration in the transformer code.
Can I mix static and rotating providers for different LLM backends?
Yes. The auth.APIKeyProvider interface allows each transformer to use a different provider implementation. In conf/conf.go, you can configure a static provider for one LLM service (e.g., a single Anthropic key) while using a rotating provider for another (e.g., multiple OpenAI keys). The transformer remains agnostic to the provider type, calling only the standard Get(ctx) method.
Does AxonHub support weighted or priority-based API key selection?
The current implementation in internal/pkg/auth/rotating_key_provider.go uses a simple round-robin algorithm that treats all keys equally. There is no built-in support for weighted distribution or priority tiers in the open-source codebase. To implement priority-based selection, you would need to extend the APIKeyProvider interface with a custom implementation that tracks key priorities or usage quotas.
How does the round-robin mechanism handle concurrent requests?
The RotatingKeyProvider uses an atomic counter (atomic.AddUint64) to manage the current index position. When multiple goroutines call Get(ctx) simultaneously, the atomic operation ensures each request receives a unique key from the pool without race conditions. The modulo operation (% uint64(len(p.keys))) wraps the index back to zero after reaching the end of the key slice, creating a continuous rotation cycle safe for high-concurrency workloads.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →