How Channel Health Probing and Auto-Disable Work in AxonHub
AxonHub protects downstream LLM providers by continuously monitoring channel health through circuit-breaker-style probing that tests recovery after failures, while auto-disable permanently removes channels or API keys after repeated specific HTTP errors.
The looplj/axonhub repository implements a dual-layer resilience strategy to prevent cascading failures when proxying requests to large language model APIs. Channel health probing and auto-disable work at different granularities—one handles transient outages through automated recovery testing, while the other handles persistent authentication or rate-limit issues through administrative shutdown.
Understanding Channel Health Probing
Channel health probing operates on individual models within a channel using a circuit-breaker pattern. When a model exceeds failure thresholds, AxonHub enters a probationary state where it allows exactly one test request after a cooldown period to verify recovery.
State Tracking and Failure Thresholds
Each model maintains its health state in the ModelCircuitBreakerStats struct defined in internal/server/biz/model_circuit_breaker.go. The struct tracks State (closed, half_open, or open), ConsecutiveFailures, and NextProbeAt.
When RecordError detects that ConsecutiveFailures exceeds the OpenThreshold (default 5), it transitions the state to StateOpen and schedules the next probe:
// internal/server/biz/model_circuit_breaker.go
if stats.ConsecutiveFailures >= policy.OpenThreshold {
stats.State = StateOpen
stats.NextProbeAt = time.Now().Add(policy.ProbeInterval) // default 5 minutes
}
The Probing Window
While a model is in StateOpen, GetEffectiveWeight returns 0.0 to the load balancer, dropping all traffic. However, once time.Now() exceeds NextProbeAt and no probe is currently in progress, the weight temporarily increases to baseWeight * HalfOpenWeight (default 0.3), allowing exactly one request through:
// internal/server/biz/model_circuit_breaker.go:88-96
if time.Now().After(stats.NextProbeAt) {
if atomic.LoadInt32(&stats.probingInProgress) == 0 {
return baseWeight * policy.HalfOpenWeight // allow single probe
}
}
return 0.0 // otherwise drop traffic
Probe Execution and State Recovery
The orchestrator in internal/server/orchestrator/model_circuit_breaker.go manages the probe lifecycle. TryBeginProbe atomically sets probingInProgress from 0 to 1, ensuring only one request becomes the probe even under concurrent load:
// internal/server/orchestrator/model_circuit_breaker.go:2-16
func (m *ModelCircuitBreaker) TryBeginProbe(ctx context.Context, channelID int, modelID string) bool {
// ... state checks ...
return atomic.CompareAndSwapInt32(&stats.probingInProgress, 0, 1)
}
When the probe request completes successfully, RecordSuccess resets the circuit breaker to StateClosed, clears ConsecutiveFailures, and removes the probe schedule, immediately restoring full traffic capacity.
How Auto-Disable Protects Your Infrastructure
While probing handles transient failures, auto-disable handles persistent error patterns that indicate configuration issues (invalid API keys) or upstream rate limits. After a configurable number of consecutive errors with a specific HTTP status code, AxonHub permanently disables the channel or API key until manual operator intervention.
Configurable Error Thresholds
The RetryPolicy struct in internal/server/biz/system.go defines AutoDisableChannel, which maps HTTP status codes to failure thresholds:
// internal/server/biz/system.go
type AutoDisableChannel struct {
Enabled bool
Statuses []AutoDisableChannelStatus // e.g., {Status: 401, Times: 3}
}
Per-Status Error Counting
The ChannelService in internal/server/biz/channel_auto_disable.go maintains in-memory counters (channelErrorCounts and apiKeyErrorCounts) that track consecutive errors per status code. The checkAndHandleChannelError method increments counters when a request fails:
// internal/server/biz/channel_auto_disable.go:47-55
svc.channelErrorCounts[perf.ChannelID][perf.ErrorStatusCode]++
count := svc.channelErrorCounts[perf.ChannelID][perf.ErrorStatusCode]
Channel and API Key Disablement
When the error count reaches the configured Times threshold, markChannelUnavailable executes a database update setting channel.Status = StatusDisabled, effectively removing the channel from the load-balancer pool:
// internal/server/biz/channel_auto_disable.go:58-63
if count >= statusConfig.Times {
svc.markChannelUnavailable(ctx, perf.ChannelID, perf.ErrorStatusCode)
// reset counters to prevent duplicate disablement
}
The same logic applies to API keys via DisableAPIKey, which generates a reason string such as "Auto-disabled after 3 consecutive errors with status 401" for audit trails.
Interaction Between Probing and Auto-Disable
These mechanisms operate at different layers and time scales:
- Probing functions at the model level within
internal/server/biz/model_circuit_breaker.go. It activates only after a channel has entered theOpenstate due to general failures, allowing temporary recovery without administrative action. - Auto-disable functions at the channel or API-key level within
internal/server/biz/channel_auto_disable.go. It triggers on specific HTTP status codes (401, 429, 500) that indicate non-transient issues, permanently removing the resource from rotation.
If a channel is auto-disabled, it is removed from the load-balancer entirely, so the probing mechanism never executes for that channel's models. Conversely, if a channel experiences transient 500 errors that trigger the circuit breaker but not the auto-disable threshold, probing provides the recovery path.
Configuration Examples
Enabling Auto-Disable in YAML
Configure specific status codes and thresholds in your AxonHub configuration file:
# config.example.yml
retry_policy:
auto_disable_channel:
enabled: true
statuses:
- status: 401 # Unauthorized - likely invalid API key
times: 3 # Disable after 3 consecutive 401s
- status: 429 # Rate limited
times: 5 # Allow brief spikes before disabling
- status: 500 # Internal server error
times: 5
Manual Probe Testing
For administrative endpoints or testing scenarios, you can manually trigger a probe:
// Testing probe initiation
ctx := context.Background()
svc := channelService // from dependency injection
// Attempt to begin probe for model "gpt-4" on channel 42
if svc.modelCircuitBreaker.TryBeginProbe(ctx, 42, "gpt-4") {
// Execute single probe request
resp, err := llmClient.Ping(ctx, "gpt-4")
if err == nil && resp.Healthy {
svc.modelCircuitBreaker.RecordSuccess(ctx, 42, "gpt-4")
fmt.Println("Channel recovered and returned to service")
} else {
svc.modelCircuitBreaker.RecordError(ctx, 42, "gpt-4")
fmt.Println("Probe failed, channel remains in Open state")
}
}
Summary
- Channel health probing implements a circuit-breaker pattern at the model level, using
internal/server/biz/model_circuit_breaker.goto track failures and allow single probe requests after a cooldown period. - Auto-disable provides permanent protection by monitoring specific HTTP status codes in
internal/server/biz/channel_auto_disable.go, disabling channels or API keys after configurable consecutive error thresholds. - Probing uses atomic operations (
atomic.CompareAndSwapInt32) to ensure exactly one probe request executes concurrently, preventing thundering herds during recovery attempts. - The two systems operate hierarchically: auto-disable removes channels from the load balancer entirely, while probing manages traffic distribution for channels that remain active but experience transient failures.
Frequently Asked Questions
What is the difference between channel health probing and auto-disable in AxonHub?
Channel health probing is a temporary, automated recovery mechanism that tests whether a failed model has healed after a cooldown period, using a circuit-breaker pattern to allow single probe requests. Auto-disable is a permanent administrative action that removes a channel or API key from service after repeated specific HTTP errors (such as 401 or 429) indicate configuration or rate-limit issues that require human intervention.
How does AxonHub prevent multiple simultaneous probe requests to a recovering channel?
AxonHub uses atomic integer operations in internal/server/orchestrator/model_circuit_breaker.go to ensure only one probe executes at a time. The TryBeginProbe function calls atomic.CompareAndSwapInt32 to atomically flip the probingInProgress flag from 0 to 1. If the swap succeeds, the current request becomes the probe; if another goroutine has already set the flag, the function returns false and the request is treated as normal traffic (which gets dropped while the channel is Open).
What HTTP status codes trigger auto-disable, and how are thresholds configured?
Auto-disable triggers on any HTTP status codes explicitly configured in the RetryPolicy.AutoDisableChannel.Statuses array, commonly including 401 (unauthorized), 429 (rate limited), and 500 (internal server error). Thresholds are configured via YAML in config.example.yml by specifying the status and times fields—for example, setting status: 401 with times: 3 disables the channel after three consecutive 401 errors. The ChannelService in internal/server/biz/channel_auto_disable.go maintains per-channel counters to track these consecutive failures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →