How Channel Health Probing and Auto-Disable Work in AxonHub

AxonHub protects downstream LLM providers by continuously monitoring channel health through circuit-breaker-style probing that tests recovery after failures, while auto-disable permanently removes channels or API keys after repeated specific HTTP errors.

The looplj/axonhub repository implements a dual-layer resilience strategy to prevent cascading failures when proxying requests to large language model APIs. Channel health probing and auto-disable work at different granularities—one handles transient outages through automated recovery testing, while the other handles persistent authentication or rate-limit issues through administrative shutdown.

Understanding Channel Health Probing

Channel health probing operates on individual models within a channel using a circuit-breaker pattern. When a model exceeds failure thresholds, AxonHub enters a probationary state where it allows exactly one test request after a cooldown period to verify recovery.

State Tracking and Failure Thresholds

Each model maintains its health state in the ModelCircuitBreakerStats struct defined in internal/server/biz/model_circuit_breaker.go. The struct tracks State (closed, half_open, or open), ConsecutiveFailures, and NextProbeAt.

When RecordError detects that ConsecutiveFailures exceeds the OpenThreshold (default 5), it transitions the state to StateOpen and schedules the next probe:

// internal/server/biz/model_circuit_breaker.go
if stats.ConsecutiveFailures >= policy.OpenThreshold {
    stats.State = StateOpen
    stats.NextProbeAt = time.Now().Add(policy.ProbeInterval) // default 5 minutes
}

The Probing Window

While a model is in StateOpen, GetEffectiveWeight returns 0.0 to the load balancer, dropping all traffic. However, once time.Now() exceeds NextProbeAt and no probe is currently in progress, the weight temporarily increases to baseWeight * HalfOpenWeight (default 0.3), allowing exactly one request through:

// internal/server/biz/model_circuit_breaker.go:88-96
if time.Now().After(stats.NextProbeAt) {
    if atomic.LoadInt32(&stats.probingInProgress) == 0 {
        return baseWeight * policy.HalfOpenWeight // allow single probe
    }
}
return 0.0 // otherwise drop traffic

Probe Execution and State Recovery

The orchestrator in internal/server/orchestrator/model_circuit_breaker.go manages the probe lifecycle. TryBeginProbe atomically sets probingInProgress from 0 to 1, ensuring only one request becomes the probe even under concurrent load:

// internal/server/orchestrator/model_circuit_breaker.go:2-16
func (m *ModelCircuitBreaker) TryBeginProbe(ctx context.Context, channelID int, modelID string) bool {
    // ... state checks ...
    return atomic.CompareAndSwapInt32(&stats.probingInProgress, 0, 1)
}

When the probe request completes successfully, RecordSuccess resets the circuit breaker to StateClosed, clears ConsecutiveFailures, and removes the probe schedule, immediately restoring full traffic capacity.

How Auto-Disable Protects Your Infrastructure

While probing handles transient failures, auto-disable handles persistent error patterns that indicate configuration issues (invalid API keys) or upstream rate limits. After a configurable number of consecutive errors with a specific HTTP status code, AxonHub permanently disables the channel or API key until manual operator intervention.

Configurable Error Thresholds

The RetryPolicy struct in internal/server/biz/system.go defines AutoDisableChannel, which maps HTTP status codes to failure thresholds:

// internal/server/biz/system.go
type AutoDisableChannel struct {
    Enabled  bool
    Statuses []AutoDisableChannelStatus // e.g., {Status: 401, Times: 3}
}

Per-Status Error Counting

The ChannelService in internal/server/biz/channel_auto_disable.go maintains in-memory counters (channelErrorCounts and apiKeyErrorCounts) that track consecutive errors per status code. The checkAndHandleChannelError method increments counters when a request fails:

// internal/server/biz/channel_auto_disable.go:47-55
svc.channelErrorCounts[perf.ChannelID][perf.ErrorStatusCode]++
count := svc.channelErrorCounts[perf.ChannelID][perf.ErrorStatusCode]

Channel and API Key Disablement

When the error count reaches the configured Times threshold, markChannelUnavailable executes a database update setting channel.Status = StatusDisabled, effectively removing the channel from the load-balancer pool:

// internal/server/biz/channel_auto_disable.go:58-63
if count >= statusConfig.Times {
    svc.markChannelUnavailable(ctx, perf.ChannelID, perf.ErrorStatusCode)
    // reset counters to prevent duplicate disablement
}

The same logic applies to API keys via DisableAPIKey, which generates a reason string such as "Auto-disabled after 3 consecutive errors with status 401" for audit trails.

Interaction Between Probing and Auto-Disable

These mechanisms operate at different layers and time scales:

  • Probing functions at the model level within internal/server/biz/model_circuit_breaker.go. It activates only after a channel has entered the Open state due to general failures, allowing temporary recovery without administrative action.
  • Auto-disable functions at the channel or API-key level within internal/server/biz/channel_auto_disable.go. It triggers on specific HTTP status codes (401, 429, 500) that indicate non-transient issues, permanently removing the resource from rotation.

If a channel is auto-disabled, it is removed from the load-balancer entirely, so the probing mechanism never executes for that channel's models. Conversely, if a channel experiences transient 500 errors that trigger the circuit breaker but not the auto-disable threshold, probing provides the recovery path.

Configuration Examples

Enabling Auto-Disable in YAML

Configure specific status codes and thresholds in your AxonHub configuration file:


# config.example.yml

retry_policy:
  auto_disable_channel:
    enabled: true
    statuses:
      - status: 401   # Unauthorized - likely invalid API key

        times: 3      # Disable after 3 consecutive 401s

      - status: 429   # Rate limited

        times: 5      # Allow brief spikes before disabling

      - status: 500   # Internal server error

        times: 5

Manual Probe Testing

For administrative endpoints or testing scenarios, you can manually trigger a probe:

// Testing probe initiation
ctx := context.Background()
svc := channelService // from dependency injection

// Attempt to begin probe for model "gpt-4" on channel 42
if svc.modelCircuitBreaker.TryBeginProbe(ctx, 42, "gpt-4") {
    // Execute single probe request
    resp, err := llmClient.Ping(ctx, "gpt-4")
    if err == nil && resp.Healthy {
        svc.modelCircuitBreaker.RecordSuccess(ctx, 42, "gpt-4")
        fmt.Println("Channel recovered and returned to service")
    } else {
        svc.modelCircuitBreaker.RecordError(ctx, 42, "gpt-4")
        fmt.Println("Probe failed, channel remains in Open state")
    }
}

Summary

  • Channel health probing implements a circuit-breaker pattern at the model level, using internal/server/biz/model_circuit_breaker.go to track failures and allow single probe requests after a cooldown period.
  • Auto-disable provides permanent protection by monitoring specific HTTP status codes in internal/server/biz/channel_auto_disable.go, disabling channels or API keys after configurable consecutive error thresholds.
  • Probing uses atomic operations (atomic.CompareAndSwapInt32) to ensure exactly one probe request executes concurrently, preventing thundering herds during recovery attempts.
  • The two systems operate hierarchically: auto-disable removes channels from the load balancer entirely, while probing manages traffic distribution for channels that remain active but experience transient failures.

Frequently Asked Questions

What is the difference between channel health probing and auto-disable in AxonHub?

Channel health probing is a temporary, automated recovery mechanism that tests whether a failed model has healed after a cooldown period, using a circuit-breaker pattern to allow single probe requests. Auto-disable is a permanent administrative action that removes a channel or API key from service after repeated specific HTTP errors (such as 401 or 429) indicate configuration or rate-limit issues that require human intervention.

How does AxonHub prevent multiple simultaneous probe requests to a recovering channel?

AxonHub uses atomic integer operations in internal/server/orchestrator/model_circuit_breaker.go to ensure only one probe executes at a time. The TryBeginProbe function calls atomic.CompareAndSwapInt32 to atomically flip the probingInProgress flag from 0 to 1. If the swap succeeds, the current request becomes the probe; if another goroutine has already set the flag, the function returns false and the request is treated as normal traffic (which gets dropped while the channel is Open).

What HTTP status codes trigger auto-disable, and how are thresholds configured?

Auto-disable triggers on any HTTP status codes explicitly configured in the RetryPolicy.AutoDisableChannel.Statuses array, commonly including 401 (unauthorized), 429 (rate limited), and 500 (internal server error). Thresholds are configured via YAML in config.example.yml by specifying the status and times fields—for example, setting status: 401 with times: 3 disables the channel after three consecutive 401 errors. The ChannelService in internal/server/biz/channel_auto_disable.go maintains per-channel counters to track these consecutive failures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →