# How Channel Health Probing and Auto-Disable Work in AxonHub

> Learn how AxonHub uses channel health probing and auto disable to protect LLM providers. Discover circuit breaker style testing and permanent removal of failing channels or API keys.

- Repository: [Loop/axonhub](https://github.com/looplj/axonhub)
- Tags: internals
- Published: 2026-03-06

---

**AxonHub protects downstream LLM providers by continuously monitoring channel health through circuit-breaker-style probing that tests recovery after failures, while auto-disable permanently removes channels or API keys after repeated specific HTTP errors.**

The `looplj/axonhub` repository implements a dual-layer resilience strategy to prevent cascading failures when proxying requests to large language model APIs. Channel health probing and auto-disable work at different granularities—one handles transient outages through automated recovery testing, while the other handles persistent authentication or rate-limit issues through administrative shutdown.

## Understanding Channel Health Probing

Channel health probing operates on individual **models** within a channel using a circuit-breaker pattern. When a model exceeds failure thresholds, AxonHub enters a probationary state where it allows exactly one test request after a cooldown period to verify recovery.

### State Tracking and Failure Thresholds

Each model maintains its health state in the `ModelCircuitBreakerStats` struct defined in [`internal/server/biz/model_circuit_breaker.go`](https://github.com/looplj/axonhub/blob/main/internal/server/biz/model_circuit_breaker.go). The struct tracks `State` (closed, half_open, or open), `ConsecutiveFailures`, and `NextProbeAt`.

When `RecordError` detects that `ConsecutiveFailures` exceeds the `OpenThreshold` (default 5), it transitions the state to `StateOpen` and schedules the next probe:

```go
// internal/server/biz/model_circuit_breaker.go
if stats.ConsecutiveFailures >= policy.OpenThreshold {
    stats.State = StateOpen
    stats.NextProbeAt = time.Now().Add(policy.ProbeInterval) // default 5 minutes
}

```

### The Probing Window

While a model is in `StateOpen`, `GetEffectiveWeight` returns `0.0` to the load balancer, dropping all traffic. However, once `time.Now()` exceeds `NextProbeAt` and no probe is currently in progress, the weight temporarily increases to `baseWeight * HalfOpenWeight` (default 0.3), allowing exactly one request through:

```go
// internal/server/biz/model_circuit_breaker.go:88-96
if time.Now().After(stats.NextProbeAt) {
    if atomic.LoadInt32(&stats.probingInProgress) == 0 {
        return baseWeight * policy.HalfOpenWeight // allow single probe
    }
}
return 0.0 // otherwise drop traffic

```

### Probe Execution and State Recovery

The orchestrator in [`internal/server/orchestrator/model_circuit_breaker.go`](https://github.com/looplj/axonhub/blob/main/internal/server/orchestrator/model_circuit_breaker.go) manages the probe lifecycle. `TryBeginProbe` atomically sets `probingInProgress` from 0 to 1, ensuring only one request becomes the probe even under concurrent load:

```go
// internal/server/orchestrator/model_circuit_breaker.go:2-16
func (m *ModelCircuitBreaker) TryBeginProbe(ctx context.Context, channelID int, modelID string) bool {
    // ... state checks ...
    return atomic.CompareAndSwapInt32(&stats.probingInProgress, 0, 1)
}

```

When the probe request completes successfully, `RecordSuccess` resets the circuit breaker to `StateClosed`, clears `ConsecutiveFailures`, and removes the probe schedule, immediately restoring full traffic capacity.

## How Auto-Disable Protects Your Infrastructure

While probing handles transient failures, **auto-disable** handles persistent error patterns that indicate configuration issues (invalid API keys) or upstream rate limits. After a configurable number of consecutive errors with a specific HTTP status code, AxonHub permanently disables the channel or API key until manual operator intervention.

### Configurable Error Thresholds

The `RetryPolicy` struct in [`internal/server/biz/system.go`](https://github.com/looplj/axonhub/blob/main/internal/server/biz/system.go) defines `AutoDisableChannel`, which maps HTTP status codes to failure thresholds:

```go
// internal/server/biz/system.go
type AutoDisableChannel struct {
    Enabled  bool
    Statuses []AutoDisableChannelStatus // e.g., {Status: 401, Times: 3}
}

```

### Per-Status Error Counting

The `ChannelService` in [`internal/server/biz/channel_auto_disable.go`](https://github.com/looplj/axonhub/blob/main/internal/server/biz/channel_auto_disable.go) maintains in-memory counters (`channelErrorCounts` and `apiKeyErrorCounts`) that track consecutive errors per status code. The `checkAndHandleChannelError` method increments counters when a request fails:

```go
// internal/server/biz/channel_auto_disable.go:47-55
svc.channelErrorCounts[perf.ChannelID][perf.ErrorStatusCode]++
count := svc.channelErrorCounts[perf.ChannelID][perf.ErrorStatusCode]

```

### Channel and API Key Disablement

When the error count reaches the configured `Times` threshold, `markChannelUnavailable` executes a database update setting `channel.Status = StatusDisabled`, effectively removing the channel from the load-balancer pool:

```go
// internal/server/biz/channel_auto_disable.go:58-63
if count >= statusConfig.Times {
    svc.markChannelUnavailable(ctx, perf.ChannelID, perf.ErrorStatusCode)
    // reset counters to prevent duplicate disablement
}

```

The same logic applies to API keys via `DisableAPIKey`, which generates a reason string such as "Auto-disabled after 3 consecutive errors with status 401" for audit trails.

## Interaction Between Probing and Auto-Disable

These mechanisms operate at different layers and time scales:

*   **Probing** functions at the **model level** within [`internal/server/biz/model_circuit_breaker.go`](https://github.com/looplj/axonhub/blob/main/internal/server/biz/model_circuit_breaker.go). It activates only after a channel has entered the `Open` state due to general failures, allowing temporary recovery without administrative action.
*   **Auto-disable** functions at the **channel or API-key level** within [`internal/server/biz/channel_auto_disable.go`](https://github.com/looplj/axonhub/blob/main/internal/server/biz/channel_auto_disable.go). It triggers on specific HTTP status codes (401, 429, 500) that indicate non-transient issues, permanently removing the resource from rotation.

If a channel is auto-disabled, it is removed from the load-balancer entirely, so the probing mechanism never executes for that channel's models. Conversely, if a channel experiences transient 500 errors that trigger the circuit breaker but not the auto-disable threshold, probing provides the recovery path.

## Configuration Examples

### Enabling Auto-Disable in YAML

Configure specific status codes and thresholds in your AxonHub configuration file:

```yaml

# config.example.yml

retry_policy:
  auto_disable_channel:
    enabled: true
    statuses:
      - status: 401   # Unauthorized - likely invalid API key

        times: 3      # Disable after 3 consecutive 401s

      - status: 429   # Rate limited

        times: 5      # Allow brief spikes before disabling

      - status: 500   # Internal server error

        times: 5

```

### Manual Probe Testing

For administrative endpoints or testing scenarios, you can manually trigger a probe:

```go
// Testing probe initiation
ctx := context.Background()
svc := channelService // from dependency injection

// Attempt to begin probe for model "gpt-4" on channel 42
if svc.modelCircuitBreaker.TryBeginProbe(ctx, 42, "gpt-4") {
    // Execute single probe request
    resp, err := llmClient.Ping(ctx, "gpt-4")
    if err == nil && resp.Healthy {
        svc.modelCircuitBreaker.RecordSuccess(ctx, 42, "gpt-4")
        fmt.Println("Channel recovered and returned to service")
    } else {
        svc.modelCircuitBreaker.RecordError(ctx, 42, "gpt-4")
        fmt.Println("Probe failed, channel remains in Open state")
    }
}

```

## Summary

*   **Channel health probing** implements a circuit-breaker pattern at the model level, using [`internal/server/biz/model_circuit_breaker.go`](https://github.com/looplj/axonhub/blob/main/internal/server/biz/model_circuit_breaker.go) to track failures and allow single probe requests after a cooldown period.
*   **Auto-disable** provides permanent protection by monitoring specific HTTP status codes in [`internal/server/biz/channel_auto_disable.go`](https://github.com/looplj/axonhub/blob/main/internal/server/biz/channel_auto_disable.go), disabling channels or API keys after configurable consecutive error thresholds.
*   Probing uses atomic operations (`atomic.CompareAndSwapInt32`) to ensure exactly one probe request executes concurrently, preventing thundering herds during recovery attempts.
*   The two systems operate hierarchically: auto-disable removes channels from the load balancer entirely, while probing manages traffic distribution for channels that remain active but experience transient failures.

## Frequently Asked Questions

### What is the difference between channel health probing and auto-disable in AxonHub?

Channel health probing is a temporary, automated recovery mechanism that tests whether a failed model has healed after a cooldown period, using a circuit-breaker pattern to allow single probe requests. Auto-disable is a permanent administrative action that removes a channel or API key from service after repeated specific HTTP errors (such as 401 or 429) indicate configuration or rate-limit issues that require human intervention.

### How does AxonHub prevent multiple simultaneous probe requests to a recovering channel?

AxonHub uses atomic integer operations in [`internal/server/orchestrator/model_circuit_breaker.go`](https://github.com/looplj/axonhub/blob/main/internal/server/orchestrator/model_circuit_breaker.go) to ensure only one probe executes at a time. The `TryBeginProbe` function calls `atomic.CompareAndSwapInt32` to atomically flip the `probingInProgress` flag from 0 to 1. If the swap succeeds, the current request becomes the probe; if another goroutine has already set the flag, the function returns false and the request is treated as normal traffic (which gets dropped while the channel is Open).

### What HTTP status codes trigger auto-disable, and how are thresholds configured?

Auto-disable triggers on any HTTP status codes explicitly configured in the `RetryPolicy.AutoDisableChannel.Statuses` array, commonly including 401 (unauthorized), 429 (rate limited), and 500 (internal server error). Thresholds are configured via YAML in [`config.example.yml`](https://github.com/looplj/axonhub/blob/main/config.example.yml) by specifying the `status` and `times` fields—for example, setting `status: 401` with `times: 3` disables the channel after three consecutive 401 errors. The `ChannelService` in [`internal/server/biz/channel_auto_disable.go`](https://github.com/looplj/axonhub/blob/main/internal/server/biz/channel_auto_disable.go) maintains per-channel counters to track these consecutive failures.