How the Model Circuit Breaker Prevents Cascading Failures in AxonHub
The model circuit breaker in AxonHub isolates failing LLM providers by tracking per-model failure states and dynamically adjusting load-balancing weights, stopping error cascades before they propagate through the gateway.
AxonHub is an open-source LLM gateway that routes requests across multiple providers. To prevent a single failing model from overwhelming the system with retries and timeouts, the repository implements a model circuit breaker integrated directly into the load-balancing layer. This mechanism tracks health metrics per (channel, model) pair and automatically removes unhealthy endpoints from the rotation.
Per-Model State Tracking
ModelCircuitBreakerStats Structure
At the core of the system is the ModelCircuitBreakerStats struct defined in internal/server/biz/model_circuit_breaker.go. Each instance tracks the health of a specific model on a specific channel using a sync.RWMutex for thread-safe access.
type ModelCircuitBreakerStats struct {
sync.RWMutex
ChannelID int
ModelID string
State CircuitBreakerState
ConsecutiveFailures int
LastFailureAt time.Time
LastSuccessAt time.Time
NextProbeAt time.Time // for the Open → Half‑Open transition
probingInProgress int32 // atomic flag to allow a single probe request
}
Source: model_circuit_breaker.go
The stats are stored in an in-memory map (modelStats) keyed by (channelID, modelID), ensuring that circuit breaker decisions are localized to specific provider-model combinations. This prevents a failure in one model from affecting traffic to other models on the same channel.
Policy-Driven State Transitions
The model circuit breaker follows the classic Closed-Half-Open-Open pattern, with state transitions governed by a ModelCircuitBreakerPolicy that defines thresholds, time-to-live (TTL), probe intervals, and half-open weights.
Closed to Open Transition
When RecordError detects that ConsecutiveFailures has reached the OpenThreshold, the breaker immediately transitions to the Open state and schedules a probe time:
if stats.ConsecutiveFailures >= policy.OpenThreshold {
stats.State = StateOpen
stats.NextProbeAt = now.Add(policy.ProbeInterval)
}
Source: model_circuit_breaker.go – RecordError
Open to Half-Open Recovery
Once NextProbeAt expires, the breaker transitions to Half-Open and allows a single probe request with reduced weight (HalfOpenWeight). This tests whether the provider has recovered without risking full traffic volume.
Half-Open to Closed Reset
When RecordSuccess is called while the state is not Closed, the breaker immediately resets to Closed and clears all failure counters:
if stats.State != StateClosed && success {
stats.State = StateClosed
stats.ConsecutiveFailures = 0
}
Source: model_circuit_breaker.go – RecordSuccess
Load Balancer Integration
ModelAwareCircuitBreakerStrategy
The circuit breaker integrates with AxonHub's load balancer through the ModelAwareCircuitBreakerStrategy defined in internal/server/orchestrator/lb_strategy_model_aware_circuit_breaker.go. For each candidate channel, the strategy queries the breaker for an effective weight via GetEffectiveWeight.
effectiveWeight := s.cbProvider.GetEffectiveWeight(ctx, channel.ID, modelID, 1.0)
score := effectiveWeight * s.maxScore
Source: lb_strategy_model_aware_circuit_breaker.go – Score
The effective weight mapping works as follows:
- Closed →
baseWeight(full traffic allowed) - Half-Open →
baseWeight * HalfOpenWeight(reduced traffic for probing) - Open →
0(complete isolation unless a single probe is permitted)
When a channel is Open, its weight becomes 0, causing the load balancer to skip it entirely and route traffic exclusively to healthy providers.
Single-Probe Guard Mechanism
To prevent a "thundering herd" of probe requests when a model transitions from Open to Half-Open, the breaker uses an atomic flag (probingInProgress). Only one concurrent request is allowed to probe the failing provider:
if atomic.LoadInt32(&stats.probingInProgress) == 0 {
return baseWeight * policy.HalfOpenWeight // allow the probe
}
Source: model_circuit_breaker.go – GetEffectiveWeight (Open case)
This atomic guard ensures that even if multiple requests arrive simultaneously during the probe window, only one actually tests the provider, preventing overload during recovery attempts.
Orchestrator Error Handling
When the outbound transformer encounters a model in the Open state, it returns errSkipCandidateByCircuitBreaker. This signals the orchestrator to immediately skip the candidate and select the next healthy channel without wasting time on retries:
var errSkipCandidateByCircuitBreaker = errors.New("skip candidate by circuit breaker")
...
if errors.Is(err, errSkipCandidateByCircuitBreaker) { return false }
Source: outbound.go – errSkipCandidateByCircuitBreaker & CanRetry
This fast-fail mechanism ensures that the retry logic only attempts recovery on viable candidates, eliminating latency penalties from contacting known-bad providers.
Summary
- Fault Isolation: The
ModelCircuitBreakerStatsstruct tracks consecutive failures per(channel, model)pair, immediately transitioning unhealthy providers to theOpenstate where they receive zero traffic weight. - Controlled Recovery: The
Half-Openstate allows only a single probe request via theprobingInProgressatomic flag, testing provider health without risking traffic floods. - Automatic Reset: Successful probes immediately reset the breaker to
Closed, clearing failure counters and restoring full traffic capacity without manual intervention. - Load Balancer Integration: The
ModelAwareCircuitBreakerStrategydynamically adjusts channel scores based onGetEffectiveWeight, ensuring the orchestrator never selects a provider marked asOpen.
Frequently Asked Questions
What triggers the model circuit breaker to open?
The breaker transitions from Closed to Open when ConsecutiveFailures reaches the OpenThreshold defined in the ModelCircuitBreakerPolicy. This threshold is checked in the RecordError method within internal/server/biz/model_circuit_breaker.go, ensuring rapid isolation of consistently failing LLM providers.
How does AxonHub test if a failed model has recovered?
AxonHub uses the Half-Open state to test recovery. When the NextProbeAt timestamp expires, the breaker allows a single request through with reduced weight (HalfOpenWeight). The probingInProgress atomic flag in ModelCircuitBreakerStats ensures only one concurrent probe occurs, preventing overload during the health check.
What prevents multiple simultaneous probe requests when a model is in half-open state?
The probingInProgress field in the ModelCircuitBreakerStats struct acts as an atomic flag. Before allowing a probe request in the Open or Half-Open state, the GetEffectiveWeight method checks this flag using atomic.LoadInt32. If a probe is already in progress, subsequent requests receive a weight of zero until the current probe completes.
How does the circuit breaker integrate with AxonHub's load balancing?
The integration occurs through the ModelAwareCircuitBreakerStrategy in internal/server/orchestrator/lb_strategy_model_aware_circuit_breaker.go. This strategy calls GetEffectiveWeight for each candidate channel during the scoring phase. The effective weight multiplies the base score: full weight for Closed, reduced weight for Half-Open, and zero for Open states, ensuring the load balancer never routes traffic to isolated providers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →