How the Circuit Breaker’s Steer-Constrain-Stop Escalation Ladder Prevents Runaway Agents in Munder Difflin
The circuit breaker’s steer-constrain-stop escalation ladder is a three-tiered guardrail system that incrementally escalates agent restrictions from gentle behavioral nudges to strict resource limits and finally termination, preventing autonomous agents from spiraling into infinite loops or uncontrolled token consumption.
The circuit breaker’s steer-constrain-stop escalation ladder serves as the core safety mechanism in the Munder Difflin autonomous agent framework. Implemented in the CircuitBreaker class within the heartbeat mission, this system monitors agent behavior in real-time and applies progressive sanctions when agents exhibit signs of runaway behavior such as repetitive tool use, API error storms, or token-velocity spikes.
The Three Levels of the Escalation Ladder
The ladder defines four discrete states—healthy, steering, constrained, and stopped—through which agents transition based on per-beat evaluations in src/main/breaker.ts.
Level 1: Steering (Corrective Nudges)
The steering level represents the first warning stage. When the breaker detects early signs of trouble—such as initial instances of repetitive tool calls—it sets action = "steer" and emits a corrective chat message to the agent. This gentle nudge encourages the agent to alter its strategy without imposing hard limits, often resolving transient glitches without human intervention.
Level 2: Constrained (Strict Resource Limits)
If warning conditions persist across subsequent heartbeats, the ladder escalates to constrained. At this level, the breaker applies action = "constrain" and enforces stricter policies such as token-velocity caps and tighter budget limits. The agent receives a tighter policy command that throttles consumption while allowing useful work to continue.
Level 3: Stopped (Termination and Archival)
The final stopped level triggers only when the optional hardStop flag is enabled in src/main/config.ts. At this stage, the breaker returns action = "stop", causing the heartbeat loop in src/main/index.ts to terminate the agent’s PTY and archive the session. This serves as the ultimate safety net for truly runaway agents that would otherwise burn unlimited tokens.
How the Ladder Evaluates Agent Health
The evaluation engine resides in src/main/breaker.ts, where the CircuitBreaker.tick method processes a list of BreakerInput objects for each active agent. Each input aggregates three signal sources:
- Usage samples from the
UsageProvider(token counts, costs, timestamps) - Hook events recorded by the
HookServer(repeated identical tool calls, API error storms) - File-mtime progress indicators showing whether the agent touched coordination files during the last beat
The evaluate method compares these inputs against thresholds defined in CircuitBreakerConfig at src/main/config.ts (lines 44–61). When any trip condition returns tripping: true, the ladder escalates exactly one level—never jumping directly from healthy to stopped. Conversely, when a beat passes without violations, the ladder de-escalates one step, allowing agents to recover gradually.
Trip Conditions That Trigger Escalation
The breaker monitors seven specific failure modes defined in src/main/breaker.ts:
- Repeated identical tool calls: Trips when
repeatCount ≥ repeatedToolLimit(lines 94–98), reporting"looping: … identical tool call" - Error storms: Triggers when
errorCount ≥ errorStormLimit(lines 99–101), reporting"error storm: …" - Per-agent token caps: Fires when
tokensOf(sample) > agentTokenCaps[agentId](lines 103–106), reporting"token limit: … over the agent cap" - Floor-wide cost caps: Activates when the floor’s total spending exceeds limits (lines 108–112), reporting
"cost cap: floor total over $…" - Floor-wide token caps: Triggers on aggregate token consumption (lines 114–118), reporting
"token cap: floor total over … tokens" - Token-velocity spikes: Detects when
velocity > tokenVelocityPerMin(lines 124–127), reporting"token velocity …/min > …/min" - No-progress states: Identifies agents generating output tokens without updating coordination files (lines 132–138), reporting
"no-progress: generating tokens without coordinating"
Each condition updates the UI via state.level and state.reason, displaying status badges in the agent card through src/renderer/src/hooks/useHive.ts so operators can diagnose why an agent was nudged, limited, or killed.
Implementation in the Heartbeat Loop
The heartbeat mission in src/main/index.ts (lines 1124–1194) serves as the sole orchestrator of the circuit breaker. Each beat calls CircuitBreaker.tick with fresh inputs, then inspects the returned action field:
- steer: Sends a corrective message to the agent’s chat interface
- constrain: Applies tighter policy restrictions via the constrain command
- stop: Kills the PTY and archives the agent (only when
hardStopis enabled)
This centralized enforcement ensures that escalation decisions remain consistent across all agents while maintaining separation between detection logic and execution.
Configuration and Customization
All thresholds and behavior flags live in src/main/config.ts under the circuitBreaker namespace. The default configuration adopts a conservative posture with hardStop disabled to prevent unintended terminations.
Basic Configuration
{
"circuitBreaker": {
"enabled": true,
"hardStop": false,
"repeatedToolLimit": 8,
"errorStormLimit": 5,
"tokenVelocityPerMin": 60000
}
}
Per-Agent Token Caps
Apply hard limits to individual agents using the configuration API:
import { setAgentTokenCap } from '../main/config';
// Limit agent "creed-xyz" to 1,000,000 tokens
await setAgentTokenCap('creed-xyz', 1_000_000);
Runtime Integration
Instantiate the breaker in the renderer process with dynamic configuration:
import { CircuitBreaker } from '../main/breaker';
import { readConfig } from '../main/config';
const breaker = new CircuitBreaker(() => readConfig().circuitBreaker ?? {});
Then execute the evaluation loop:
const inputs = agents.map(agent => ({
agentId: agent.id,
sample: usageProvider.sampleFor(agent.id),
progressing: fileMtimeChanged(agent.id)
}));
const decisions = breaker.tick(inputs, Date.now());
for (const d of decisions) {
if (d.changed && d.action !== 'none') {
sendBreakerMessage(d.state.agentId, d.action, d.state.reason);
}
}
Summary
- The circuit breaker’s steer-constrain-stop escalation ladder provides a graduated response to agent misbehavior, moving from gentle corrections to absolute termination.
- Implemented in
src/main/breaker.ts, the system evaluates agents against seven trip conditions including repetitive tool use, error storms, and token-velocity spikes. - Escalation occurs incrementally—one level per heartbeat—preventing false positives while guaranteeing eventual containment of rogue agents.
- De-escalation happens automatically when agents return to healthy behavior, allowing recovery without manual intervention.
- Configuration resides in
src/main/config.ts, with an optionalhardStopflag controlling whether the ladder can reach the terminal stopped state.
Frequently Asked Questions
What triggers the first level of the escalation ladder?
The steering level triggers when the evaluate method in src/main/breaker.ts detects any initial trip condition, such as repetitive identical tool calls exceeding the repeatedToolLimit threshold or the start of an API error storm. At this stage, the breaker emits a corrective message to guide the agent back to healthy behavior without imposing resource restrictions.
Can an agent jump directly from healthy to stopped?
No. The ladder enforces incremental escalation—agents can only advance one level per heartbeat tick. An agent must progress through steering and constrained states before reaching stopped, and even then, only if the hardStop configuration flag is enabled. This design prevents immediate termination due to transient spikes.
How does the circuit breaker detect runaway token consumption?
The breaker monitors token velocity (tokens per minute) against the tokenVelocityPerMin threshold and tracks aggregate consumption against per-agent caps (agentTokenCaps) and floor-wide limits. These checks occur in src/main/breaker.ts lines 103–127, comparing real-time usage samples from the UsageProvider against configured maximums.
What happens when an agent recovers from a constrained state?
When an agent completes a heartbeat without triggering any trip conditions, the ladder de-escalates one level automatically. An agent in the constrained state moves back to steering on the next healthy beat, then returns to healthy on the subsequent clean beat. This symmetrical recovery path ensures agents can resume full autonomy once they demonstrate stable behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →