How Prime Agent Autonomous Mode Uses Turn, Token, and Time Budgets with Quality Gates
Prime Agent's autonomous mode enables LLM sessions to continue working without human intervention by tracking turn, token, and time consumption against configurable limits, while external quality gates determine when the task is complete or needs retry.
Prime Agent, developed by PrimeIntellect-ai, implements a deterministic autonomous execution loop in packages/coding-agent/src/core/autonomous.ts that allows coding agents to self-direct until budgets exhaust or quality criteria pass. This system prevents runaway token consumption and infinite loops through strict accounting of resources and configurable shell-command validation.
Configuration Interface and Default Budget Limits
The autonomous behavior is governed by the AgentAutonomousConfig interface, which defines hard stops for resource consumption and quality validation parameters.
export interface AgentAutonomousConfig {
enabled?: boolean;
maxContinuations?: number; // autonomous "continue" messages allowed
maxTurns?: number; // LLM turns (assistant → user) allowed
maxTokens?: number; // cumulative token budget
timeoutMs?: number; // wall-clock time limit
continuationPrompt?: string;
gates?: AgentAutonomousGateConfig;
}
When users omit configuration, DEFAULT_AUTONOMOUS_LIMITS applies conservative defaults defined at line 45 of autonomous.ts:
export const DEFAULT_AUTONOMOUS_LIMITS = {
maxContinuations: 3,
maxTurns: 12,
maxTokens: 80_000,
timeoutMs: 30 * 60 * 1000, // 30 minutes
};
Runtime State and Budget Accounting
createAutonomousRuntimeState() constructs a mutable AutonomousRuntimeState object that tracks consumption against these limits. The function normalizes undefined values via normalizeLimit() (lines 6-33) and initializes counters for continuationsUsed, turnsUsed, tokensUsed, and startedAt.
The system updates metrics through specific accounting functions:
addAutonomousUsage(state, usage)(lines 71-77): IncrementsturnsUsedand adds the token delta (input + output + cache-write) totokensUsedafter each LLM turn.addAutonomousContinuation(state)(lines 79-85): IncrementscontinuationsUsedwhen injecting synthetic continue messages.- Time tracking:
startedAtis set on initialization;autonomousLimitReason()(lines 58-69) checks elapsed time againstlimits.timeoutMson each decision cycle.
Quality Gates and Retry Logic
Quality gates are external shell commands—such as test suites or linters—that validate work before allowing continuation. These execute via runAutonomousQualityGates (lines 73-82).
Git Worktree Optimization To prevent infinite retry loops, the system captures a Git worktree snapshot (status, diff, and hash of untracked files) before executing gate commands. If the snapshot matches a previous failure state, the gate skips execution and increments the retry counter. This prevents redundant computation when the workspace hasn't changed.
Execution Parameters Commands run with:
- Configurable timeout (
state.gates.timeoutMs) - Output truncation at
MAX_GATE_OUTPUT_CHARS - Execution within the session's working directory (
cwd)
On failure, truncated output populates lastGateFailure. When retries exceed gates.maxRetries, the result becomes "retry_exhausted".
Autonomous Decision Flow
The continuation logic centers on nextAutonomousContinuation(), called by AgentSession after each assistant message.
Continuation Criteria
shouldAutonomouslyContinue() (lines 27-52) evaluates:
- Abort conditions (error or aborted messages)
- Quality gate results via
refreshAutonomousQualityGates - Budget exhaustion via
autonomousLimitReason
Decision Outcomes
- Gate passes: Returns
reason: "not_needed"and stops autonomous execution - Gate fails with retries remaining: Returns
shouldContinue: truewithreason: "gate_failed" - Budget exhausted: Returns
reason: "limit_reached"for any exceeded limit (turns, tokens, time, or continuations)
When continuing, the system generates a synthetic user message using either buildGateFailureContinuation() (incorporating gate error output) or the generic continuationPrompt. This message injects back into the session as user input, driving the next LLM turn without human intervention.
Practical Implementation Example
Configure and initialize autonomous mode with custom budgets and quality gates:
import { createAutonomousRuntimeState, nextAutonomousContinuation } from "packages/coding-agent/src/core/autonomous.js";
const config = {
enabled: true,
maxContinuations: 5,
maxTurns: 20,
maxTokens: 120_000,
timeoutMs: 15 * 60 * 1000, // 15 minutes
gates: {
commands: ["npm test --silent", "npx tsc --noEmit"],
maxRetries: 2,
timeoutMs: 2 * 60 * 1000, // 2 minutes per gate
},
};
const state = createAutonomousRuntimeState(config);
// Inside session processing:
const autonomousMessage = await nextAutonomousContinuation(
state,
assistantMessage,
{ cwd: process.cwd() }
);
if (autonomousMessage) {
await session.sendUserMessage(autonomousMessage);
}
Summary
- Prime Agent autonomous mode operates through
packages/coding-agent/src/core/autonomous.ts, managing self-directed LLM sessions via strict resource accounting. - Budget types include turn count (
maxTurns), cumulative tokens (maxTokens), wall-clock time (timeoutMs), and continuation limits (maxContinuations), tracked inAutonomousRuntimeState. - Quality gates execute external shell commands with Git worktree snapshot caching to prevent infinite loops, supporting configurable retry limits.
- Decision logic in
shouldAutonomouslyContinuehalts execution when gates pass, budgets exhaust, or retry limits hit, ensuring deterministic termination. - Session integration occurs through
AgentSession.nextAutonomousContinuation, which injects synthetic user messages to drive continued execution without human prompts.
Frequently Asked Questions
How does Prime Agent prevent infinite loops in autonomous mode?
Prime Agent prevents infinite loops through three mechanisms: budget limits enforced by autonomousLimitReason(), quality gate retry limits that cap attempts per gate configuration, and Git worktree snapshots that skip gate re-execution when the workspace hasn't changed since the last failure. These safeguards ensure the autonomous loop terminates deterministically.
What counts toward the token budget in Prime Agent autonomous mode?
The token budget aggregates input tokens, output tokens, and cache-write tokens consumed during the session. The addAutonomousUsage() function updates tokensUsed with this cumulative delta after each LLM turn. When tokensUsed exceeds maxTokens (defaulting to 80,000), autonomousLimitReason() returns "maxTokens" and halts autonomous execution.
Can quality gates use any shell command, and how are they configured?
Yes, quality gates can execute any shell command—such as npm test, pytest, or custom linters—configured via the gates.commands array in AgentAutonomousConfig. Each command runs in the session's working directory with independent timeout (gates.timeoutMs) and retry (gates.maxRetries) settings. Failed gate output truncates to MAX_GATE_OUTPUT_CHARS and feeds back into the continuation prompt via buildGateFailureContinuation().
What happens when multiple budget limits are exceeded simultaneously?
The autonomousLimitReason() function checks budgets in priority order—maxContinuations, maxTurns, maxTokens, then timeoutMs—returning the first exceeded limit as the termination reason. While multiple limits may trigger simultaneously, the function reports the primary constraint that halted execution, and the session ceases autonomous continuation regardless of other remaining budget headroom.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →