How the Paperclip Heartbeat System Triggers Agent Execution: A Deep Architectural Guide
The Paperclip heartbeat system triggers agent execution through a three-phase pipeline: external events create wakeup requests, a background queue processor claims ready runs while respecting concurrency limits, and an execution engine launches agents with full runtime instrumentation.
The heartbeat ecosystem in Paperclip transforms asynchronous events—issue comments, scheduled timers, manual approvals—into running agent executions. This article examines the implementation in the paperclipai/paperclip repository, tracing the path from wakeup signal to completed run.
Core Architecture: Three Phases of Execution
The heartbeat service operates as a continuous loop with three tightly coupled phases: wakeup reception, queue processing, and run execution. Each phase is implemented in server/src/services/heartbeat.ts with specific responsibilities for reliability and observability.
Phase 1: Wakeup Reception and Registration
External events enter the system through heartbeat.trackWakeup(agentId, opts). This method registers the wakeup in activeWakeupPromises and forwards to enqueueWakeup for persistence.
// server/src/services/heartbeat.ts
// trackWakeup implementation around L13935
public trackWakeup(
agentId: string,
opts: {
reason?: AgentWakeupReason;
priority?: number;
metadata?: Record<string, unknown>;
} = {}
): void {
const promise = this.enqueueWakeup(agentId, opts);
this.activeWakeupPromises.add(promise);
promise.finally(() => this.activeWakeupPromises.delete(promise));
}
The wakeup request is stored in the agentWakeupRequests table, creating a durable signal that survives process restarts. Priority and metadata allow fine-grained control over execution ordering.
Phase 2: Queue Processing and Run Claiming
A background timer (started via heartbeat.start()) repeatedly invokes startNextQueuedRunForAgent(agentId). This function implements the core dispatch logic with strict concurrency controls.
// Pseudocode based on implementation around L13822
async startNextQueuedRunForAgent(agentId: string): Promise<void> {
// Single-dispatcher guarantee per agent
await this.withAgentStartLock(agentId, async () => {
const agent = await this.loadAgent(agentId);
const freeSlots = this.computeFreeSlots(agent.heartbeatPolicy.maxConcurrentRuns);
if (freeSlots <= 0) return;
const readyRuns = await this.selectReadyRuns(agentId, freeSlots);
for (const run of readyRuns) {
const claimed = await this.claimQueuedRun(run.id);
if (claimed) {
const execPromise = this.executeRun(run.id);
this.activeRunExecutionPromises.add(execPromise);
execPromise.finally(() => this.activeRunExecutionPromises.delete(execPromise));
}
}
});
}
Key safeguards in this phase include:
- Scheduling suppression — prevents runs outside allowed time windows
- Work-tree cut-off — blocks execution when repository state is unstable
- Agent start lock —
withAgentStartLockguarantees exactly one dispatcher per agent - Dependency sorting — runs are ordered by dependency readiness, issue priority, and creation time
Phase 3: Run Execution and Runtime Instrumentation
Once claimed, executeRun(runId) transforms the queued run into an active agent execution. This is the heavyweight phase involving database state transitions, runtime preparation, and adapter invocation.
// Core execution flow beginning around L13952
async executeRun(runId: string): Promise<void> {
// 1. State validation and claim
const run = await this.reloadAndValidateRun(runId);
// 2. Runtime preparation
const agent = await this.loadAgent(run.agentId);
const runtimeState = await this.ensureRuntimeState(run);
// 3. Environment construction
const scratch = await this.prepareHeartbeatRunScratch(run, agent);
const { taskKey, codec, issueContext } = this.deriveExecutionParams(run, agent);
// 4. Adapter invocation
const adapter = this.getServerAdapter(agent.adapterType);
const execution = await adapter.start({
runtimeState,
scratch,
taskKey,
codec,
streamingCallbacks: this.createLogStreamers(runId)
});
// 5. Completion and event publishing
await this.finalizeRun(runId, execution.result);
await this.publishLiveEvent({ type: 'run_completed', runId, result: execution.result });
}
Critical persistence guarantee: All database writes complete before the returned promise resolves. This enables deterministic test teardown via drainActiveRunExecutions().
Coordination Mechanisms for Reliability
Graceful Shutdown and Test Isolation
The heartbeat system tracks all active work through three internal sets:
activeRunExecutions— currently running agent processesactiveRunExecutionPromises— in-flightexecuteRuncallsactiveWakeupPromises— pending wakeup registrations
// Awaiting complete quiescence
async drainActiveRunExecutions(): Promise<void> {
await Promise.all([...this.activeWakeupPromises]);
while (
this.activeRunExecutions.size > 0 ||
this.activeRunExecutionPromises.size > 0
) {
await Promise.race([
...this.activeRunExecutionPromises,
sleep(100)
]);
}
}
Workspace Contention Handling
When multiple agents target the same workspace, WorkspaceBusyDeferral implements bounded exponential backoff:
// Retry configuration from source
const WORKSPACE_BUSY_RETRY_BASE_DELAY_MS = 1000;
const WORKSPACE_BUSY_RETRY_MAX_DELAY_MS = 30000;
const WORKSPACE_BUSY_RETRY_MAX_ATTEMPTS = 5;
Policy Enforcement Injection
Before execution, multiple services validate and prepare the run context:
| Service | Responsibility |
|---|---|
costService |
Estimates and tracks token consumption |
budgetService |
Enforces spending limits per agent/organization |
secretService |
Binds and injects configured secrets |
trustService |
Validates low-trust environment requirements |
Practical Usage Examples
Programmatic Wakeup Trigger
import { heartbeat } from '@paperclipai/server';
// Fire-and-forget wakeup from external event handler
app.post('/github/webhook', async (req, res) => {
const { agentId, issueNumber } = parseWebhook(req.body);
heartbeat.trackWakeup(agentId, {
reason: 'issue_commented',
priority: issueNumber,
metadata: { commentId: req.body.comment.id }
});
res.status(202).send('Wakeup queued');
});
CLI Invocation
The paperclipai heartbeat-run command uses the same internal path:
# Direct agent execution via CLI
$ paperclipai heartbeat-run --agent-id=agent-abc123 --wait
# With priority override
$ paperclipai heartbeat-run --agent-id=agent-abc123 --priority=10 --reason=manual_trigger
Test Harness with Drain
import { createTestHeartbeat } from '@paperclipai/server/testing';
describe('agent execution', () => {
let heartbeat: HeartbeatService;
beforeEach(() => {
heartbeat = createTestHeartbeat();
});
afterEach(async () => {
// Ensures no orphaned database writes
await heartbeat.drainActiveRunExecutions();
});
it('processes wakeup through completion', async () => {
heartbeat.trackWakeup('test-agent', { reason: 'test' });
// Wait for natural completion
await heartbeat.drainActiveRunExecutions();
const runs = await listRuns('test-agent');
expect(runs[0].status).toBe('completed');
});
});
Key Source Files
| File | Role | Link |
|---|---|---|
server/src/services/heartbeat.ts |
Core service: wakeup handling, queue dispatch, execution orchestration | github.com/paperclipai/paperclip/blob/master/server/src/services/heartbeat.ts |
server/src/services/heartbeat-run-summary.ts |
Result formatting, output trimming, PR comment generation | heartbeat-run-summary.ts |
packages/shared/src/types/heartbeat.ts |
Public TypeScript definitions for payloads and status enums | heartbeat types |
cli/src/commands/heartbeat-run.ts |
CLI interface wrapping trackWakeup |
CLI command |
ui/src/api/heartbeats.ts |
Frontend polling and real-time status subscription | UI API |
Summary
The Paperclip heartbeat system triggers agent execution through:
- Durable wakeup registration via
trackWakeup, storing signals inagentWakeupRequestswith priority metadata - Lock-protected queue processing in
startNextQueuedRunForAgent, respectingmaxConcurrentRunsand dependency ordering - Fully instrumented execution through
executeRun, with guaranteed persistence before promise resolution - Clean coordination primitives including
drainActiveRunExecutionsfor graceful shutdown and test isolation
This architecture ensures that every external event—whether from webhooks, schedules, or manual triggers—reliably converts to an executed agent run with full observability and policy enforcement.
Frequently Asked Questions
What happens if an agent is already at max concurrency when a wakeup arrives?
The wakeup is stored in agentWakeupRequests and the queue processor will attempt to claim runs on the next polling cycle. No events are dropped; the system continuously retries until capacity becomes available or the request expires based on configured TTL policies.
How does the heartbeat system handle process crashes during execution?
Run state is persisted to the database at each phase transition (queued → running → completed/failed). On restart, the queue processor rescues runs stuck in running state by validating heartbeats from the underlying adapter; orphaned runs are marked failed and can be retried.
Can multiple heartbeat processes run simultaneously for high availability?
Yes, the withAgentStartLock mechanism ensures only one dispatcher operates per agent across all processes. The lock is implemented via database advisory locks or distributed lock service depending on deployment configuration.
What is the difference between trackWakeup and directly calling executeRun?
trackWakeup enqueues a durable request and returns immediately, allowing the background processor to handle scheduling and concurrency. Direct executeRun calls bypass the queue and are used only internally after a run has been claimed and validated.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →