Paperclip AI Heartbeat Queue Coalescing and Orphan Run Recovery: A Technical Deep Dive
Paperclip AI uses policy-driven heartbeat run coalescing to eliminate duplicate work and employs proactive orphan detection with automatic lock recovery to ensure agent execution remains reliable even after process crashes.
The heartbeat subsystem in paperclipai/paperclip orchestrates autonomous agent execution through a state machine of heartbeat runs tracked in the heartbeatRuns table. These runs progress through queued → running → scheduled_retry → terminated states (defined in server/src/services/heartbeat.ts#L367‑L382). Two critical reliability mechanisms—queue coalescing and orphan run recovery—keep this system efficient and self-healing.
Heartbeat Queue Coalescing
When a routine triggers, Paperclip must decide whether to spawn a new run or merge into an existing one. This decision flows from the routine's concurrency policy, configured at declaration time in files like server/src/services/plugin-managed-routines.ts#L50.
Concurrency Policy Options
Paperclip implements three concurrency behaviors:
always_enqueue— Every trigger creates a fresh run; no deduplication occurs.skip_if_active— If any run isqueuedorrunning, the trigger is silently dropped.coalesce_if_active(default) — If a live run exists, the trigger merges into it, returning acoalescedstatus.
The actual coalescing logic resides in server/src/services/routines.ts#L262‑L282. When enqueueWakeup processes a trigger, it queries for matching live runs; if found and policy permits, it returns the coalesced run without creating a new database row.
Why Coalescing Matters
Coalescing delivers three operational benefits:
- Resource efficiency — Prevents redundant work from rapid successive triggers (e.g., user double-clicks, webhook retries).
- Idempotency guarantees — The same logical operation executes once, eliminating race conditions.
- Back-pressure protection — Downstream services like LLM providers receive consolidated rather than bursty requests.
Frontend Visibility
The UI surfaces coalescing behavior through the cron-fire helper in ui/src/lib/cron-fires.ts#L216‑L228. This maps internal dispositions to user-facing labels:
import { previewFirePolicies } from "@/ui/src/lib/cron-fires";
const fires = previewFirePolicies("* * * * *", "coalesce_if_active");
console.log(fires.map(f => f.disposition));
// Output: ["queued", "coalesced", "coalesced"] — second and third triggers merge
Orphan Run Detection and Recovery
An orphan run has status = 'running' in the database but its underlying OS process or process group no longer exists. These occur after server crashes, kill -9 terminations, or container evictions.
Startup Reaper: First Line of Defense
Before any timer ticks begin, the server runs a startup orphan reaper (server/src/index.ts#L1024‑L1055). This sequencing is deliberate: stale runs must be cleared so new triggers cannot accidentally coalesce with them.
The reaper executes four steps:
- Identify orphans — Query
heartbeatRunsfor rows whoseprocessPid/processGroupIdno longer correspond to live OS processes. - Mark terminal — Update status with error codes
orphaned_running_runororphaned_running_run_issue_terminal. - Release execution locks — Clear
executionRunId,executionAgentNameKey,executionLockedAtcolumns (server/src/services/heartbeat.ts#L16455‑L16457). - Queue process-loss retry — Optionally enqueue a retry so the agent resumes after clean start (
server/src/services/heartbeat.ts#L10633‑L10639).
Periodic Orphan Sweep
A 5-minute periodic sweep (server/src/index.ts#L1266 comment) captures orphans that appear post-startup. This background task maintains system health during long-running server lifetimes.
Process-Group Cleanup Events
When heartbeat detects a dead process group, it emits a synthetic event:
{ "kind": "orphaned_process_group_cleanup", "runId": "...", "agentId": "..." }
This event originates at server/src/services/heartbeat.ts#L13197 and propagates to the plugin-worker-manager, which closes lingering SSE streams (server/src/services/plugin-worker-manager.ts#L1228).
Lock Table Recovery
Orphaned runs hold execution locks on issues that would otherwise block other agents indefinitely. The reaper explicitly clears these in server/src/services/routines.ts#L290‑L306, unblocking the work queue.
End-to-End Execution Flow
The complete lifecycle integrates both mechanisms:
- Trigger arrival →
enqueueWakeupcreates aqueuedrun (or returnscoalescedper policy). - Process launch → Run transitions to
runningwithprocessPidandprocessGroupIdrecorded. - Normal completion → Run terminates, locks release, no cleanup needed.
- Orphan detection → Startup reaper or periodic sweep identifies stale
runningrows, marks terminal, releases locks, optionally retries.
Practical Code Examples
Enqueue with default coalescing:
import { enqueueWakeup } from "@paperclipai/server/src/services/heartbeat.js";
const run = await enqueueWakeup(agentId, {
action: "my.routine.fire",
payload: { task: "generate-report" }
});
// "queued" if no live run exists, "coalesced" if merged into existing run
console.log(run.status);
Manual orphan recovery:
import { getRunById, markRunTerminal, enqueueWakeup }
from "@paperclipai/server/src/services/heartbeat.js";
const orphan = await getRunById(orphanRunId);
if (orphan?.status === "running" && !(await processExists(orphan.processPid))) {
await markRunTerminal(orphan.id, "orphaned_running_run");
await enqueueWakeup(orphan.agentId, {
action: "retry.after.orphan",
originalRunId: orphan.id
});
}
Preview coalescing behavior in UI:
import { previewFirePolicies } from "@/ui/src/lib/cron-fires";
const schedule = "*/5 * * * *"; // Every 5 minutes
const fires = previewFirePolicies(schedule, "coalesce_if_active");
// Visualize which fires queue fresh vs. merge into active runs
Summary
- Queue coalescing via
coalesce_if_active(default) prevents duplicate heartbeat runs by merging triggers into live executions. - Orphan detection operates at startup (
server/src/index.ts#L1024‑L1055) and every 5 minutes to identify and terminate stalerunningrows. - Lock recovery ensures orphaned runs cannot indefinitely block other agents from executing the same work.
- Retry semantics allow automatic resumption after process loss through process-loss retry enqueuing.
- Observability is built in: all orphan operations log explicit error codes and emit cleanup events for downstream handling.
Frequently Asked Questions
How does Paperclip AI prevent duplicate routine executions?
Paperclip AI applies concurrency policies at trigger time. The default coalesce_if_active policy checks server/src/services/routines.ts#L262‑L282 for existing queued or running runs; if found, the trigger merges into that run rather than creating a new row. This deduplication happens transparently before any OS process spawns.
What happens to heartbeat runs when the server crashes?
Orphaned runs are reclaimed at the next startup. The reaper in server/src/index.ts#L1024‑L1055 executes before timer initialization, finding running rows with dead PIDs, marking them terminal with error code orphaned_running_run, releasing execution locks, and optionally queueing retries. A 5-minute periodic sweep handles crashes during runtime.
Can I disable heartbeat queue coalescing for a specific routine?
Yes, set the concurrency policy to always_enqueue. This is configured where routines are declared—typically server/src/services/plugin-managed-routines.ts or pipeline definitions. With this policy, server/src/services/routines.ts bypasses live-run checks and always inserts new heartbeatRuns rows.
How does Paperclip AI clean up resources from orphaned runs?
Through lock clearing and synthetic events. The reaper updates execution-lock columns (executionRunId, etc.) in server/src/services/routines.ts#L290‑L306 and emits orphaned_process_group_cleanup events from server/src/services/heartbeat.ts#L13197. The plugin-worker-manager consumes these to close SSE streams at server/src/services/plugin-worker-manager.ts#L1228.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →