# Paperclip AI Heartbeat Queue Coalescing and Orphan Run Recovery: A Technical Deep Dive

> Explore Paperclip AI heartbeat queue coalescing and orphan run recovery. Discover how policy-driven coalescing and proactive orphan detection ensure reliable agent execution after crashes.

- Repository: [Paperclip/paperclip](https://github.com/paperclipai/paperclip)
- Tags: deep-dive
- Published: 2026-08-12

---

**Paperclip AI uses policy-driven heartbeat run coalescing to eliminate duplicate work and employs proactive orphan detection with automatic lock recovery to ensure agent execution remains reliable even after process crashes.**

The heartbeat subsystem in [paperclipai/paperclip](https://github.com/paperclipai/paperclip) orchestrates autonomous agent execution through a state machine of **heartbeat runs** tracked in the `heartbeatRuns` table. These runs progress through `queued → running → scheduled_retry → terminated` states (defined in `server/src/services/heartbeat.ts#L367‑L382`). Two critical reliability mechanisms—**queue coalescing** and **orphan run recovery**—keep this system efficient and self-healing.

## Heartbeat Queue Coalescing

When a routine triggers, Paperclip must decide whether to spawn a new run or merge into an existing one. This decision flows from the routine's **concurrency policy**, configured at declaration time in files like `server/src/services/plugin-managed-routines.ts#L50`.

### Concurrency Policy Options

Paperclip implements three concurrency behaviors:

- **`always_enqueue`** — Every trigger creates a fresh run; no deduplication occurs.
- **`skip_if_active`** — If any run is `queued` or `running`, the trigger is silently dropped.
- **`coalesce_if_active`** (default) — If a live run exists, the trigger merges into it, returning a `coalesced` status.

The actual coalescing logic resides in `server/src/services/routines.ts#L262‑L282`. When `enqueueWakeup` processes a trigger, it queries for matching live runs; if found and policy permits, it returns the coalesced run without creating a new database row.

### Why Coalescing Matters

Coalescing delivers three operational benefits:

- **Resource efficiency** — Prevents redundant work from rapid successive triggers (e.g., user double-clicks, webhook retries).
- **Idempotency guarantees** — The same logical operation executes once, eliminating race conditions.
- **Back-pressure protection** — Downstream services like LLM providers receive consolidated rather than bursty requests.

### Frontend Visibility

The UI surfaces coalescing behavior through the **cron-fire helper** in `ui/src/lib/cron-fires.ts#L216‑L228`. This maps internal dispositions to user-facing labels:

```typescript
import { previewFirePolicies } from "@/ui/src/lib/cron-fires";

const fires = previewFirePolicies("* * * * *", "coalesce_if_active");
console.log(fires.map(f => f.disposition));
// Output: ["queued", "coalesced", "coalesced"] — second and third triggers merge

```

## Orphan Run Detection and Recovery

An **orphan run** has `status = 'running'` in the database but its underlying OS process or process group no longer exists. These occur after server crashes, `kill -9` terminations, or container evictions.

### Startup Reaper: First Line of Defense

Before any timer ticks begin, the server runs a **startup orphan reaper** (`server/src/index.ts#L1024‑L1055`). This sequencing is deliberate: stale runs must be cleared so new triggers cannot accidentally coalesce with them.

The reaper executes four steps:

1. **Identify orphans** — Query `heartbeatRuns` for rows whose `processPid`/`processGroupId` no longer correspond to live OS processes.
2. **Mark terminal** — Update status with error codes `orphaned_running_run` or `orphaned_running_run_issue_terminal`.
3. **Release execution locks** — Clear `executionRunId`, `executionAgentNameKey`, `executionLockedAt` columns (`server/src/services/heartbeat.ts#L16455‑L16457`).
4. **Queue process-loss retry** — Optionally enqueue a retry so the agent resumes after clean start (`server/src/services/heartbeat.ts#L10633‑L10639`).

### Periodic Orphan Sweep

A **5-minute periodic sweep** (`server/src/index.ts#L1266` comment) captures orphans that appear post-startup. This background task maintains system health during long-running server lifetimes.

### Process-Group Cleanup Events

When heartbeat detects a dead process group, it emits a synthetic event:

```json
{ "kind": "orphaned_process_group_cleanup", "runId": "...", "agentId": "..." }

```

This event originates at `server/src/services/heartbeat.ts#L13197` and propagates to the **plugin-worker-manager**, which closes lingering SSE streams (`server/src/services/plugin-worker-manager.ts#L1228`).

### Lock Table Recovery

Orphaned runs hold **execution locks** on issues that would otherwise block other agents indefinitely. The reaper explicitly clears these in `server/src/services/routines.ts#L290‑L306`, unblocking the work queue.

## End-to-End Execution Flow

The complete lifecycle integrates both mechanisms:

1. **Trigger arrival** → `enqueueWakeup` creates a `queued` run (or returns `coalesced` per policy).
2. **Process launch** → Run transitions to `running` with `processPid` and `processGroupId` recorded.
3. **Normal completion** → Run terminates, locks release, no cleanup needed.
4. **Orphan detection** → Startup reaper or periodic sweep identifies stale `running` rows, marks terminal, releases locks, optionally retries.

### Practical Code Examples

**Enqueue with default coalescing:**

```typescript
import { enqueueWakeup } from "@paperclipai/server/src/services/heartbeat.js";

const run = await enqueueWakeup(agentId, {
  action: "my.routine.fire",
  payload: { task: "generate-report" }
});
// "queued" if no live run exists, "coalesced" if merged into existing run
console.log(run.status);

```

**Manual orphan recovery:**

```typescript
import { getRunById, markRunTerminal, enqueueWakeup } 
  from "@paperclipai/server/src/services/heartbeat.js";

const orphan = await getRunById(orphanRunId);
if (orphan?.status === "running" && !(await processExists(orphan.processPid))) {
  await markRunTerminal(orphan.id, "orphaned_running_run");
  await enqueueWakeup(orphan.agentId, { 
    action: "retry.after.orphan",
    originalRunId: orphan.id 
  });
}

```

**Preview coalescing behavior in UI:**

```typescript
import { previewFirePolicies } from "@/ui/src/lib/cron-fires";

const schedule = "*/5 * * * *";  // Every 5 minutes
const fires = previewFirePolicies(schedule, "coalesce_if_active");
// Visualize which fires queue fresh vs. merge into active runs

```

## Summary

- **Queue coalescing** via `coalesce_if_active` (default) prevents duplicate heartbeat runs by merging triggers into live executions.
- **Orphan detection** operates at startup (`server/src/index.ts#L1024‑L1055`) and every 5 minutes to identify and terminate stale `running` rows.
- **Lock recovery** ensures orphaned runs cannot indefinitely block other agents from executing the same work.
- **Retry semantics** allow automatic resumption after process loss through process-loss retry enqueuing.
- **Observability** is built in: all orphan operations log explicit error codes and emit cleanup events for downstream handling.

## Frequently Asked Questions

### How does Paperclip AI prevent duplicate routine executions?

**Paperclip AI applies concurrency policies at trigger time.** The default `coalesce_if_active` policy checks `server/src/services/routines.ts#L262‑L282` for existing `queued` or `running` runs; if found, the trigger merges into that run rather than creating a new row. This deduplication happens transparently before any OS process spawns.

### What happens to heartbeat runs when the server crashes?

**Orphaned runs are reclaimed at the next startup.** The reaper in `server/src/index.ts#L1024‑L1055` executes before timer initialization, finding `running` rows with dead PIDs, marking them terminal with error code `orphaned_running_run`, releasing execution locks, and optionally queueing retries. A 5-minute periodic sweep handles crashes during runtime.

### Can I disable heartbeat queue coalescing for a specific routine?

**Yes, set the concurrency policy to `always_enqueue`.** This is configured where routines are declared—typically [`server/src/services/plugin-managed-routines.ts`](https://github.com/paperclipai/paperclip/blob/main/server/src/services/plugin-managed-routines.ts) or pipeline definitions. With this policy, [`server/src/services/routines.ts`](https://github.com/paperclipai/paperclip/blob/main/server/src/services/routines.ts) bypasses live-run checks and always inserts new `heartbeatRuns` rows.

### How does Paperclip AI clean up resources from orphaned runs?

**Through lock clearing and synthetic events.** The reaper updates execution-lock columns (`executionRunId`, etc.) in `server/src/services/routines.ts#L290‑L306` and emits `orphaned_process_group_cleanup` events from `server/src/services/heartbeat.ts#L13197`. The plugin-worker-manager consumes these to close SSE streams at `server/src/services/plugin-worker-manager.ts#L1228`.