# How Paperclip Task Watchdog Recovery Handles Execution Failures: A Deep Dive

> Learn how Paperclip task watchdog recovery handles execution failures by detecting stopped subtrees, creating fingerprints, and resuming failed tasks automatically.

- Repository: [Paperclip/paperclip](https://github.com/paperclipai/paperclip)
- Tags: deep-dive
- Published: 2026-08-16

---

**Paperclip's task watchdog recovery system detects stopped execution subtrees, generates deterministic fingerprints, and automatically creates or reopens watchdog issues to resume failed tasks.**

Paperclip's execution engine relies on a **task-watchdog subsystem** to ensure that issue subtrees remain healthy. When execution stalls, the system must surface failures, track decisions, and eventually resume work. This article explains how the [`task-watchdogs.ts`](https://github.com/paperclipai/paperclip/blob/main/task-watchdogs.ts) and [`recovery/service.ts`](https://github.com/paperclipai/paperclip/blob/main/recovery/service.ts) modules collaborate to handle execution failures end-to-end.

## Understanding the Task Watchdog Architecture

The task watchdog operates as a monitoring layer above issue execution. It periodically evaluates whether a watched subtree has active execution paths. When no live runs, queued wake-ups, or scheduled retries exist, the subtree is classified as **"stopped"** — triggering the recovery flow.

The watchdog does not attempt recovery directly. Instead, it delegates to a dedicated **recovery service** that makes durable decisions about issue creation, reopening, and wake-up scheduling. This separation of concerns ensures that recovery logic is centralized and auditable.

## Step 1: Detecting a Stopped Subtree with classifyTaskWatchdogSubtree

The classification phase begins in [`task-watchdogs.ts`](https://github.com/paperclipai/paperclip/blob/main/task-watchdogs.ts). The `classifyTaskWatchdogSubtree` function (lines 71-85) builds a snapshot of the watched subtree and determines its execution state:

```typescript
// Simplified representation of the classification result
{
  state: "stopped",
  stopFingerprint: "sha256:abc123...", // Deterministic identifier
  materialLeaves: [...],              // Terminal nodes in the subtree
  pendingWaits: [...]                 // Awaiting interactions/approvals
}

```

A **stopped state** indicates that no forward progress is possible without external intervention. The function returns a `TaskWatchdogClassifierResult` that carries both the state and a **stop fingerprint** for deduplication.

## Step 2: Generating a Stable Stop Fingerprint

Before invoking recovery, the watchdog computes a deterministic identifier via `stableStopFingerprint` (lines 103-113 in [`task-watchdogs.ts`](https://github.com/paperclipai/paperclip/blob/main/task-watchdogs.ts)):

```typescript
// SHA-256 hash over material leaves + pending waits
const fingerprint = stableStopFingerprint({
  materialLeaves: snapshot.leaves,
  pendingWaits: snapshot.waits
});

```

This fingerprint serves two critical purposes:

- **Deduplication** – The same stopped state observed multiple times produces an identical hash, preventing redundant recovery actions.
- **Auditability** – Each recovery decision links to an immutable snapshot of the failure condition.

## Step 3: Delegating to the Recovery Service

With classification complete, the watchdog hands control to `recoveryService.recordWatchdogDecision`. The call site in [`task-watchdogs.ts`](https://github.com/paperclipai/paperclip/blob/main/task-watchdogs.ts) (around line 1040) passes three key inputs:

1. The **watchdog row** being evaluated
2. The **source issue** that owns the watchdog
3. The **classification result** including the stop fingerprint

This boundary ensures that the watchdog (concerned with detection) remains separate from the recovery service (concerned with durable action).

## Step 4: Managing the Watchdog Issue

The recovery service's central function, `recordWatchdogDecision`, first locates or creates a suitable issue for tracking. The `ensureReusableWatchdogIssue` helper (lines 1223-1275 in [`recovery/service.ts`](https://github.com/paperclipai/paperclip/blob/main/recovery/service.ts)) implements the following logic:

- **Search** – Call `findTaskWatchdogIssue` to locate an existing watchdog issue for this source
- **Evaluate** – If found, check whether the issue is terminal or needs fresh wake-up
- **Reopen or update** – Reopen terminal issues; update fingerprints on active ones

This reuse strategy prevents issue sprawl while ensuring that recurring failures accumulate context in a single location.

## Step 5: Documenting the Failure Context

Human operators need visibility into why execution stopped. The `buildStoppedFingerprintComment` function (lines 642-650 in [`recovery/service.ts`](https://github.com/paperclipai/paperclip/blob/main/recovery/service.ts)) generates a structured comment attached to the watchdog issue:

- Up to **12 stopped leaves** with their specific failure modes
- **Pending interactions or approvals** blocking progress
- The **fingerprint hash** for cross-referencing with logs

This comment transforms the raw stop snapshot into an actionable incident record.

## Step 6: Persisting the Recovery Decision

Durability is enforced through two writes in `recordWatchdogDecision` (around line 850 in [`recovery/service.ts`](https://github.com/paperclipai/paperclip/blob/main/recovery/service.ts)):

```typescript
// 1. Update the watchdog row with observed state
await updateWatchdog(watchdogId, {
  lastObservedFingerprint: fingerprint,
  lastObservedStopSnapshot: snapshot
});

// 2. Create auditable activity log
await logActivity({
  type: "watchdog_stop_observed",
  watchdogId,
  fingerprint,
  decision: "created_issue" | "reopened_issue" | "updated_fingerprint"
});

```

The **watchdog activity log** provides a complete audit trail for post-mortem analysis and compliance purposes.

## Step 7: Enqueuing Automatic Recovery

Recovery completes only when execution resumes. Inside `recordWatchdogDecision` (around line 860 in [`recovery/service.ts`](https://github.com/paperclipai/paperclip/blob/main/recovery/service.ts)), the service calls `enqueueWakeup` to schedule task continuation:

```typescript
// Idempotent wake-up request
await enqueueWakeup({
  issueId: watchdogIssueId,
  idempotencyKey: taskWatchdogWakeIdempotencyKey(watchdogId, fingerprint)
});

```

The wake-up mechanism respects any rate limits or backoff policies configured in the queue, ensuring that recovery does not overwhelm downstream systems.

## Step 8: Enforcing Idempotency

Duplicate watchdog evaluations must not spawn duplicate wake-ups. The `taskWatchdogWakeIdempotencyKey` function (lines 160-167 in [`task-watchdogs.ts`](https://github.com/paperclipai/paperclip/blob/main/task-watchdogs.ts)) constructs a unique key:

```typescript
function taskWatchdogWakeIdempotencyKey(
  watchdogId: string,
  fingerprint: string
): string {
  return `task-watchdog-wake:${watchdogId}:${fingerprint}`;
}

```

This key is passed to `enqueueWakeup`, which guarantees **exactly-once execution** of the recovery action for each unique stopped state.

## Execution Failure Recovery Flow Summary

The complete failure handling sequence operates as follows:

1. **Detection** – `classifyTaskWatchdogSubtree` identifies stopped subtrees
2. **Fingerprinting** – `stableStopFingerprint` creates a deterministic hash
3. **Delegation** – Watchdog calls `recoveryService.recordWatchdogDecision`
4. **Issue resolution** – `ensureReusableWatchdogIssue` finds or creates tracking issue
5. **Documentation** – `buildStoppedFingerprintComment` adds human-readable context
6. **Persistence** – Watchdog row and activity log are updated
7. **Recovery** – `enqueueWakeup` schedules task resumption
8. **Safety** – Idempotency key prevents duplicate actions

## Summary

Paperclip's task watchdog recovery transforms silent execution failures into tracked, actionable incidents:

- **Deterministic detection** via SHA-256 fingerprints prevents duplicate handling
- **Centralized recovery logic** in [`recovery/service.ts`](https://github.com/paperclipai/paperclip/blob/main/recovery/service.ts) ensures consistent decision-making
- **Reusable watchdog issues** consolidate failure context and reduce noise
- **Idempotent wake-ups** guarantee safe automatic recovery without side effects
- **Full audit logging** supports operational review and compliance requirements

The system's design reflects a core principle: **every stopped subtree must surface as a concrete, resumable issue** rather than disappearing into log noise.

## Frequently Asked Questions

### What triggers the task watchdog recovery flow?

The recovery flow triggers when `classifyTaskWatchdogSubtree` returns `state: "stopped"` — meaning no live runs, queued wake-ups, or scheduled retries exist for any issue in the watched subtree. This classification occurs during periodic watchdog evaluations.

### How does Paperclip prevent duplicate recovery actions for the same failure?

The `stableStopFingerprint` function generates a deterministic SHA-256 hash of the stopped state. This fingerprint becomes part of the idempotency key passed to `enqueueWakeup`, ensuring that repeated evaluations with identical snapshots produce exactly one wake-up request.

### What happens if a watchdog issue already exists for a recurring failure?

The `ensureReusableWatchdogIssue` helper locates existing issues via `findTaskWatchdogIssue`. Terminal issues are reopened with fresh context; active issues receive updated fingerprints. This reuse strategy prevents issue proliferation while maintaining historical context.

### Where is the recovery decision recorded for operational audit?

Two durables are updated in `recordWatchdogDecision`: the **watchdog row** (`lastObservedFingerprint`, `lastObservedStopSnapshot`) and the **watchdog activity log** via `logActivity`. Together, these provide complete traceability from detection through resolution.