How Paperclip Task Watchdog Recovery Handles Execution Failures: A Deep Dive
Paperclip's task watchdog recovery system detects stopped execution subtrees, generates deterministic fingerprints, and automatically creates or reopens watchdog issues to resume failed tasks.
Paperclip's execution engine relies on a task-watchdog subsystem to ensure that issue subtrees remain healthy. When execution stalls, the system must surface failures, track decisions, and eventually resume work. This article explains how the task-watchdogs.ts and recovery/service.ts modules collaborate to handle execution failures end-to-end.
Understanding the Task Watchdog Architecture
The task watchdog operates as a monitoring layer above issue execution. It periodically evaluates whether a watched subtree has active execution paths. When no live runs, queued wake-ups, or scheduled retries exist, the subtree is classified as "stopped" — triggering the recovery flow.
The watchdog does not attempt recovery directly. Instead, it delegates to a dedicated recovery service that makes durable decisions about issue creation, reopening, and wake-up scheduling. This separation of concerns ensures that recovery logic is centralized and auditable.
Step 1: Detecting a Stopped Subtree with classifyTaskWatchdogSubtree
The classification phase begins in task-watchdogs.ts. The classifyTaskWatchdogSubtree function (lines 71-85) builds a snapshot of the watched subtree and determines its execution state:
// Simplified representation of the classification result
{
state: "stopped",
stopFingerprint: "sha256:abc123...", // Deterministic identifier
materialLeaves: [...], // Terminal nodes in the subtree
pendingWaits: [...] // Awaiting interactions/approvals
}
A stopped state indicates that no forward progress is possible without external intervention. The function returns a TaskWatchdogClassifierResult that carries both the state and a stop fingerprint for deduplication.
Step 2: Generating a Stable Stop Fingerprint
Before invoking recovery, the watchdog computes a deterministic identifier via stableStopFingerprint (lines 103-113 in task-watchdogs.ts):
// SHA-256 hash over material leaves + pending waits
const fingerprint = stableStopFingerprint({
materialLeaves: snapshot.leaves,
pendingWaits: snapshot.waits
});
This fingerprint serves two critical purposes:
- Deduplication – The same stopped state observed multiple times produces an identical hash, preventing redundant recovery actions.
- Auditability – Each recovery decision links to an immutable snapshot of the failure condition.
Step 3: Delegating to the Recovery Service
With classification complete, the watchdog hands control to recoveryService.recordWatchdogDecision. The call site in task-watchdogs.ts (around line 1040) passes three key inputs:
- The watchdog row being evaluated
- The source issue that owns the watchdog
- The classification result including the stop fingerprint
This boundary ensures that the watchdog (concerned with detection) remains separate from the recovery service (concerned with durable action).
Step 4: Managing the Watchdog Issue
The recovery service's central function, recordWatchdogDecision, first locates or creates a suitable issue for tracking. The ensureReusableWatchdogIssue helper (lines 1223-1275 in recovery/service.ts) implements the following logic:
- Search – Call
findTaskWatchdogIssueto locate an existing watchdog issue for this source - Evaluate – If found, check whether the issue is terminal or needs fresh wake-up
- Reopen or update – Reopen terminal issues; update fingerprints on active ones
This reuse strategy prevents issue sprawl while ensuring that recurring failures accumulate context in a single location.
Step 5: Documenting the Failure Context
Human operators need visibility into why execution stopped. The buildStoppedFingerprintComment function (lines 642-650 in recovery/service.ts) generates a structured comment attached to the watchdog issue:
- Up to 12 stopped leaves with their specific failure modes
- Pending interactions or approvals blocking progress
- The fingerprint hash for cross-referencing with logs
This comment transforms the raw stop snapshot into an actionable incident record.
Step 6: Persisting the Recovery Decision
Durability is enforced through two writes in recordWatchdogDecision (around line 850 in recovery/service.ts):
// 1. Update the watchdog row with observed state
await updateWatchdog(watchdogId, {
lastObservedFingerprint: fingerprint,
lastObservedStopSnapshot: snapshot
});
// 2. Create auditable activity log
await logActivity({
type: "watchdog_stop_observed",
watchdogId,
fingerprint,
decision: "created_issue" | "reopened_issue" | "updated_fingerprint"
});
The watchdog activity log provides a complete audit trail for post-mortem analysis and compliance purposes.
Step 7: Enqueuing Automatic Recovery
Recovery completes only when execution resumes. Inside recordWatchdogDecision (around line 860 in recovery/service.ts), the service calls enqueueWakeup to schedule task continuation:
// Idempotent wake-up request
await enqueueWakeup({
issueId: watchdogIssueId,
idempotencyKey: taskWatchdogWakeIdempotencyKey(watchdogId, fingerprint)
});
The wake-up mechanism respects any rate limits or backoff policies configured in the queue, ensuring that recovery does not overwhelm downstream systems.
Step 8: Enforcing Idempotency
Duplicate watchdog evaluations must not spawn duplicate wake-ups. The taskWatchdogWakeIdempotencyKey function (lines 160-167 in task-watchdogs.ts) constructs a unique key:
function taskWatchdogWakeIdempotencyKey(
watchdogId: string,
fingerprint: string
): string {
return `task-watchdog-wake:${watchdogId}:${fingerprint}`;
}
This key is passed to enqueueWakeup, which guarantees exactly-once execution of the recovery action for each unique stopped state.
Execution Failure Recovery Flow Summary
The complete failure handling sequence operates as follows:
- Detection –
classifyTaskWatchdogSubtreeidentifies stopped subtrees - Fingerprinting –
stableStopFingerprintcreates a deterministic hash - Delegation – Watchdog calls
recoveryService.recordWatchdogDecision - Issue resolution –
ensureReusableWatchdogIssuefinds or creates tracking issue - Documentation –
buildStoppedFingerprintCommentadds human-readable context - Persistence – Watchdog row and activity log are updated
- Recovery –
enqueueWakeupschedules task resumption - Safety – Idempotency key prevents duplicate actions
Summary
Paperclip's task watchdog recovery transforms silent execution failures into tracked, actionable incidents:
- Deterministic detection via SHA-256 fingerprints prevents duplicate handling
- Centralized recovery logic in
recovery/service.tsensures consistent decision-making - Reusable watchdog issues consolidate failure context and reduce noise
- Idempotent wake-ups guarantee safe automatic recovery without side effects
- Full audit logging supports operational review and compliance requirements
The system's design reflects a core principle: every stopped subtree must surface as a concrete, resumable issue rather than disappearing into log noise.
Frequently Asked Questions
What triggers the task watchdog recovery flow?
The recovery flow triggers when classifyTaskWatchdogSubtree returns state: "stopped" — meaning no live runs, queued wake-ups, or scheduled retries exist for any issue in the watched subtree. This classification occurs during periodic watchdog evaluations.
How does Paperclip prevent duplicate recovery actions for the same failure?
The stableStopFingerprint function generates a deterministic SHA-256 hash of the stopped state. This fingerprint becomes part of the idempotency key passed to enqueueWakeup, ensuring that repeated evaluations with identical snapshots produce exactly one wake-up request.
What happens if a watchdog issue already exists for a recurring failure?
The ensureReusableWatchdogIssue helper locates existing issues via findTaskWatchdogIssue. Terminal issues are reopened with fresh context; active issues receive updated fingerprints. This reuse strategy prevents issue proliferation while maintaining historical context.
Where is the recovery decision recorded for operational audit?
Two durables are updated in recordWatchdogDecision: the watchdog row (lastObservedFingerprint, lastObservedStopSnapshot) and the watchdog activity log via logActivity. Together, these provide complete traceability from detection through resolution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →