How Atomic Task Checkout Prevents Double-Work in Distributed Agent Execution

Atomic task checkout prevents double-work by using a database-backed single-writer lock where only one agent can successfully claim an issue at a time, with conflicting attempts receiving an immediate 409 Conflict response.

The paperclip open-source project implements a robust distributed coordination mechanism for AI agents. When multiple agents poll for work simultaneously, the system must guarantee that no two agents process the same issue—this is the "double-work" problem that atomic task checkout solves.

The Lock Mechanism: checkoutRunId Column

At the heart of the system is the checkoutRunId column on the issues table. This field acts as a lease token: when populated, it contains the unique run ID of the agent currently processing that issue.

The checkout flow in server/src/services/issues.ts follows five atomic steps:

  1. Row-level lock acquisition – The transaction begins with SELECT ... FOR UPDATE to block concurrent modifications
  2. Same-run validation – The sameRunLock helper verifies the lock is free or already held by the requesting run
  3. Conditional update – An UPDATE ... WHERE clause atomically writes the new run ID only if validation passes
  4. Conflict detection – Zero affected rows triggers a ConflictError with HTTP 409
  5. Lock release – Completion clears checkoutRunId to NULL, allowing re-acquisition

The Same-Run Validation Logic

The sameRunLock function enforces ownership semantics:

// server/src/services/issues.ts#L73-L76
function sameRunLock(checkoutRunId: string | null, actorRunId: string | null) {
  if (actorRunId) return checkoutRunId === actorRunId;
  return checkoutRunId == null;
}

This helper serves two purposes. It allows idempotent retries when the same run ID attempts checkout again (useful for network timeouts), while rejecting cross-run claims when another agent already holds the lock.

Database-Level Atomicity

The actual checkout uses a conditional UPDATE that cannot be interleaved:

// server/src/services/issues.ts#L8033-L8050 (simplified)
const result = await db
  .update(issues)
  .set({ checkoutRunId: actorRunId })
  .where(and(
    eq(issues.id, issueId),
    or(
      isNull(issues.checkoutRunId),
      eq(issues.checkoutRunId, actorRunId)
    )
  ))
  .returning()
  .execute();

if (result.length === 0) {
  throw new ConflictError('Issue already checked out');
}

The WHERE clause encodes the sameRunLock logic directly in SQL. Because PostgreSQL evaluates this atomically, there is no race window between the read check and the write operation.

Why the 409 Conflict Matters

When the conditional update matches zero rows, the service responds with 409 Conflict rather than silent failure or polling loops. This explicit signal allows agents to implement intelligent backoff strategies.

Client code in ui/src/api/issues.ts handles this explicitly:

// ui/src/api/issues.ts
export const checkout = async (id, agentId, expectedStatuses, checkoutRunId) => {
  const res = await api.post<Issue>(`/issues/${id}/checkout`, {
    agentId,
    expectedStatuses,
    checkoutRunId,
  });

  if (!res.ok() && res.status() === 409) {
    throw new Error('Checkout conflict');
  }
  return res.json();
};

Agent implementations can then distinguish between transient conflicts (retry later) and actual errors (escalate).

Lock Release and Cleanup

The lock lifecycle completes when the run finishes. The service in server/src/services/issues.ts#L8400-L8420 executes:

UPDATE issues SET checkoutRunId = NULL WHERE id = $issueId

This unblocks the issue for the next available agent. The release is typically triggered by run completion handlers regardless of success, failure, or cancellation status—preventing zombie locks from orphaning work.

Practical Agent Integration Pattern

A complete agent implementation combines checkout with work execution and conflict handling:

async function tryToWorkOnIssue(issueId, agent) {
  try {
    const issue = await checkout(issueId, agent.id, ['open'], agent.runId);
    // Exclusive ownership granted—perform work
    await doWork(issue);
    // Lock released automatically on completion
  } catch (e) {
    if (e.message.includes('conflict')) {
      // Another agent won—exponential backoff and retry
      await sleep(5000 * Math.random());
      await tryToWorkOnIssue(issueId, agent);
    } else {
      throw e;
    }
  }
}

This pattern ensures exactly-once semantics for issue processing without requiring external coordination services like Redis or ZooKeeper.

Key Source Files

File Responsibility
server/src/services/issues.ts Core checkout logic, sameRunLock validation, conditional updates
server/src/routes/issues.ts HTTP endpoint POST /issues/:id/checkout
packages/db/src/schema/issues.ts checkoutRunId column definition in Drizzle schema
ui/src/api/issues.ts Client-side checkout wrapper with 409 handling

Summary

  • Atomic task checkout uses database transactions and conditional writes to guarantee single-agent ownership
  • The checkoutRunId column in packages/db/src/schema/issues.ts serves as the distributed lock token
  • sameRunLock validation enables both exclusive acquisition and idempotent retries by the same run
  • Zero-row updates from the conditional WHERE clause trigger immediate 409 Conflict responses
  • No external lock service is required—PostgreSQL's row-level locking provides linearizable semantics

Frequently Asked Questions

What happens if two agents try to checkout the same issue simultaneously?

Only one succeeds. The first agent to commit its transaction wins and writes its actorRunId to checkoutRunId. The second agent's conditional UPDATE matches zero rows because the WHERE clause requires checkoutRunId IS NULL OR checkoutRunId = $actorRunId, and neither condition holds. This results in a 409 Conflict error returned immediately.

Can the same agent retry checkout if its first attempt times out?

Yes. The sameRunLock function in server/src/services/issues.ts explicitly allows this: if checkoutRunId already equals the requesting actorRunId, the checkout succeeds. This idempotency prevents false conflicts from network timeouts without creating a security hole for cross-run takeover.

What prevents zombie locks if an agent crashes mid-execution?

The checkoutRunId column contains the run ID, not the agent ID. When a run completes—successfully, with errors, or via cancellation—the completion handler executes UPDATE issues SET checkoutRunId = NULL. Since crashed processes cannot continue their run, the cleanup logic (typically a separate reconciler or timeout-based expiry in production deployments) eventually releases the lock.

Why use database locking instead of Redis or another distributed lock service?

Paperclip achieves linearizable semantics without additional infrastructure. PostgreSQL's FOR UPDATE row locks provide the same guarantees as a dedicated lock service while keeping the architecture simpler. The conditional UPDATE WHERE pattern is also portable across most SQL databases, reducing operational complexity for self-hosted deployments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →