How the SyncRetryScheduler Handles Durable Pending-Sync Retries After Container Restarts

The SyncRetryScheduler persists a single pending-sync intent per backend in Durable Object storage, then recovers post-command pulls after computerd container restarts by reconciling watermarks, validating runtime identity, and retrying with bounded exponential backoff.

The @cloudflare/computer workspace uses the SyncRetryScheduler to guarantee that post-command synchronization eventually succeeds, even when the backing container process crashes or restarts between command execution and sync completion. This article walks through the exact implementation in the Cloudflare Computer repository, tracing how intents are created, persisted, recovered, and resolved across container lifecycles.

What the SyncRetryScheduler Does

The SyncRetryScheduler interface (packages/computer/src/workspace.ts defines the durable boundary for retry intents. Unlike typical workspace operations, the host (Durable Object) owns persistence because the workspace library cannot set Durable Object alarms directly.

The scheduler stores three pieces of state per backend:

  • get(backend) — retrieve the current pending intent
  • schedule(intent) — persist a new or updated intent
  • clear(backend) — remove the intent after success or permanent failure

Tests demonstrate a simple in-memory implementation in MemoryRetryScheduler (packages/computer/src/retry.test.ts, while production deployments use Durable Object KV storage.

Creating Retry Intents After Failed Pulls

When CommandExecutor finishes a command via exec, it schedules a pending sync if the post-command pull fails.

The #schedulePendingSync method (L707-718) creates the intent:

// From workspace.ts — simplified structure
interface PendingSyncIntent {
  backend: string;           // target backend id
  runtime?: string;          // runtime that produced the command
  attempt: number;           // current retry count (starts at 0)
  notBefore: number;         // timestamp for next attempt (ms since epoch)
}

The #retryIntent helper (L221-231) computes notBefore using bounded exponential backoff:

  • initialDelayMs — first retry wait time (default 1 second)
  • maxDelayMs — cap on retry interval (default 60 seconds)
  • maxAttempts — hard limit before exhaustion
// Example: Creating a retry intent with exponential backoff
const intent = this.#retryIntent(backendId, runtimeId, attemptCount);
// attempt 0 → notBefore = now + 1_000ms
// attempt 1 → notBefore = now + 2_000ms
// attempt 2 → notBefore = now + 4_000ms
// ...capped at maxDelayMs
await this.retryScheduler?.schedule(intent);

Recovering State After Container Restart

When a Durable Object re-instantiates with a fresh container, the workspace reconnects via #handleFor. Before returning a handle, it reconciles watermarks to detect state loss.

The reconcileWatermarks function (packages/rpc/src/sync-driver.ts compares the host's persisted cursor against the remote's actual log length. If the remote's log is shorter (indicating container restart with lost memory), the sync resets to revision 0, establishing a clean baseline for the retry.

// Conceptual flow during reconnection
const handle = await ws.runtime.exec("npm test", { backend: "container-shell" });
// ...container crashes, DO persists intent...
// ...new container starts, DO re-instantiates...

// Inside #handleFor — automatic recovery
await this.#reconcileWatermarks(backend);  // sync-driver.ts L95-115
// Detects: remote log shorter than host cursor
// Action: reset to rev 0, prepare for retry from clean state

Executing Durable Retries

The Durable Object's alarm (or external trigger) calls retryPendingSync(backendId) to attempt resolution. This method follows a strict state machine:

Step 1: Validate Intent Eligibility

Lines 24-41 in workspace.ts fetch the stored intent and check exhaustion:

const intent = await this.retryScheduler?.get(backendId);
if (!intent || Date.now() < intent.notBefore) {
  return { status: "pending", played: 0, skipped: 0 };
}
if (intent.attempt >= this.#retry.maxAttempts) {
  await this.retryScheduler?.clear(backendId);
  return { status: "exhausted", played: 0, skipped: 0 };
}

Step 2: Attempt Pull with Ordering Guarantees

Lines 83-99 wrap the pull in withSpan and #serialize, ensuring FIFO ordering per-backend:

return await this.#serialize(backendId, async () => {
  return await withSpan("sync_retry", async (span) => {
    span.setAttribute("backend", backendId);
    const { played, skipped } = await this.#pullResolved(...);
    // ...
  });
});

Step 3: Handle Success, Loss, or Reschedule

Three outcomes are possible:

Outcome Condition Action
Success Pull completes Clear intent, return {status: "complete"} (L44-53)
Lost Runtime EEXEC_LOST error Clear intent, return {status: "lost"} (L55-66)
Retryable Failure Other error Reschedule with attempt+1, return {status: "pending"} (L76-78)

The lost runtime case (L55-66) is critical: if the container that produced the original command no longer exists, retrying the pull is futile. The scheduler clears the intent immediately, allowing a fresh runtime to establish new state.

Complete Implementation Example

// packages/computer/src/workspace.ts — durable scheduler integration

import { Workspace, SyncRetryScheduler } from "@cloudflare/computer";

class DurableObjectScheduler implements SyncRetryScheduler {
  constructor(private storage: DurableObjectStorage) {}

  async get(backend: string): Promise<PendingSyncIntent | undefined> {
    return this.storage.get(`pendingSync:${backend}`);
  }

  async schedule(intent: PendingSyncIntent): Promise<void> {
    await this.storage.put(`pendingSync:${intent.backend}`, intent);
  }

  async clear(backend: string): Promise<void> {
    await this.storage.delete(`pendingSync:${backend}`);
  }
}

// Workspace configuration
const ws = new Workspace({
  storage: ctx.storage,
  backends: [new ContainerBackend()],
  retryScheduler: new DurableObjectScheduler(ctx.storage),
  retry: {
    initialDelayMs: 1_000,
    maxDelayMs: 60_000,
    maxAttempts: 5
  },
});
// DO alarm handler — triggers durable retries
export default {
  async alarm(state: DurableObjectState, env: Env) {
    const ws = await getWorkspace(state, env);
    
    for (const backend of ws.getBackendIds()) {
      const result = await ws.retryPendingSync(backend);
      
      if (result.status === "pending") {
        // Reschedule alarm for notBefore timestamp
        const intent = await ws.retryScheduler!.get(backend);
        state.setAlarm(intent!.notBefore);
      }
      // "complete", "lost", or "exhausted" — no further alarm needed
    }
  }
};

Test Coverage for Durable Behavior

The retry.test.ts suite validates three critical scenarios:

Lost runtime clearing (L49-62): A WorkspaceTransportError during pull produces "lost" status and removes the stale intent, preventing infinite retries against a dead container.

Stale intent replacement (L71-84): When a new runtime successfully reconnects, it replaces any pending intent tied to the previous runtime, ensuring coherence.

Exponential backoff with exhaustion (L28-54): Confirms the delay schedule and maxAttempts enforcement, with visible exhaustion status.

Summary

  • Single intent per backend: The SyncRetryScheduler persists one PendingSyncIntent per backend, bounded by the Durable Object's storage.
  • Runtime-tied validity: Each intent tracks the runtime id that produced the command; retries only proceed while that container remains alive.
  • Watermark reconciliation: On every reconnection, reconcileWatermarks detects container restarts and resets sync baselines as needed.
  • Three terminal states: Success (complete), permanent failure (lost runtime), and exhaustion (exhausted) all properly clear persisted state.
  • FIFO retry execution: The #serialize semaphore guarantees that retries don't interleave with concurrent pushes or pulls.

Frequently Asked Questions

How does the SyncRetryScheduler know if a container has restarted?

The scheduler itself doesn't detect restarts directly. Instead, the reconcileWatermarks function in sync-driver.ts compares the host's persisted cursor against the remote container's actual log length. A shorter remote log indicates state loss, triggering a reset to revision 0 before any retry proceeds.

What happens if a retry exceeds maxAttempts?

When intent.attempt >= maxAttempts, retryPendingSync clears the intent from storage and returns {status: "exhausted"}. The pending sync is abandoned; the application must handle this as a permanent failure, typically by surfacing the error to operators or triggering manual intervention.

Can multiple runtimes race to retry the same pending sync?

No. The #serialize semaphore in workspace.ts ensures FIFO ordering per backend. Additionally, each intent is tied to a specific runtime id. If a new runtime schedules its own pending sync, it replaces any stale intent from the previous runtime (verified in retry.test.ts.

Is the SyncRetryScheduler required for basic workspace operation?

No. The retryScheduler parameter is optional. Without it, post-command pulls that fail return immediately with sync.status === "pending", and the application must implement its own retry logic. The scheduler is specifically for applications requiring durable, automatic retries across Durable Object invocations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →