How Watermarks Enable Incremental Sync in Cloudflare Computer
Watermarks enable incremental sync by persisting revision cursors in SQLite, allowing Cloudflare Computer to resume synchronization from the exact point where the last successful transfer ended, transmitting only new changes rather than full dataset snapshots.
Cloudflare Computer synchronizes Durable Object (DO)-backed SQLite databases with remote backends like the computerd FUSE daemon through a delta-based protocol. The system achieves efficient, restart-safe synchronization by storing watermarks in special SQLite tables that act as durable bookmarks for the sync stream.
Understanding Watermark Types
The sync driver maintains two distinct watermark types in packages/dofs/src/sync/watermarks.ts to track bilateral data flow. Each watermark is scoped to a specific backend identifier, allowing multiple remotes to sync with the same DO independently.
Push Revision Watermark (pushRev)
The pushRev watermark stores the highest revision number that the local DO has successfully pushed to a specific remote backend. Persisted in the _vfs_watermark table with key 'pushRev', this value ensures that subsequent pushes transmit only mutations with rev > pushRev. The driver reads this value via readWatermark(db, "pushRev", backend) and updates it via writeWatermark() only after the remote acknowledges successful receipt.
Fetch Cursor Watermark (fetchRev)
The fetch cursor tracks how much data the DO has received from the remote. Unlike the simple integer used for pushes, this watermark combines a revision number with an optional path string stored across _vfs_watermark and _vfs_fetch_cursor tables. When path is null, the revision is fully drained; when populated, it indicates a partial resume point within that revision. The driver manipulates this structure through readFetchCursor(db, backend) and writeFetchCursor(db, { rev, path }, backend).
Incremental Pull: Fetching Remote Changes
The pull process (DO ← Remote) leverages the fetch cursor to request only unseen data. According to the implementation in packages/rpc/src/sync-driver.ts, specifically within pullOnceImpl, the driver executes a four-phase protocol:
- Retrieve cursor:
readFetchCursorreturns the last processed(rev, path)tuple, which becomes theafterparameter for the next request. - Request delta:
remote.fetchChanges({ after })streams entries occurring strictly after the cursor position. - Apply batch: The driver applies received entries in batches bounded by
PULL_BATCH_SIZE(256 entries), keeping memory usage constant regardless of total backlog size. - Advance cursor: Upon successful application,
writeFetchCursorpersists the new position, ensuring idempotent resume capability after crashes or network interruptions.
If the remote reports a cursor that predates the DO's local watermark—indicating the remote "forgot" revisions—the driver detects this divergence and resets both watermarks to 0, forcing a full resync to prevent data loss.
Incremental Push: Sending Local Changes
The push process (DO → Remote) operates symmetrically using the pushRev watermark to avoid retransmitting historical data:
- Read high-water mark:
readWatermark(db, "pushRev", backend)retrieves the last confirmed revision. - Coalesce mutations:
coalesceChanges(db, lastPushed + 1)gathers all local changes with revision numbers exceeding the watermark, compacting them into an efficient delta package. - Transmit delta:
remote.pushChangessends the co batch along with thebaseRevfor server-side validation. - Commit progress: Only after receiving acknowledgment does the driver call
writeWatermark(db, "pushRev", newRev, backend)to advance the persistent cursor.
The driver also validates the remote's appliedPushCursor—the remote's view of how much data it has processed. If this cursor diverges from the DO's local expectation, the watermarks reset to ensure consistency.
Resilience and Fine-Grained Resumption
Watermarks provide more than simple revision counters; they enable sub-revision resumption through the path component of the fetch cursor. When a single revision spans thousands of files, the driver can pause mid-revision, persist the specific file path last processed, and resume from that exact location without re-requesting earlier files in the same revision.
This mechanism survives process restarts because watermarks are committed to the same SQLite transaction as the data changes. After a crash, the driver simply reads _vfs_watermark and _vfs_fetch_cursor tables to reconstruct the exact sync state, eliminating the need for expensive full synchronization scans.
Code Implementation Examples
The following examples demonstrate the core watermark interactions used by the sync driver:
// Incremental Pull Implementation
import { readFetchCursor, writeFetchCursor } from "@cloudflare/dofs";
import type { Database } from "@cloudflare/dofs";
import type { SyncRPC } from "@cloudflare/computer-rpc";
async function incrementalPull(
db: Database,
remote: SyncRPC,
backend = "default"
) {
// Retrieve last processed position
const after = readFetchCursor(db, backend);
// Fetch only changes after our cursor
const { currentCursor } = await remote.fetchChanges({ after });
// Apply changes (simplified; actual driver batches internally)
// ...
// Persist new position for next iteration
writeFetchCursor(db, currentCursor, backend);
}
// Incremental Push Implementation
import { readWatermark, writeWatermark, coalesceChanges } from "@cloudflare/dofs";
async function incrementalPush(
db: Database,
remote: SyncRPC,
backend = "default"
) {
// Find last confirmed revision
const lastPushed = readWatermark(db, "pushRev", backend);
// Collect only new mutations
const { changes, newRev } = coalesceChanges(db, lastPushed + 1);
// Send delta
await remote.pushChanges({ changes, baseRev: lastPushed });
// Update watermark only after confirmation
writeWatermark(db, "pushRev", newRev, backend);
}
Summary
- Watermarks are persistent cursors stored in
_vfs_watermarkand_vfs_fetch_cursorSQLite tables that record the sync position for each backend. - Incremental pull uses
readFetchCursorandwriteFetchCursorto resume fetching remote changes from the exact revision and path where the last pull ended. - Incremental push relies on
pushRevwatermarks to transmit only local mutations occurring after the last confirmed push, usingcoalesceChangesto gather deltas. - Divergence detection compares local and remote cursor states; mismatches trigger automatic watermark reset to
0, ensuring data consistency through full resynchronization when necessary. - Batch boundedness combines with watermarks to limit memory usage (
PULL_BATCH_SIZE = 256) while guaranteeing exactly-once processing semantics across crashes.
Frequently Asked Questions
What happens if the remote backend loses its watermark state?
If the remote backend's cursor diverges from the DO's local watermark—such as when the remote "forgets" previously processed revisions—the driver detects this mismatch in packages/rpc/src/sync-driver.ts. Upon detection, the driver resets both pushRev and fetch cursors to 0, triggering a full resynchronization that guarantees no data loss while restoring consistency between the peers.
How does Cloudflare Computer handle revisions too large to fit in memory?
The sync driver processes large revisions using batching with a fixed PULL_BATCH_SIZE of 256 entries. When a revision exceeds this size, the driver stores the specific file path reached within that revision into the fetch cursor via writeFetchCursor. This allows the system to resume mid-revision after a restart or network timeout without reprocessing already-received entries, effectively paginating through massive change sets without unbounded memory growth.
Can multiple remote backends synchronize with the same Durable Object?
Yes. All watermark read and write operations accept a backend parameter that scopes the cursor values to specific remote endpoints. The schema in packages/dofs/src/sync/watermarks.ts maintains separate entries for each backend identifier, enabling simultaneous incremental sync with multiple remotes (such as edge nodes and central servers) while keeping their synchronization states isolated.
Where are watermarks physically stored in the database?
Watermarks persist in two internal SQLite tables created by the DO virtual file system: _vfs_watermark stores the pushRev integer and fetch revision number, while _vfs_fetch_cursor stores the optional path component for partial revision resumption. These tables are manipulated through the helper functions defined in packages/dofs/src/sync/watermarks.ts and are committed atomically with data changes to ensure crash consistency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →