Content‑Addressed Blobs in Cloudflare Computer: How SHA‑256 Deduplication Works
Content‑addressed blobs in Cloudflare Computer are SHA‑256 hashed data chunks stored in SQLite tables, enabling automatic deduplication where identical file segments are stored only once regardless of how many files reference them.
Cloudflare Computer uses a content‑addressed blob store embedded within Durable Objects to minimize storage overhead and sync bandwidth. By splitting file metadata from raw bytes and indexing data by cryptographic hash, the system eliminates redundancy at both the chunk and manifest levels.
What Are Content‑Addressed Blobs?
In Cloudflare Computer, content‑addressed blobs are immutable byte sequences identified by the SHA‑256 hash of their contents rather than by location or filename. The architecture stores these blobs inside the Durable Object’s SQLite database across four specialized tables defined in packages/dofs/src/schema/core.ts:
vfs_blobs– Stores each blob’s hash, byte size, and a garbage‑collection clock (last_seen)vfs_blob_bytes– Holds the raw blob data, keyed by the same SHA‑256 hashvfs_chunks– Maps file inodes and chunk indices to blob hashesvfs_manifests– Stores ordered lists of blob hashes that compose entire files
This separation of metadata from bytes allows multiple files to reference the same blob without data duplication.
How Deduplication Works in Cloudflare Computer
The deduplication pipeline operates through five distinct stages when writing files:
1. Fixed‑Size Chunking
Files are split into 512 KiB chunks defined by the CHUNK_SIZE constant. This granularity balances storage efficiency with lookup performance.
2. Cryptographic Hashing
Each chunk’s bytes are hashed using SHA‑256. The resulting digest serves as the primary key in vfs_blobs and vfs_blob_bytes.
3. Insert‑or‑Ignore Semantics
When staging a chunk, the system executes an INSERT with ON CONFLICT DO NOTHING. As implemented in packages/dofs/src/sync/blobs.ts, this guarantees that identical data is stored only once:
// packages/dofs/src/sync/blobs.ts
await db.run(`
INSERT INTO vfs_blob_bytes (hash, bytes) VALUES (?, ?)
ON CONFLICT(hash) DO UPDATE SET bytes = excluded.bytes
`, hash, chunkBytes);
If the hash already exists, the write is skipped, providing automatic deduplication at the storage layer.
4. Reference Counting via GC Clocks
Each row in vfs_blobs maintains a last_seen timestamp refreshed whenever the blob is referenced. Multiple files can point to the same hash through vfs_chunks rows, and unused blobs are reclaimed when their clock expires during garbage collection.
5. Manifest‑Level Deduplication
Files with identical content generate identical ordered lists of chunk hashes. These manifests are stored in vfs_manifests and themselves content‑addressed, allowing entire files to be deduplicated without comparing bytes.
The Write Path: Staging and Persistence
When writeFile.ts processes a file, it streams data into 512 KiB buffers, computes SHA‑256 hashes, and stages chunks into the blob store. The implementation in packages/dofs/src/fs/writeFile.ts (lines 322–526) creates the vfs_chunks mapping rows only after confirming the blob exists in vfs_blob_bytes.
Reading and Caching Content‑Addressed Blobs
On read operations, packages/dofs/src/fs/blobCache.ts implements lazy loading. It queries vfs_blob_bytes by hash and caches the result in memory for subsequent accesses:
// packages/dofs/src/fs/blobCache.ts
const bytes = await db.get<ArrayBuffer>(
"SELECT bytes FROM vfs_blob_bytes WHERE hash = ?",
hash
);
This design ensures that frequently accessed chunks are served from memory while maintaining a single source of truth in SQLite.
Sync Protocol and Network Deduplication
The content‑addressed model enables efficient synchronization. The RPC interface defined in packages/rpc/src/interface.ts exposes three key methods:
hasObjects– Returns which hashes the remote Durable Object is missingfetchObjects– Retrieves specific blob bytes by hashpushObjects– Transmits missing blobs to the remote side
When syncing, the system calls hasObjects with the list of chunk hashes from vfs_manifests. Only hashes unknown to the remote (triggering EUNKNOWN_HASH errors) are transferred, reducing network traffic to unique content only.
The push implementation in packages/dofs/src/sync/push.ts retrieves missing bytes via:
// packages/dofs/src/sync/push.ts
const rows = await db.all<ArrayBuffer[]>(
"SELECT bytes FROM vfs_blob_bytes WHERE hash = ?",
hash
);
Summary
- Content‑addressed blobs use SHA‑256 hashes as primary keys, ensuring identical chunks map to the same database row.
- The
vfs_blob_bytestable enforces deduplication throughON CONFLICT DO NOTHINGsemantics during writes. - Reference counting via
last_seentimestamps allows safe garbage collection of unreferenced blobs. - The sync protocol transfers only missing hashes, eliminating redundant network traffic during replication.
- Manifest deduplication extends content addressing to entire files, not just individual chunks.
Frequently Asked Questions
What hash algorithm does Cloudflare Computer use for content addressing?
Cloudflare Computer uses SHA‑256 to generate content addresses. Every 512 KiB chunk is hashed, and the resulting 256‑bit digest serves as the primary key in vfs_blobs and vfs_blob_bytes.
How does Cloudflare Computer handle duplicate chunks during file writes?
When writing files, the system attempts to insert each chunk’s bytes into vfs_blob_bytes with an ON CONFLICT DO NOTHING clause. If the SHA‑256 hash already exists, the insert is skipped, and the existing blob is referenced instead of storing duplicate data.
Can identical files across different directories share the same storage?
Yes. Because chunks are content‑addressed, files with overlapping content—whether in the same directory or different paths—reference the same rows in vfs_blob_bytes. Additionally, identical entire files generate the same manifest hash in vfs_manifests, enabling complete file‑level deduplication.
What happens to blobs when files are deleted?
Deleted files remove their entries from vfs_chunks and vfs_manifests, but the underlying blobs in vfs_blob_bytes persist temporarily. The last_seen timestamp in vfs_blobs tracks usage; blobs unreferenced for extended periods are reclaimed by the garbage collection process based on this clock.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →