How Cloudflare Computer Handles Large Files and Deduplicates Data Using Chunking

Cloudflare Computer stores every file as a series of fixed-size 512 KiB chunks, using SHA-256 content addressing to deduplicate identical data across files and minimize sync traffic by transferring only missing chunks.

The cloudflare/computer repository implements a distributed object filesystem that treats large files as collections of content-addressed chunks. By splitting files into 512 KiB windows and storing each chunk independently, the system achieves efficient storage deduplication and optimized network synchronization. This architecture ensures that identical data blocks shared between files or file versions are stored only once, while sync operations transmit strictly necessary data.

The 512 KiB Chunking Architecture

Cloudflare Computer processes every file write by dividing the byte stream into fixed 512 KiB windows. In packages/dofs/src/fs/writeFile.ts, the constant CHUNK_SIZE = 512 * 1024 (lines 27-30) defines this boundary, ensuring consistent chunk boundaries regardless of file size.

Fixed-Size Windowing and SHA-256 Hashing

When a file is written, the chunksOf() function (lines 98-107 in writeFile.ts) iterates through the byte array and computes a SHA-256 hash for each 512 KiB window. This produces a content-addressed chunk hash that uniquely identifies the data content. Each chunk results in a {hash, bytes, size} object, where the hash serves as the primary key for storage.

The raw chunk bytes are stored in vfs_blob_bytes, while their metadata (hash-size pairs) are recorded in vfs_chunks. Because storage is content-addressed, two identical 512 KiB blocks from different files generate identical hashes and point to the same underlying blob row.

Content-Addressed Storage and Manifests

Rather than storing files as monolithic objects, Cloudflare Computer constructs manifests that represent an ordered list of chunk hashes. The buildManifest() function in packages/dofs/src/sync/manifests.ts (lines 5-13) serializes the chunk list and computes a manifest hash that uniquely identifies the exact sequence of chunks.

Staging Blobs Without Loading Files Into Memory

The implementation never loads entire large files into RAM during the write path. The writeFileStreaming() function (lines 26-34 in writeFile.ts) processes ReadableStream inputs, windows them into 512 KiB pieces, and immediately stages each chunk via stageBlob() (lines 60-63).

After staging, linkStagedChunksSync() (lines 40-48) performs an atomic transaction that inserts the chunk references into vfs_chunks and updates the inode's manifest hash. This O(number of chunks) operation ensures that even multi-gigabyte files are handled with constant memory overhead.

Sync-Time Deduplication

During synchronization, the driver minimizes network traffic by determining which chunks already exist on the remote peer before transmitting data. This deduplication on sync ensures that only novel content crosses the network.

Remote Probing with hasObjects

Before pushing data, the sync driver calls hasObjects() (lines 59-62 in packages/rpc/src/sync-driver.ts) to query the remote side's existing chunk hashes. The driver transmits only the list of chunk hashes, not the data itself, allowing the remote peer to respond with a bitmask of which chunks it already possesses.

Selective Transfer of Missing Chunks

For push operations, the driver (lines 14-22 in sync-driver.ts) streams only chunks marked as missing via pushObjects. Conversely, during pull operations (lines 191-199), the driver gathers required hashes and invokes fetchObjects() exclusively for chunks absent locally. Because manifests are transferred as hash lists, the receiver can reconstruct files identical to the sender's without redundant data transfer.

Practical Implementation

The following examples demonstrate the chunking and deduplication workflow using the Cloudflare Computer API.

To split a buffer into content-addressed chunks:

import { chunksOf } from "@cloudflare/dofs/fs/writeFile";

const bytes = await Deno.readFile("large.bin");
const chunkInfos = chunksOf(bytes);   // [{hash, bytes, size}, …]

To persist chunks and link them to a file path:

import { linkStagedChunksSync } from "@cloudflare/dofs/fs/writeFile";

const db = await openDatabase();               // @cloudflare/dofs Database
const parts = ["data", "large.bin"];           // canonicalized path parts
const chunkRefs = chunkInfos.map(c => ({hash: c.hash, size: c.size}));

linkStagedChunksSync(
  db,
  "/data/large.bin",   // canonical path
  parts,
  chunkRefs,
  { mode: 0o644 },     // write options
  Date.now(),
);

To synchronize with remote deduplication:

import { pushOnce, pullOnce } from "@cloudflare/rpc/sync-driver";

await pushOnce(db, remote);   // remote.hasObjects() filters out existing chunks
await pullOnce(db, remote);   // remote.fetchObjects() brings in only unknown chunks

Summary

  • Cloudflare Computer splits every file into 512 KiB chunks addressed by SHA-256 hashes, defined in packages/dofs/src/fs/writeFile.ts.
  • Content-addressed storage in vfs_blob_bytes ensures identical chunks across different files share a single storage row.
  • Manifests stored in vfs_manifests represent ordered chunk lists; identical files share manifest hashes.
  • The sync protocol probes remote chunk possession via hasObjects() and transfers only missing chunks through pushObjects or fetchObjects.
  • Streaming write operations process files in constant memory, staging chunks via stageBlob() without loading entire files into RAM.

Frequently Asked Questions

What is the chunk size used in Cloudflare Computer?

Cloudflare Computer uses a fixed chunk size of 512 KiB (524,288 bytes). This constant is defined as CHUNK_SIZE = 512 * 1024 in packages/dofs/src/fs/writeFile.ts (lines 27-30) and is applied consistently during file writes, chunk hashing, and sync operations.

How does content addressing enable deduplication?

Content addressing uses SHA-256 hashes to identify chunks. When two files contain identical 512 KiB blocks, they produce identical hashes. The system stores these chunks in vfs_blob_bytes keyed by hash, so duplicate data references the same underlying storage row. This applies across different files and different versions of the same file.

How does the sync protocol minimize bandwidth?

The sync protocol minimizes bandwidth through selective chunk transfer. Before sending data, the driver queries the remote peer using hasObjects() to identify which chunks already exist. Only chunks absent from the remote are transmitted via pushObjects or fetchObjects. Because manifest transfers use hash lists rather than full content, sync traffic is proportional to new data rather than total file size.

Can Cloudflare Computer handle files larger than available memory?

Yes. The streaming write path in writeFileStreaming() (lines 26-34) processes files as ReadableStream inputs without loading them entirely into memory. Chunks are staged immediately via stageBlob() and linked atomically through linkStagedChunksSync(), maintaining constant memory overhead regardless of file size.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →