# How Cloudflare Computer Handles Large Files and Deduplicates Data Using Chunking

> Discover how Cloudflare Computer handles large files and deduplicates data with fixed-size chunks and SHA-256 content addressing, minimizing sync traffic by transferring only missing chunks.

- Repository: [Cloudflare/computer](https://github.com/cloudflare/computer)
- Tags: internals
- Published: 2026-08-13

---

**Cloudflare Computer stores every file as a series of fixed-size 512 KiB chunks, using SHA-256 content addressing to deduplicate identical data across files and minimize sync traffic by transferring only missing chunks.**

The [cloudflare/computer](https://github.com/cloudflare/computer) repository implements a distributed object filesystem that treats large files as collections of content-addressed chunks. By splitting files into 512 KiB windows and storing each chunk independently, the system achieves efficient storage deduplication and optimized network synchronization. This architecture ensures that identical data blocks shared between files or file versions are stored only once, while sync operations transmit strictly necessary data.

## The 512 KiB Chunking Architecture

Cloudflare Computer processes every file write by dividing the byte stream into fixed 512 KiB windows. In [`packages/dofs/src/fs/writeFile.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/writeFile.ts), the constant `CHUNK_SIZE = 512 * 1024` (lines 27-30) defines this boundary, ensuring consistent chunk boundaries regardless of file size.

### Fixed-Size Windowing and SHA-256 Hashing

When a file is written, the `chunksOf()` function (lines 98-107 in [`writeFile.ts`](https://github.com/cloudflare/computer/blob/main/writeFile.ts)) iterates through the byte array and computes a SHA-256 hash for each 512 KiB window. This produces a content-addressed **chunk hash** that uniquely identifies the data content. Each chunk results in a `{hash, bytes, size}` object, where the hash serves as the primary key for storage.

The raw chunk bytes are stored in `vfs_blob_bytes`, while their metadata (hash-size pairs) are recorded in `vfs_chunks`. Because storage is content-addressed, two identical 512 KiB blocks from different files generate identical hashes and point to the same underlying blob row.

## Content-Addressed Storage and Manifests

Rather than storing files as monolithic objects, Cloudflare Computer constructs **manifests** that represent an ordered list of chunk hashes. The `buildManifest()` function in [`packages/dofs/src/sync/manifests.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/sync/manifests.ts) (lines 5-13) serializes the chunk list and computes a manifest hash that uniquely identifies the exact sequence of chunks.

### Staging Blobs Without Loading Files Into Memory

The implementation never loads entire large files into RAM during the write path. The `writeFileStreaming()` function (lines 26-34 in [`writeFile.ts`](https://github.com/cloudflare/computer/blob/main/writeFile.ts)) processes `ReadableStream` inputs, windows them into 512 KiB pieces, and immediately stages each chunk via `stageBlob()` (lines 60-63).

After staging, `linkStagedChunksSync()` (lines 40-48) performs an atomic transaction that inserts the chunk references into `vfs_chunks` and updates the inode's manifest hash. This O(number of chunks) operation ensures that even multi-gigabyte files are handled with constant memory overhead.

## Sync-Time Deduplication

During synchronization, the driver minimizes network traffic by determining which chunks already exist on the remote peer before transmitting data. This **deduplication on sync** ensures that only novel content crosses the network.

### Remote Probing with hasObjects

Before pushing data, the sync driver calls `hasObjects()` (lines 59-62 in [`packages/rpc/src/sync-driver.ts`](https://github.com/cloudflare/computer/blob/main/packages/rpc/src/sync-driver.ts)) to query the remote side's existing chunk hashes. The driver transmits only the list of chunk hashes, not the data itself, allowing the remote peer to respond with a bitmask of which chunks it already possesses.

### Selective Transfer of Missing Chunks

For **push** operations, the driver (lines 14-22 in [`sync-driver.ts`](https://github.com/cloudflare/computer/blob/main/sync-driver.ts)) streams only chunks marked as missing via `pushObjects`. Conversely, during **pull** operations (lines 191-199), the driver gathers required hashes and invokes `fetchObjects()` exclusively for chunks absent locally. Because manifests are transferred as hash lists, the receiver can reconstruct files identical to the sender's without redundant data transfer.

## Practical Implementation

The following examples demonstrate the chunking and deduplication workflow using the Cloudflare Computer API.

To split a buffer into content-addressed chunks:

```typescript
import { chunksOf } from "@cloudflare/dofs/fs/writeFile";

const bytes = await Deno.readFile("large.bin");
const chunkInfos = chunksOf(bytes);   // [{hash, bytes, size}, …]

```

To persist chunks and link them to a file path:

```typescript
import { linkStagedChunksSync } from "@cloudflare/dofs/fs/writeFile";

const db = await openDatabase();               // @cloudflare/dofs Database
const parts = ["data", "large.bin"];           // canonicalized path parts
const chunkRefs = chunkInfos.map(c => ({hash: c.hash, size: c.size}));

linkStagedChunksSync(
  db,
  "/data/large.bin",   // canonical path
  parts,
  chunkRefs,
  { mode: 0o644 },     // write options
  Date.now(),
);

```

To synchronize with remote deduplication:

```typescript
import { pushOnce, pullOnce } from "@cloudflare/rpc/sync-driver";

await pushOnce(db, remote);   // remote.hasObjects() filters out existing chunks
await pullOnce(db, remote);   // remote.fetchObjects() brings in only unknown chunks

```

## Summary

- Cloudflare Computer splits every file into **512 KiB chunks** addressed by SHA-256 hashes, defined in [`packages/dofs/src/fs/writeFile.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/writeFile.ts).
- **Content-addressed storage** in `vfs_blob_bytes` ensures identical chunks across different files share a single storage row.
- **Manifests** stored in `vfs_manifests` represent ordered chunk lists; identical files share manifest hashes.
- The **sync protocol** probes remote chunk possession via `hasObjects()` and transfers only missing chunks through `pushObjects` or `fetchObjects`.
- **Streaming write operations** process files in constant memory, staging chunks via `stageBlob()` without loading entire files into RAM.

## Frequently Asked Questions

### What is the chunk size used in Cloudflare Computer?

Cloudflare Computer uses a **fixed chunk size of 512 KiB** (524,288 bytes). This constant is defined as `CHUNK_SIZE = 512 * 1024` in [`packages/dofs/src/fs/writeFile.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/writeFile.ts) (lines 27-30) and is applied consistently during file writes, chunk hashing, and sync operations.

### How does content addressing enable deduplication?

Content addressing uses **SHA-256 hashes** to identify chunks. When two files contain identical 512 KiB blocks, they produce identical hashes. The system stores these chunks in `vfs_blob_bytes` keyed by hash, so duplicate data references the same underlying storage row. This applies across different files and different versions of the same file.

### How does the sync protocol minimize bandwidth?

The sync protocol minimizes bandwidth through **selective chunk transfer**. Before sending data, the driver queries the remote peer using `hasObjects()` to identify which chunks already exist. Only chunks absent from the remote are transmitted via `pushObjects` or `fetchObjects`. Because manifest transfers use hash lists rather than full content, sync traffic is proportional to new data rather than total file size.

### Can Cloudflare Computer handle files larger than available memory?

Yes. The **streaming write path** in `writeFileStreaming()` (lines 26-34) processes files as `ReadableStream` inputs without loading them entirely into memory. Chunks are staged immediately via `stageBlob()` and linked atomically through `linkStagedChunksSync()`, maintaining constant memory overhead regardless of file size.