# How the Content-Addressed Blob Cache Reduces Repeated SQLite Reads in Cloudflare Computer

> Discover how Cloudflare's content-addressed blob cache minimizes repeated SQLite reads by storing data in memory keyed by hash, boosting performance and efficiency.

- Repository: [Cloudflare/computer](https://github.com/cloudflare/computer)
- Tags: internals
- Published: 2026-08-16

---

**The content-addressed blob cache eliminates redundant SQLite queries by storing blob bytes in an in-memory Map keyed by cryptographic hash, serving subsequent reads directly from memory while only querying the database on cache misses or after explicit invalidation.**

The Cloudflare Computer project implements a virtual file system that treats file contents as immutable, **content-addressed** objects stored within SQLite. To minimize expensive disk reads when accessing frequently used blobs, the runtime maintains an in-process cache layer that intercepts lookups before they reach the database, dramatically reducing I/O contention for read-heavy workloads.

## What Is the Content-Addressed Blob Cache?

At the storage layer, Computer persists blobs across two SQLite tables defined in [`packages/dofs/src/schema/core.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/schema/core.ts):

- **`vfs_blobs`** – Stores metadata including the cryptographic hash, byte size, and last-seen timestamp.
- **`vfs_blob_bytes`** – Stores the actual binary payload indexed strictly by the content hash (see the schema definition lines 58‑66)【https://github.com/cloudflare/computer/blob/main/packages/dofs/src/schema/core.ts#L58-L66】.

Because the hash itself serves as the primary key, identical content automatically deduplicates to a single row. The **blob cache** is an in-memory JavaScript `Map` that mirrors these rows during the process lifetime, mapping each hash to its corresponding `Uint8Array` of bytes.

## How `getBlobBytes` Minimizes Database Access

The core caching logic lives in [`packages/dofs/src/fs/blobCache.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/blobCache.ts). The `getBlobBytes(db, hash)` function implements a read-through cache pattern:

1. **Memory lookup** – It first checks the internal `Map` for the requested hash. If present, it returns the cached `Uint8Array` immediately without touching SQLite.
2. **Database fallback** – On cache miss, it executes `SELECT bytes FROM vfs_blob_bytes WHERE hash = ?`, stores the resulting buffer in the Map, and returns the data (lines 20‑38)【https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/blobCache.ts#L20-L38】.

After the first successful retrieval for any given hash, subsequent reads for that content—including reads from different file paths referencing the same blob—are served entirely from heap memory, bypassing SQLite’s B‑tree and page cache entirely.

## Cache Invalidation During Writes

To maintain consistency when content changes, the system explicitly invalidates the cache whenever it mutates blob data. The `clearBlobCache(db)` function (also in [`blobCache.ts`](https://github.com/cloudflare/computer/blob/main/blobCache.ts)) wipes the in-memory Map, ensuring that no stale bytes survive a write operation.

High-level write paths such as `stageBlob` in [`packages/dofs/src/sync/blobs.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/sync/blobs.ts) invoke `clearBlobCache` before inserting or updating records (lines 1‑4)【https://github.com/cloudflare/computer/blob/main/packages/dofs/src/sync/blobs.ts#L1-L4】. This guarantees that any subsequent read observes the freshly written bytes rather than a cached version from the previous state.

## Performance Benefits and Deduplication

Because the cache key is the cryptographic hash of the content itself, the system achieves **perfect deduplication**. Two files with identical contents share the same hash and therefore hit the same cache entry. Tests in [`packages/dofs/src/sync/blobs.test.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/sync/blobs.test.ts) verify this behavior, demonstrating that writing a file whose content already exists reuses the existing blob row and cache entry without issuing redundant `INSERT` or `SELECT` statements (lines 84‑95)【https://github.com/cloudflare/computer/blob/main/packages/dofs/src/sync/blobs.test.ts#L84-L95】.

This architecture is particularly effective for workloads containing many small duplicate files—such as node modules or build artifacts—where the cache hit rate approaches 100% after the initial warm-up phase.

## Implementation Example

The following patterns illustrate how the cache integrates into read and write operations:

Reading a file returns cached bytes when available:

```typescript
// packages/dofs/src/provider.ts (conceptual)
import { getBlobBytes } from "./fs/blobCache.js";

async function readFileContent(
  db: Database,
  hash: Uint8Array
): Promise<Uint8Array> {
  // First call may query SQLite; subsequent calls hit the Map.
  const bytes = await getBlobBytes(db, hash);
  if (!bytes) {
    throw new Error(`missing blob bytes for ${hash}`);
  }
  return bytes;
}

```

Writing a blob clears the cache to ensure consistency:

```typescript
// packages/dofs/src/sync/blobs.ts
import { clearBlobCache } from "../fs/blobCache.js";

export async function stageBlob(
  db: Database,
  hash: Uint8Array,
  data: Uint8Array
): Promise<void> {
  // Invalidate before mutation.
  clearBlobCache(db);

  db.run(
    `INSERT INTO vfs_blobs (hash, size, last_seen)
     VALUES (?, ?, ?)
     ON CONFLICT(hash) DO UPDATE SET
       size = excluded.size,
       last_seen = excluded.last_seen`,
    hash,
    data.length,
    Date.now()
  );

  db.run(
    `INSERT INTO vfs_blob_bytes (hash, bytes) VALUES (?, ?)
     ON CONFLICT(hash) DO UPDATE SET bytes = excluded.bytes`,
    hash,
    data
  );
}

```

The test suite verifies cache reuse and invalidation:

```typescript
// packages/dofs/src/fs/blobCache.test.ts (conceptual)
import { getBlobBytes, clearBlobCache } from "./blobCache.js";

test("subsequent reads hit the cache, not SQLite", async () => {
  const db = await makeTestDb();
  const hash = await writeRandomBlob(db, new Uint8Array([1, 2, 3]));

  // First read populates the cache via SQLite.
  const first = await getBlobBytes(db, hash);
  expect(first).toEqual(new Uint8Array([1, 2, 3]));

  // Second read returns the identical object from the Map.
  const second = await getBlobBytes(db, hash);
  expect(second).toBe(first); // Same reference => cache hit.

  // After clearing, the next read fetches from SQLite again.
  clearBlobCache(db);
  const third = await getBlobBytes(db, hash);
  expect(third).not.toBe(first);
});

```

## Summary

- The **content-addressed blob cache** stores bytes in an in-process `Map` keyed by cryptographic hash, eliminating repeated `SELECT` queries against `vfs_blob_bytes`.
- **`getBlobBytes`** checks memory first and only falls back to SQLite on cache misses, as implemented in [`packages/dofs/src/fs/blobCache.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/blobCache.ts).
- **`clearBlobCache`** invalidates the entire cache before any write operation to guarantee read-after-write consistency.
- Content addressing provides inherent deduplication, allowing identical files to share both database rows and cached memory without additional overhead.

## Frequently Asked Questions

### How does `getBlobBytes` decide whether to query SQLite or use the cache?

The function checks an internal `Map` for the requested hash. If the hash exists as a key, it returns the associated `Uint8Array` immediately. Only when the hash is absent does it execute the SQL `SELECT bytes FROM vfs_blob_bytes WHERE hash = ?`, storing the result in the Map before returning it to the caller.

### When is the blob cache cleared to prevent stale data?

The cache is cleared via `clearBlobCache` whenever the system stages or writes new blob data. Specifically, functions in [`packages/dofs/src/sync/blobs.ts`](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/sync/blobs.ts) call this routine before performing `INSERT` or `UPDATE` operations on the `vfs_blob_bytes` table, ensuring that subsequent reads fetch the latest bytes from disk rather than an outdated cached version.

### Why does content addressing improve cache efficiency compared to path-based caching?

Because the cache key is the hash of the content itself rather than a file path, any number of files or references pointing to identical data automatically share the same cache entry. This eliminates redundant storage and lookup operations for duplicate content, whereas a path-based scheme would cache the same bytes multiple times under different keys.

### What happens if a requested blob hash does not exist in the database?

If `getBlobBytes` encounters a cache miss and the subsequent SQLite query returns no rows, the function returns `undefined` (or `null` depending on the wrapper) without inserting an entry into the cache. This prevents the system from caching negative results, allowing future attempts to succeed if the blob is written later.