How the Content-Addressed Blob Cache Reduces Repeated SQLite Reads in Cloudflare Computer

The content-addressed blob cache eliminates redundant SQLite queries by storing blob bytes in an in-memory Map keyed by cryptographic hash, serving subsequent reads directly from memory while only querying the database on cache misses or after explicit invalidation.

The Cloudflare Computer project implements a virtual file system that treats file contents as immutable, content-addressed objects stored within SQLite. To minimize expensive disk reads when accessing frequently used blobs, the runtime maintains an in-process cache layer that intercepts lookups before they reach the database, dramatically reducing I/O contention for read-heavy workloads.

What Is the Content-Addressed Blob Cache?

At the storage layer, Computer persists blobs across two SQLite tables defined in packages/dofs/src/schema/core.ts:

Because the hash itself serves as the primary key, identical content automatically deduplicates to a single row. The blob cache is an in-memory JavaScript Map that mirrors these rows during the process lifetime, mapping each hash to its corresponding Uint8Array of bytes.

How getBlobBytes Minimizes Database Access

The core caching logic lives in packages/dofs/src/fs/blobCache.ts. The getBlobBytes(db, hash) function implements a read-through cache pattern:

  1. Memory lookup – It first checks the internal Map for the requested hash. If present, it returns the cached Uint8Array immediately without touching SQLite.
  2. Database fallback – On cache miss, it executes SELECT bytes FROM vfs_blob_bytes WHERE hash = ?, stores the resulting buffer in the Map, and returns the data (lines 20‑38)【https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/blobCache.ts#L20-L38】.

After the first successful retrieval for any given hash, subsequent reads for that content—including reads from different file paths referencing the same blob—are served entirely from heap memory, bypassing SQLite’s B‑tree and page cache entirely.

Cache Invalidation During Writes

To maintain consistency when content changes, the system explicitly invalidates the cache whenever it mutates blob data. The clearBlobCache(db) function (also in blobCache.ts) wipes the in-memory Map, ensuring that no stale bytes survive a write operation.

High-level write paths such as stageBlob in packages/dofs/src/sync/blobs.ts invoke clearBlobCache before inserting or updating records (lines 1‑4)【https://github.com/cloudflare/computer/blob/main/packages/dofs/src/sync/blobs.ts#L1-L4】. This guarantees that any subsequent read observes the freshly written bytes rather than a cached version from the previous state.

Performance Benefits and Deduplication

Because the cache key is the cryptographic hash of the content itself, the system achieves perfect deduplication. Two files with identical contents share the same hash and therefore hit the same cache entry. Tests in packages/dofs/src/sync/blobs.test.ts verify this behavior, demonstrating that writing a file whose content already exists reuses the existing blob row and cache entry without issuing redundant INSERT or SELECT statements (lines 84‑95)【https://github.com/cloudflare/computer/blob/main/packages/dofs/src/sync/blobs.test.ts#L84-L95】.

This architecture is particularly effective for workloads containing many small duplicate files—such as node modules or build artifacts—where the cache hit rate approaches 100% after the initial warm-up phase.

Implementation Example

The following patterns illustrate how the cache integrates into read and write operations:

Reading a file returns cached bytes when available:

// packages/dofs/src/provider.ts (conceptual)
import { getBlobBytes } from "./fs/blobCache.js";

async function readFileContent(
  db: Database,
  hash: Uint8Array
): Promise<Uint8Array> {
  // First call may query SQLite; subsequent calls hit the Map.
  const bytes = await getBlobBytes(db, hash);
  if (!bytes) {
    throw new Error(`missing blob bytes for ${hash}`);
  }
  return bytes;
}

Writing a blob clears the cache to ensure consistency:

// packages/dofs/src/sync/blobs.ts
import { clearBlobCache } from "../fs/blobCache.js";

export async function stageBlob(
  db: Database,
  hash: Uint8Array,
  data: Uint8Array
): Promise<void> {
  // Invalidate before mutation.
  clearBlobCache(db);

  db.run(
    `INSERT INTO vfs_blobs (hash, size, last_seen)
     VALUES (?, ?, ?)
     ON CONFLICT(hash) DO UPDATE SET
       size = excluded.size,
       last_seen = excluded.last_seen`,
    hash,
    data.length,
    Date.now()
  );

  db.run(
    `INSERT INTO vfs_blob_bytes (hash, bytes) VALUES (?, ?)
     ON CONFLICT(hash) DO UPDATE SET bytes = excluded.bytes`,
    hash,
    data
  );
}

The test suite verifies cache reuse and invalidation:

// packages/dofs/src/fs/blobCache.test.ts (conceptual)
import { getBlobBytes, clearBlobCache } from "./blobCache.js";

test("subsequent reads hit the cache, not SQLite", async () => {
  const db = await makeTestDb();
  const hash = await writeRandomBlob(db, new Uint8Array([1, 2, 3]));

  // First read populates the cache via SQLite.
  const first = await getBlobBytes(db, hash);
  expect(first).toEqual(new Uint8Array([1, 2, 3]));

  // Second read returns the identical object from the Map.
  const second = await getBlobBytes(db, hash);
  expect(second).toBe(first); // Same reference => cache hit.

  // After clearing, the next read fetches from SQLite again.
  clearBlobCache(db);
  const third = await getBlobBytes(db, hash);
  expect(third).not.toBe(first);
});

Summary

  • The content-addressed blob cache stores bytes in an in-process Map keyed by cryptographic hash, eliminating repeated SELECT queries against vfs_blob_bytes.
  • getBlobBytes checks memory first and only falls back to SQLite on cache misses, as implemented in packages/dofs/src/fs/blobCache.ts.
  • clearBlobCache invalidates the entire cache before any write operation to guarantee read-after-write consistency.
  • Content addressing provides inherent deduplication, allowing identical files to share both database rows and cached memory without additional overhead.

Frequently Asked Questions

How does getBlobBytes decide whether to query SQLite or use the cache?

The function checks an internal Map for the requested hash. If the hash exists as a key, it returns the associated Uint8Array immediately. Only when the hash is absent does it execute the SQL SELECT bytes FROM vfs_blob_bytes WHERE hash = ?, storing the result in the Map before returning it to the caller.

When is the blob cache cleared to prevent stale data?

The cache is cleared via clearBlobCache whenever the system stages or writes new blob data. Specifically, functions in packages/dofs/src/sync/blobs.ts call this routine before performing INSERT or UPDATE operations on the vfs_blob_bytes table, ensuring that subsequent reads fetch the latest bytes from disk rather than an outdated cached version.

Why does content addressing improve cache efficiency compared to path-based caching?

Because the cache key is the hash of the content itself rather than a file path, any number of files or references pointing to identical data automatically share the same cache entry. This eliminates redundant storage and lookup operations for duplicate content, whereas a path-based scheme would cache the same bytes multiple times under different keys.

What happens if a requested blob hash does not exist in the database?

If getBlobBytes encounters a cache miss and the subsequent SQLite query returns no rows, the function returns undefined (or null depending on the wrapper) without inserting an entry into the cache. This prevents the system from caching negative results, allowing future attempts to succeed if the blob is written later.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →