# How @cloudflare/dofs Implements a Content-Addressed Blob Cache to Optimize Reads

> Explore the @cloudflare/dofs content-addressed blob cache architecture. Discover how its in-process LRU cache optimizes reads by reducing redundant SQLite queries and speeding up access.

- Repository: [Cloudflare/computer](https://github.com/cloudflare/computer)
- Tags: architecture
- Published: 2026-08-14

---

**The @cloudflare/dofs package uses a per-Database, in-process LRU cache that stores up to 16 recently-accessed blobs in memory to eliminate redundant SQLite queries during sequential reads.**

The **content-addressed blob cache** is a critical optimization layer in `@cloudflare/dofs`, Cloudflare's distributed object filesystem. By caching raw `Uint8Array` payloads indexed by cryptographic hash, the system avoids repeated database round-trips when the kernel requests the same file chunks multiple times per read pass.

## Core Architecture

The cache implementation resides in [[`blobCache.ts`](https://github.com/cloudflare/computer/blob/main/blobCache.ts)](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/blobCache.ts). It combines content-addressed storage with a bounded, per-database LRU strategy to balance performance against memory constraints.

### Storage Foundation

File contents are stored in a SQLite table named `vfs_blob_bytes`. Each entry is immutable:

- **32-byte cryptographic hash** serves as the primary key (content address)
- **`INSERT … ON CONFLICT DO NOTHING`** ensures idempotent writes
- Blobs never change after being written

This immutability property makes caching safe—once a blob is cached, its contents remain valid until explicitly invalidated.

### In-Process LRU Cache

To optimize reads, dofs maintains a **WeakMap\<Database, Map\<string, Uint8Array\>>** called `caches`:

| Component | Implementation Detail |
|-----------|----------------------|
| **Per-database isolation** | Each `Database` instance gets its own `Map`, preventing test contamination |
| **Hash encoding** | Blobs are keyed by a **Latin-1 string** (`hashKey`) to avoid hex encoding overhead |
| **Size limit** | **`CHUNK_CACHE_MAX_ENTRIES = 16`** entries (≈ 8 MiB at 512 KiB per blob) |
| **Eviction policy** | LRU via `Map` insertion order; hits trigger remove-and-reinsert to make entries most-recent |

### Cache Operations

**On hit:** Return cached `Uint8Array` directly—no SQLite query.

**On miss:** Execute `SELECT bytes FROM vfs_blob_bytes WHERE hash = ?`, cache result, return payload.

**On blob rewrite:** `clearBlobCache(db)` drops the entire cache for that database instance. This is called from `stageBlob` in the sync receiver to prevent serving stale data.

## Why This Optimizes Reads

The cache design targets specific access patterns in FUSE filesystem operations:

- **Chunked reads at 512 KiB boundaries** — The kernel may request identical chunks up to four times per read pass. Without caching, each request becomes a separate SQL query.

- **Repeated blob patterns** — Large files filled with identical content (e.g., zero-filled sparse files) would otherwise trigger thousands of identical queries. The LRU stores the blob once and serves it from memory.

- **Bounded memory growth** — The 16-entry cap keeps per-database memory usage modest while covering the "hot" working set of sequential reads.

## Using the Cache API

The public surface consists of two functions from [`blobCache.ts`](https://github.com/cloudflare/computer/blob/main/blobCache.ts):

```typescript
import { getBlobBytes, clearBlobCache } from "@cloudflare/dofs/src/fs/blobCache.js";
import { Database } from "@cloudflare/dofs/src/storage.js";

// Fast path: checks in-process LRU before querying SQLite
async function readChunk(db: Database, hash: Uint8Array): Promise<Uint8Array> {
  const bytes = getBlobBytes(db, hash);  // Uint8Array | undefined
  
  if (bytes === undefined) {
    throw new Error("Blob not found in database");
  }
  
  // Return view directly; callers must not mutate
  return bytes;
}

// Invalidate after repair operations that rewrite blobs
function handleRepairCompletion(db: Database): void {
  clearBlobCache(db);  // Guarantees fresh reads from SQLite
}

```

Note that `getBlobBytes` returns a `Uint8Array` view of cached memory. **Callers must treat this as read-only**—mutations would corrupt the cached value for subsequent reads.

## Key Source Files

| File | Purpose |
|------|---------|
| [[`src/fs/blobCache.ts`](https://github.com/cloudflare/computer/blob/main/src/fs/blobCache.ts)](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/blobCache.ts) | LRU implementation, `getBlobBytes`, `clearBlobCache` |
| [[`src/provider.ts`](https://github.com/cloudflare/computer/blob/main/src/provider.ts)](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/provider.ts) | Higher-level provider calling `getBlobBytes` during chunk reads |
| [[`src/fs/blobCache.test.ts`](https://github.com/cloudflare/computer/blob/main/src/fs/blobCache.test.ts)](https://github.com/cloudflare/computer/blob/main/packages/dofs/src/fs/blobCache.test.ts) | Unit tests for LRU behavior, hits, and eviction |

## Summary

- **Content-addressed design**: 32-byte hashes in `vfs_blob_bytes` provide immutable, deduplicated blob storage
- **Per-database LRU**: `WeakMap`-isolated caches prevent cross-contamination with a 16-entry limit
- **Latin-1 key encoding**: Avoids hex overhead while preserving exact hash bytes
- **Explicit invalidation**: `clearBlobCache` ensures consistency after blob rewrites
- **FUSE-optimized**: Eliminates redundant queries for repeatedly-requested 512 KiB chunks

## Frequently Asked Questions

### How does the cache handle concurrent access from multiple readers?

The cache is **synchronous and single-threaded** by design. Since JavaScript's execution model is single-threaded and the cache stores `Uint8Array` references (not copies), concurrent reads of the same cached blob return shared views of the same underlying memory buffer.

### What happens when the cache exceeds 16 entries?

The **oldest entry is evicted** via `cache.keys().next().value`. This uses JavaScript's guaranteed `Map` iteration order—entries are iterated in insertion order, so the first key is the least-recently inserted. On cache hits, entries are removed and re-inserted to make them most-recent.

### Why Latin-1 encoding for the hash key instead of hex or base64?

**Performance and correctness.** Hex encoding doubles key length (64 characters). Base64 introduces padding and character set complexity. Latin-1 (`String.fromCharCode(...hash)`) preserves all 32 bytes exactly as a 32-character string with minimal overhead, since `Uint8Array` values map 1:1 to Latin-1 code points 0-255.

### When should `clearBlobCache` be called manually?

Only when **repair operations rewrite existing blobs**. The normal write path (`INSERT … ON CONFLICT DO NOTHING`) never overwrites, so cache invalidation is unnecessary. The sync receiver's `stageBlob` calls `clearBlobCache` after repairs to guarantee readers see the new blob content rather than stale cached data.