# Dexie IndexedDB Schema Design in Read Frog: Efficient Caching for Translations and AI Data

> Explore the Dexie IndexedDB schema design in Read Frog for efficient caching. Learn how versioning, primary keys, and type safety speed up translation and AI data lookups.

- Repository: [MengXi/read-frog](https://github.com/mengxi-ream/read-frog)
- Tags: architecture
- Published: 2026-03-07

---

**Read Frog implements a versioned Dexie IndexedDB schema with deterministic primary keys and type-safe EntityTables to cache translations, LLM batch requests, article summaries, and AI segmentation data, enabling O(1) lookups and seamless zero-downtime migrations.**

Read Frog persists all transient AI-driven operations—translations, LLM-generated summaries, batch-request metadata, and video-segmentation results—in a single IndexedDB database powered by **Dexie**. The schema architecture, defined in [`src/utils/db/dexie/app-db.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/app-db.ts), employs incremental versioning to add new cache tables without breaking existing user data, ensuring a robust offline-first experience for the browser extension.

## Core Database Architecture

### The AppDB Class Structure

The database logic centers on an `AppDB` class that extends Dexie's base functionality. Located in [`src/utils/db/dexie/app-db.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/app-db.ts), this class exposes strongly-typed `EntityTable` properties for each cache category:

```typescript
import Dexie, { EntityTable } from "dexie";
import { APP_NAME } from "@/utils/constants/app";
import TranslationCache from "./tables/translation-cache";
import BatchRequestRecord from "./tables/batch-request-record";
import ArticleSummaryCache from "./tables/article-summary-cache";
import AiSegmentationCache from "./tables/ai-segmentation-cache";

export default class AppDB extends Dexie {
  translationCache!: EntityTable<TranslationCache, "key">;
  batchRequestRecord!: EntityTable<BatchRequestRecord, "key">;
  articleSummaryCache!: EntityTable<ArticleSummaryCache, "key">;
  aiSegmentationCache!: EntityTable<AiSegmentationCache, "key">;

  constructor() {
    super(`${upperCamelCase(APP_NAME)}DB`);
    // Version definitions follow...
  }
}

```

* The `EntityTable<TableType, "key">` generic provides **compile-time type safety** for all CRUD operations.
* The database name is dynamically derived from `APP_NAME`, producing **ReadFrogDB** in production environments.

### Singleton Pattern for Global Access

To ensure a single source of truth across background scripts, content scripts, and UI components, the extension instantiates one shared database object in [`src/utils/db/dexie/db.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/db.ts):

```typescript
import AppDB from "./app-db";

export const db = new AppDB();

```

All modules import this `db` singleton to read or write cached data, preventing connection leaks and maintaining transactional consistency throughout the extension lifecycle.

## Versioned Schema Migration Strategy

The Dexie IndexedDB schema evolves through four discrete versions. Each upgrade adds a new table while preserving existing stores, guaranteeing that existing users retain their cached data when the extension updates.

| Version | Table Added | Primary Key | Indexed Fields |
|---------|-------------|-------------|----------------|
| **1** | `translationCache` | `key` | `translation`, `createdAt` |
| **2** | `batchRequestRecord` | `key` | `createdAt`, `originalRequestCount`, `provider`, `model` |
| **3** | `articleSummaryCache` | `key` | `createdAt` |
| **4** | `aiSegmentationCache` | `key` | `createdAt` |

The version definitions in [`src/utils/db/dexie/app-db.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/app-db.ts) explicitly repeat previous stores to maintain backward compatibility:

```typescript
// Version 1 – Initial translation caching
this.version(1).stores({
  translationCache: `key, translation, createdAt`,
});

// Version 2 – Add batch request metadata
this.version(2).stores({
  translationCache: `key, translation, createdAt`,
  batchRequestRecord: `key, createdAt, originalRequestCount, provider, model`,
});

// Version 3 – Add article summary storage
this.version(3).stores({
  translationCache: `key, translation, createdAt`,
  batchRequestRecord: `key, createdAt, originalRequestCount, provider, model`,
  articleSummaryCache: `key, createdAt`,
});

// Version 4 – Add AI video segmentation
this.version(4).stores({
  translationCache: `key, translation, createdAt`,
  batchRequestRecord: `key, createdAt, originalRequestCount, provider, model`,
  articleSummaryCache: `key, createdAt`,
  aiSegmentationCache: `key, createdAt`,
});

```

## Table Schema Design for AI Caching

### TranslationCache Schema

The `translationCache` table stores text-to-text translations with a composite deterministic key derived from source language, target language, and original text. This design eliminates duplicate network calls for identical translation requests.

- **Primary Key**: `key` (string hash)
- **Fields**: `translation` (string), `createdAt` (Date)
- **Source**: [`src/utils/db/dexie/tables/translation-cache.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/tables/translation-cache.ts)

### BatchRequestRecord Schema

To enable request throttling and usage analytics, the `batchRequestRecord` table logs metadata about batched LLM requests. It indexes multiple fields beyond the primary key to support efficient range queries.

- **Primary Key**: `key` (unique request identifier)
- **Indexed Fields**: `createdAt`, `originalRequestCount`, `provider`, `model`
- **Purpose**: Tracks batch size, AI provider, and model version for rate-limiting logic
- **Source**: [`src/utils/db/dexie/tables/batch-request-record.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/tables/batch-request-record.ts)

### ArticleSummaryCache and AiSegmentationCache Schemas

The remaining tables handle specialized AI outputs:

| Table | Content Type | Key Strategy |
|-------|--------------|--------------|
| `articleSummaryCache` | LLM-generated article summaries | SHA-256 hash of article content + provider configuration |
| `aiSegmentationCache` | VTT-formatted video subtitle segmentation | Deterministic hash of video segment identifiers |

Both tables inherit from `Dexie.Entity`, granting them native `put()`, `get()`, `delete()`, and `where()` methods with full TypeScript intellisense.

## Performance Optimization Techniques

### O(1) Lookup with Deterministic Keys

Every table uses a deterministic string `key` as its primary key. By hashing input parameters (text content, language pairs, or configuration objects), the extension generates consistent identifiers. This allows **constant-time retrieval** via `db.table.get(key)` instead of expensive full-text searches.

### Secondary Indexes for Analytics

The `batchRequestRecord` table creates secondary indexes on `createdAt`, `provider`, and `model`. These indexes enable efficient temporal queries for analytics, such as retrieving all requests made within the last 24 hours without scanning the entire object store.

### Type Safety with EntityTable

Using Dexie's `EntityTable` generic ensures that all database operations are type-checked at compile time. Attempting to insert a malformed object or querying a non-existent index results in immediate TypeScript errors, reducing runtime exceptions in the extension's background scripts.

## Practical Implementation Examples

### Caching a Translation

```typescript
import { db } from "@/utils/db/dexie/db";

async function cacheTranslation(
  source: string,
  target: string,
  originalText: string,
  translated: string
) {
  const key = `${source}:${target}:${originalText}`;
  await db.translationCache.put({
    key,
    translation: translated,
    createdAt: new Date(),
  });
}

```

### Retrieving Cached Data

```typescript
async function getCachedTranslation(
  source: string,
  target: string,
  originalText: string
) {
  const key = `${source}:${target}:${originalText}`;
  return await db.translationCache.get(key); // Returns undefined on cache miss
}

```

### Storing AI-Generated Summaries

```typescript
import { db } from "@/utils/db/dexie/db";
import { sha256 } from "@/utils/crypto";

async function cacheArticleSummary(
  text: string, 
  providerConfig: object, 
  summary: string
) {
  const hash = await sha256(text);
  const key = `${hash}:${JSON.stringify(providerConfig)}`;
  await db.articleSummaryCache.put({
    key,
    summary,
    createdAt: new Date(),
  });
}

```

### Querying Recent Batch Requests

```typescript
async function recentBatchRequests(limit = 20) {
  return await db.batchRequestRecord
    .orderBy("createdAt")
    .reverse()
    .limit(limit)
    .toArray();
}

```

## Summary

- **Versioned migrations** in [`src/utils/db/dexie/app-db.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/app-db.ts) allow additive schema changes without data loss across extension updates.
- **Deterministic primary keys** enable O(1) cache lookups and prevent duplicate AI API calls.
- **EntityTable** type definitions provide compile-time safety for all IndexedDB operations.
- **Secondary indexes** on `createdAt`, `provider`, and `model` support efficient analytics queries for batch request monitoring.
- The **singleton pattern** in [`src/utils/db/dexie/db.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/db.ts) ensures consistent database access across all extension contexts.

## Frequently Asked Questions

### Why does Read Frog use Dexie instead of the native IndexedDB API?

Dexie provides a **type-safe, Promise-based abstraction** over the verbose native IndexedDB API. The `EntityTable` interface enables compile-time validation of table schemas, while Dexie's versioning system handles complex schema migrations automatically. This reduces boilerplate code in [`src/utils/db/dexie/app-db.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/app-db.ts) and eliminates common errors like connection leaks or transaction mismanagement.

### How does the versioned schema prevent data loss during extension updates?

Each schema version explicitly lists all existing tables, ensuring that when a user upgrades from version 2 to version 4, their `translationCache` and `batchRequestRecord` data remains intact. Dexie's upgrade mechanism runs these version definitions sequentially, adding new object stores without deleting or modifying existing ones, which preserves cached translations and AI data across releases.

### What makes the cache lookups O(1) constant time?

All tables use a **deterministic string `key`** as their primary key, indexed by IndexedDB's internal B-tree structure. When retrieving data via `db.translationCache.get(key)`, the database performs a direct hash lookup rather than a table scan. Because the key is deterministically generated from input parameters (like `source:target:text`), the extension immediately knows the exact storage location without searching.

### How are deterministic keys generated for AI data like article summaries?

The extension generates keys by hashing content identifiers combined with configuration parameters. For article summaries, [`src/utils/db/dexie/tables/article-summary-cache.ts`](https://github.com/mengxi-ream/read-frog/blob/main/src/utils/db/dexie/tables/article-summary-cache.ts) uses a SHA-256 hash of the article text concatenated with a JSON-serialized provider configuration. This ensures that identical content processed with the same AI model always maps to the same database key, maximizing cache hit rates while avoiding redundant LLM API calls.