WebLLM Model Download and Caching in the Browser: A Complete Guide

WebLLM models are downloaded once per browser and cached in IndexedDB, enabling offline inference after the initial fetch by checking hasModelInCache() before triggering CreateWebWorkerMLCEngine().

The felladrin/minisearch repository demonstrates a production-ready implementation of client-side AI using the @mlc-ai/web-llm package. By handling model download and caching entirely within the browser, the application eliminates server dependencies for inference while optimizing load times through persistent storage.

How WebLLM Model Download and Caching Works

The MiniSearch client orchestrates a five-stage pipeline to manage large language models in the browser environment. When a user selects a model such as "Llama-2-7B-Chat", the system first verifies local availability before initiating any network requests.

The architecture relies on IndexedDB as the underlying storage mechanism. The @mlc-ai/web-llm library writes model weights, tokenizer configurations, and metadata to the browser's persistent storage during the first download. Subsequent sessions reuse these artifacts, reducing initialization time and enabling offline functionality.

Step-by-Step Implementation in MiniSearch

Model Selection and Cache Verification

The process begins in client/modules/textGenerationWithWebLlm.ts where the application retrieves the user's preferred model ID from settings. Before any download initiates, the code queries the cache status using the hasModelInCache() utility provided by the WebLLM library.

import { hasModelInCache } from "@mlc-ai/web-llm";
import { getSettings } from "../stores/settings";

// Lines 72-78: Import statements and configuration
const selectedModelId = getSettings().webLlmModelId;
const isCached = await hasModelInCache(selectedModelId);

Lines 89-96 in the same file determine UI behavior based on this cache check. If isCached returns true, the application skips progress indicators and proceeds directly to engine initialization.

Progress Reporting During Download

When hasModelInCache() returns false, MiniSearch configures an InitProgressCallback to stream download progress to the user interface. This callback updates a reactive store that drives a progress bar component.

// Inside textGenerationWithWebLlm.ts (lines 89-96 context)
const onProgress = (progress: { progress: number }) => {
  updateModelLoadingProgress(Math.round(progress.progress * 100));
};

The progress value represents the fraction of model weights downloaded, ranging from 0.0 to 1.0. The WebLLM library handles the underlying HTTP range requests and writes chunks to IndexedDB atomically.

Engine Initialization with Web Workers

MiniSearch implements a dual-path initialization strategy depending on browser capabilities. The primary path uses a dedicated Web Worker to isolate the ML computation from the main thread.

The worker script resides at client/modules/webLlmWorker.ts and acts as a thin wrapper that forwards messages to the MLC engine handler:

// Creating the worker (lines 109-118 context)
const worker = new Worker(
  new URL("./webLlmWorker.ts", import.meta.url),
  { type: "module" }
);

const engine = await CreateWebWorkerMLCEngine(
  worker,
  selectedModelId,
  { initProgressCallback: onProgress, logLevel: "SILENT" },
  { context_window_size: 2048 }
);

If worker instantiation fails or the environment blocks cross-origin workers, the code falls back to CreateMLCEngine(), which runs the inference engine in the main thread. This fallback maintains functionality at the cost of UI responsiveness during model loading.

Key Source Files and Functions

The implementation spans three critical files within the repository:

  • client/modules/textGenerationWithWebLlm.ts – Contains the orchestration logic at lines 72-78 (imports), 89-96 (cache checking), and 109-118 (engine instantiation).
  • client/modules/webLlmWorker.ts – Minimal worker entry point that imports @mlc-ai/web-llm and delegates to the library's internal worker handler.
  • package.json – Declares the @mlc-ai/web-llm dependency that provides CreateWebWorkerMLCEngine, hasModelInCache, and CreateMLCEngine.

After generation completes, the engine exposes runtimeStatsText(), which reports the actual download size and inference latency. MiniSearch logs these metrics to the console for debugging purposes:

// After generation
console.log(engine.runtimeStatsText());

Summary

  • Cache-first architecture: The hasModelInCache() function checks IndexedDB before any network activity, eliminating redundant downloads.
  • Worker isolation: The primary implementation uses CreateWebWorkerMLCEngine() with a dedicated worker at client/modules/webLlmWorker.ts to prevent UI blocking.
  • Progress transparency: An InitProgressCallback streams download percentage to the user interface when fetching new models.
  • Persistent storage: Model weights reside in the browser's IndexedDB, enabling offline inference after the initial download completes.
  • Graceful degradation: The system falls back to CreateMLCEngine() when Web Workers are unavailable.

Frequently Asked Questions

Where are WebLLM models stored in the browser?

WebLLM models are stored in the browser's IndexedDB database. The @mlc-ai/web-llm library manages this storage automatically, writing model weights, tokenizer files, and configuration metadata during the first download. This persistent cache survives page reloads and browser sessions, allowing subsequent initializations to load models locally without network requests.

How does MiniSearch check if a model is already cached?

MiniSearch uses the hasModelInCache() utility imported from @mlc-ai/web-llm in client/modules/textGenerationWithWebLlm.ts. Before initializing the engine, the code passes the selected model ID (retrieved from getSettings().webLlmModelId) to this function. If it returns true, the application skips the download progress UI and proceeds directly to engine creation using the cached artifacts.

What happens if the Web Worker fails to load?

If the Web Worker instantiation fails or the browser blocks cross-origin worker scripts, MiniSearch falls back to main-thread execution. The code in textGenerationWithWebLlm.ts catches worker creation failures and invokes CreateMLCEngine() instead of CreateWebWorkerMLCEngine(). While this fallback maintains functionality, it runs the inference engine on the main thread, which may cause UI stuttering during model loading and generation.

Can users run WebLLM models offline after the first download?

Yes, once a model is downloaded and cached in IndexedDB, users can run inference entirely offline. The @mlc-ai/web-llm library loads model weights from the browser's persistent storage on subsequent visits, bypassing all network requests. This offline capability persists until the user clears site data or the browser evicts the IndexedDB database due to storage pressure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →