Browser-Based AI Model Loading and WebGPU Integration in MiniSearch

MiniSearch runs large language models directly in the browser by detecting WebGPU capabilities, selecting appropriate quantized models based on shader-f16 support, initializing the WebLLM engine with progress tracking, and streaming completions through a WebWorker-backed architecture.

MiniSearch enables privacy-preserving AI search by running large language models entirely in the browser without server-side dependencies. This article examines how the felladrin/minisearch repository implements browser-based AI model loading and WebGPU integration using the @mlc-ai/web-llm library to deliver high-performance inference.

Detecting WebGPU and Shader-F16 Support

The foundation of MiniSearch's GPU acceleration lies in client/modules/webGpu.ts, which performs a runtime feature probe to determine hardware capabilities. The code first checks for the presence of navigator.gpu, then requests an adapter to inspect specific hardware features.

The critical detection focuses on the shader-f16 feature flag, which indicates support for 16-bit floating-point operations. This capability allows MiniSearch to run smaller, faster F16-quantized models instead of larger F32 variants.

// client/modules/webGpu.ts
export let isWebGPUAvailable = "gpu" in navigator;
export let isF16Supported = false;

if (isWebGPUAvailable) {
  try {
    const adapter = await (
      navigator as unknown as {
        gpu: { requestAdapter: () => Promise<{ features: Set<string> }> };
      }
    ).gpu.requestAdapter();

    if (!adapter) throw Error("Couldn't request WebGPU adapter.");
    isF16Supported = adapter.features.has("shader-f16");
  } catch {
    isWebGPUAvailable = false;
  }
}

If the adapter request fails or throws, the system gracefully degrades by setting isWebGPUAvailable to false, triggering a fallback to CPU-based inference.

Selecting the Default Model Based on GPU Capabilities

Once detection completes, client/modules/settings.ts uses the isF16Supported flag to select the appropriate default model. This conditional assignment ensures optimal performance without manual configuration.

// client/modules/settings.ts
export const defaultSettings = {
  enableWebGpu: true,
  webLlmModelId: isF16Supported
    ? VITE_WEBLLM_DEFAULT_F16_MODEL_ID
    : VITE_WEBLLM_DEFAULT_F32_MODEL_ID,
};

F16-quantized models offer reduced download sizes and faster inference on compatible hardware, while F32 models provide broader compatibility at the cost of higher VRAM requirements. The VITE_WEBLLM_DEFAULT_F16_MODEL_ID and VITE_WEBLLM_DEFAULT_F32_MODEL_ID environment variables define the specific model identifiers for each quantization type.

Initializing the WebLLM Engine

The engine initialization process resides in client/modules/textGenerationWithWebLlm.ts within the initializeWebLlmEngine function. This async function orchestrates model loading, cache verification, and engine configuration.

The implementation dynamically imports the WebLLM library only when needed, reducing initial bundle size. It checks hasModelInCache to avoid redundant downloads, then configures a progress callback that updates the UI during model fetching.

// client/modules/textGenerationWithWebLlm.ts
async function initializeWebLlmEngine() {
  const {
    CreateWebWorkerMLCEngine,
    CreateMLCEngine,
    hasModelInCache,
    prebuiltAppConfig,
  } = await import("@mlc-ai/web-llm");

  const selectedModelId = getSettings().webLlmModelId;

  updateModelSizeInMegabytes(
    prebuiltAppConfig.model_list.find(m => m.model_id === selectedModelId)
      ?.vram_required_MB || 0,
  );

  const isModelCached = await hasModelInCache(selectedModelId);
  let initProgressCallback: InitProgressCallback | undefined;

  if (!isModelCached) {
    initProgressCallback = report => {
      updateModelLoadingProgress(Math.round(report.progress * 100));
    };
  }

  const engineConfig: MLCEngineConfig = {
    initProgressCallback,
    logLevel: "SILENT",
  };

  const chatOptions: ChatOptions = {
    context_window_size: defaultContextSize,
    sliding_window_size: -1,
    attention_sink_size: -1,
  };

  return Worker
    ? await CreateWebWorkerMLCEngine(
        new Worker(new URL("./webLlmWorker.ts", import.meta.url), {
          type: "module",
        }),
        selectedModelId,
        engineConfig,
        chatOptions,
      )
    : await CreateMLCEngine(selectedModelId, engineConfig, chatOptions);
}

Key initialization steps include:

  • VRAM calculation: Extracts the model's memory requirement from prebuiltAppConfig.model_list to display download size warnings
  • Cache optimization: Skips progress reporting for cached models to improve perceived performance
  • Worker selection: Prefers CreateWebWorkerMLCEngine using webLlmWorker.ts to run inference off the main thread, falling back to CreateMLCEngine when Web Workers are unavailable
  • Context configuration: Sets context_window_size to the default value while disabling sliding window mechanisms for standard completion tasks

Streaming Text Generation

The top-level generation logic in client/modules/textGeneration.ts determines whether to use WebLLM based on the detection flags and user settings. When isWebGPUAvailable and settings.enableWebGpu are both true, it routes requests to the WebLLM implementation.

// client/modules/textGeneration.ts
if (isWebGPUAvailable && settings.enableWebGpu) {
  const { generateChatWithWebLlm } = await import("./textGenerationWithWebLlm");
  response = await generateChatWithWebLlm(lastMessages, onUpdate);
} else {
  // Fallback to wllama (CPU-only)
}

Once initialized, the engine exposes a standard OpenAI-compatible chat completion interface. The generateTextWithWebLlm function streams responses by calling engine.chat.completions.create with streaming enabled.

const completion = await engine.chat.completions.create({
  ...getDefaultChatCompletionCreateParamsStreaming(),
  messages: getDefaultChatMessages(getFormattedSearchResults(true)),
});

await handleStreamingResponse(completion, updateResponse, {
  shouldUpdateGeneratingState: true,
});

The handleStreamingResponse utility processes chunks as they arrive, updating the UI via PubSub events. After completion, engine.runtimeStatsText() provides telemetry on tokens-per-second and inference time.

Complete Implementation Workflow

For developers implementing similar architectures, the complete workflow follows this pattern:

  1. Detect capabilities using the navigator.gpu API
  2. Select models based on shader-f16 support
  3. Initialize engines with cache awareness and worker threading
  4. Stream completions through the standard chat completion interface

This architecture ensures that MiniSearch delivers GPU-accelerated AI search when hardware permits, while maintaining functionality through CPU fallbacks on unsupported browsers.

Summary

  • Capability Detection: client/modules/webGpu.ts probes for WebGPU availability and shader-f16 support to determine hardware acceleration potential
  • Intelligent Model Selection: client/modules/settings.ts automatically chooses between F16 and F32 quantized models based on detected GPU features
  • Worker-Based Initialization: client/modules/textGenerationWithWebLlm.ts initializes the WebLLM engine using Web Workers when available, with progress tracking for uncached models
  • Streaming Interface: The engine exposes a standard chat completion API that integrates with MiniSearch's existing text generation pipeline
  • Graceful Degradation: The system falls back to CPU-based inference (wllama) when WebGPU is unavailable or disabled

Frequently Asked Questions

What is WebGPU and why does MiniSearch use it?

WebGPU is a modern browser API that provides low-level access to GPU hardware for compute and rendering tasks. MiniSearch uses WebGPU through the @mlc-ai/web-llm library to run large language models directly in the browser, enabling privacy-preserving AI search without sending data to external servers. The API allows MiniSearch to leverage GPU acceleration for token generation, significantly improving inference speed compared to CPU-only implementations.

How does MiniSearch handle browsers without WebGPU support?

When client/modules/webGpu.ts detects that navigator.gpu is unavailable or the adapter request fails, it sets isWebGPUAvailable to false. This triggers client/modules/textGeneration.ts to skip the WebLLM initialization and instead import the CPU-based wllama implementation. Users still receive AI-generated responses, though with potentially slower inference speeds depending on their hardware.

What is the difference between F16 and F32 models in MiniSearch?

F16 (16-bit floating point) models use half-precision quantization, requiring approximately half the VRAM and download size of F32 (32-bit floating point) models. MiniSearch defaults to F16 models when shader-f16 support is detected in the WebGPU adapter, providing faster load times and inference. F32 models serve as the fallback for GPUs lacking 16-bit shader support, ensuring compatibility across a broader range of devices at the cost of higher resource consumption.

Why does MiniSearch use a WebWorker for the WebLLM engine?

MiniSearch instantiates the WebLLM engine via CreateWebWorkerMLCEngine in client/modules/textGenerationWithWebLlm.ts to prevent the UI from freezing during model loading and inference. Running the engine in a dedicated worker thread (defined in webLlmWorker.ts) keeps the main thread responsive, allowing progress indicators to update smoothly while the model downloads or generates text. If the browser does not support Web Workers, the code falls back to CreateMLCEngine for main-thread execution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →