# Browser-Based AI Model Loading and WebGPU Integration in MiniSearch

> Explore browser-based AI model loading with WebGPU integration in MiniSearch. Learn how this architecture enables efficient LLM execution directly in your browser.

- Repository: [Victor Nogueira/minisearch](https://github.com/felladrin/minisearch)
- Tags: deep-dive
- Published: 2026-03-01

---

**MiniSearch runs large language models directly in the browser by detecting WebGPU capabilities, selecting appropriate quantized models based on shader-f16 support, initializing the WebLLM engine with progress tracking, and streaming completions through a WebWorker-backed architecture.**

MiniSearch enables privacy-preserving AI search by running large language models entirely in the browser without server-side dependencies. This article examines how the `felladrin/minisearch` repository implements browser-based AI model loading and WebGPU integration using the **@mlc-ai/web-llm** library to deliver high-performance inference.

## Detecting WebGPU and Shader-F16 Support

The foundation of MiniSearch's GPU acceleration lies in [`client/modules/webGpu.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/webGpu.ts), which performs a runtime feature probe to determine hardware capabilities. The code first checks for the presence of `navigator.gpu`, then requests an adapter to inspect specific hardware features.

The critical detection focuses on the **shader-f16** feature flag, which indicates support for 16-bit floating-point operations. This capability allows MiniSearch to run smaller, faster F16-quantized models instead of larger F32 variants.

```typescript
// client/modules/webGpu.ts
export let isWebGPUAvailable = "gpu" in navigator;
export let isF16Supported = false;

if (isWebGPUAvailable) {
  try {
    const adapter = await (
      navigator as unknown as {
        gpu: { requestAdapter: () => Promise<{ features: Set<string> }> };
      }
    ).gpu.requestAdapter();

    if (!adapter) throw Error("Couldn't request WebGPU adapter.");
    isF16Supported = adapter.features.has("shader-f16");
  } catch {
    isWebGPUAvailable = false;
  }
}

```

If the adapter request fails or throws, the system gracefully degrades by setting `isWebGPUAvailable` to `false`, triggering a fallback to CPU-based inference.

## Selecting the Default Model Based on GPU Capabilities

Once detection completes, [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts) uses the `isF16Supported` flag to select the appropriate default model. This conditional assignment ensures optimal performance without manual configuration.

```typescript
// client/modules/settings.ts
export const defaultSettings = {
  enableWebGpu: true,
  webLlmModelId: isF16Supported
    ? VITE_WEBLLM_DEFAULT_F16_MODEL_ID
    : VITE_WEBLLM_DEFAULT_F32_MODEL_ID,
};

```

**F16-quantized models** offer reduced download sizes and faster inference on compatible hardware, while **F32 models** provide broader compatibility at the cost of higher VRAM requirements. The `VITE_WEBLLM_DEFAULT_F16_MODEL_ID` and `VITE_WEBLLM_DEFAULT_F32_MODEL_ID` environment variables define the specific model identifiers for each quantization type.

## Initializing the WebLLM Engine

The engine initialization process resides in [`client/modules/textGenerationWithWebLlm.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithWebLlm.ts) within the `initializeWebLlmEngine` function. This async function orchestrates model loading, cache verification, and engine configuration.

The implementation dynamically imports the WebLLM library only when needed, reducing initial bundle size. It checks `hasModelInCache` to avoid redundant downloads, then configures a progress callback that updates the UI during model fetching.

```typescript
// client/modules/textGenerationWithWebLlm.ts
async function initializeWebLlmEngine() {
  const {
    CreateWebWorkerMLCEngine,
    CreateMLCEngine,
    hasModelInCache,
    prebuiltAppConfig,
  } = await import("@mlc-ai/web-llm");

  const selectedModelId = getSettings().webLlmModelId;

  updateModelSizeInMegabytes(
    prebuiltAppConfig.model_list.find(m => m.model_id === selectedModelId)
      ?.vram_required_MB || 0,
  );

  const isModelCached = await hasModelInCache(selectedModelId);
  let initProgressCallback: InitProgressCallback | undefined;

  if (!isModelCached) {
    initProgressCallback = report => {
      updateModelLoadingProgress(Math.round(report.progress * 100));
    };
  }

  const engineConfig: MLCEngineConfig = {
    initProgressCallback,
    logLevel: "SILENT",
  };

  const chatOptions: ChatOptions = {
    context_window_size: defaultContextSize,
    sliding_window_size: -1,
    attention_sink_size: -1,
  };

  return Worker
    ? await CreateWebWorkerMLCEngine(
        new Worker(new URL("./webLlmWorker.ts", import.meta.url), {
          type: "module",
        }),
        selectedModelId,
        engineConfig,
        chatOptions,
      )
    : await CreateMLCEngine(selectedModelId, engineConfig, chatOptions);
}

```

**Key initialization steps include:**

- **VRAM calculation**: Extracts the model's memory requirement from `prebuiltAppConfig.model_list` to display download size warnings
- **Cache optimization**: Skips progress reporting for cached models to improve perceived performance
- **Worker selection**: Prefers `CreateWebWorkerMLCEngine` using [`webLlmWorker.ts`](https://github.com/felladrin/minisearch/blob/main/webLlmWorker.ts) to run inference off the main thread, falling back to `CreateMLCEngine` when Web Workers are unavailable
- **Context configuration**: Sets `context_window_size` to the default value while disabling sliding window mechanisms for standard completion tasks

## Streaming Text Generation

The top-level generation logic in [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts) determines whether to use WebLLM based on the detection flags and user settings. When `isWebGPUAvailable` and `settings.enableWebGpu` are both true, it routes requests to the WebLLM implementation.

```typescript
// client/modules/textGeneration.ts
if (isWebGPUAvailable && settings.enableWebGpu) {
  const { generateChatWithWebLlm } = await import("./textGenerationWithWebLlm");
  response = await generateChatWithWebLlm(lastMessages, onUpdate);
} else {
  // Fallback to wllama (CPU-only)
}

```

Once initialized, the engine exposes a standard OpenAI-compatible chat completion interface. The `generateTextWithWebLlm` function streams responses by calling `engine.chat.completions.create` with streaming enabled.

```typescript
const completion = await engine.chat.completions.create({
  ...getDefaultChatCompletionCreateParamsStreaming(),
  messages: getDefaultChatMessages(getFormattedSearchResults(true)),
});

await handleStreamingResponse(completion, updateResponse, {
  shouldUpdateGeneratingState: true,
});

```

The `handleStreamingResponse` utility processes chunks as they arrive, updating the UI via PubSub events. After completion, `engine.runtimeStatsText()` provides telemetry on tokens-per-second and inference time.

## Complete Implementation Workflow

For developers implementing similar architectures, the complete workflow follows this pattern:

1. **Detect capabilities** using the navigator.gpu API
2. **Select models** based on shader-f16 support
3. **Initialize engines** with cache awareness and worker threading
4. **Stream completions** through the standard chat completion interface

This architecture ensures that MiniSearch delivers GPU-accelerated AI search when hardware permits, while maintaining functionality through CPU fallbacks on unsupported browsers.

## Summary

- **Capability Detection**: [`client/modules/webGpu.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/webGpu.ts) probes for WebGPU availability and shader-f16 support to determine hardware acceleration potential
- **Intelligent Model Selection**: [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts) automatically chooses between F16 and F32 quantized models based on detected GPU features
- **Worker-Based Initialization**: [`client/modules/textGenerationWithWebLlm.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithWebLlm.ts) initializes the WebLLM engine using Web Workers when available, with progress tracking for uncached models
- **Streaming Interface**: The engine exposes a standard chat completion API that integrates with MiniSearch's existing text generation pipeline
- **Graceful Degradation**: The system falls back to CPU-based inference (wllama) when WebGPU is unavailable or disabled

## Frequently Asked Questions

### What is WebGPU and why does MiniSearch use it?

WebGPU is a modern browser API that provides low-level access to GPU hardware for compute and rendering tasks. MiniSearch uses WebGPU through the @mlc-ai/web-llm library to run large language models directly in the browser, enabling privacy-preserving AI search without sending data to external servers. The API allows MiniSearch to leverage GPU acceleration for token generation, significantly improving inference speed compared to CPU-only implementations.

### How does MiniSearch handle browsers without WebGPU support?

When [`client/modules/webGpu.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/webGpu.ts) detects that `navigator.gpu` is unavailable or the adapter request fails, it sets `isWebGPUAvailable` to `false`. This triggers [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts) to skip the WebLLM initialization and instead import the CPU-based wllama implementation. Users still receive AI-generated responses, though with potentially slower inference speeds depending on their hardware.

### What is the difference between F16 and F32 models in MiniSearch?

F16 (16-bit floating point) models use half-precision quantization, requiring approximately half the VRAM and download size of F32 (32-bit floating point) models. MiniSearch defaults to F16 models when `shader-f16` support is detected in the WebGPU adapter, providing faster load times and inference. F32 models serve as the fallback for GPUs lacking 16-bit shader support, ensuring compatibility across a broader range of devices at the cost of higher resource consumption.

### Why does MiniSearch use a WebWorker for the WebLLM engine?

MiniSearch instantiates the WebLLM engine via `CreateWebWorkerMLCEngine` in [`client/modules/textGenerationWithWebLlm.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithWebLlm.ts) to prevent the UI from freezing during model loading and inference. Running the engine in a dedicated worker thread (defined in [`webLlmWorker.ts`](https://github.com/felladrin/minisearch/blob/main/webLlmWorker.ts)) keeps the main thread responsive, allowing progress indicators to update smoothly while the model downloads or generates text. If the browser does not support Web Workers, the code falls back to `CreateMLCEngine` for main-thread execution.