# How to Configure and Use AI Text Generation Backends in MiniSearch

> Learn to configure and use AI text generation backends in MiniSearch. Explore WebLLM, Wllama, OpenAI, AI Horde, and internal APIs with this guide. Optimize your search experience.

- Repository: [Victor Nogueira/minisearch](https://github.com/felladrin/minisearch)
- Tags: how-to-guide
- Published: 2026-03-01

---

**MiniSearch routes text generation through a plugin-style architecture controlled by the `inferenceType` setting, supporting local browser models (WebLLM, Wllama) and remote APIs (OpenAI, AI Horde, Internal) via dynamic module imports.**

MiniSearch implements a flexible backend system that lets you switch between privacy-focused browser inference and powerful remote APIs without modifying core logic. The central dispatcher in [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts) reads `settings.inferenceType` to determine whether to load WebLLM, Wllama, OpenAI, AI Horde, or the internal API generator. All backends share a common contract: they accept an array of `ChatMessage` objects and stream partial results to the UI through the `onUpdate` callback managed by the PubSub store.

## Global Settings – Selecting the Inference Backend

The backend selection originates in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts), which defines the available `inferenceTypes` and their default configurations.

- **"browser"**: Runs models locally using WebLLM (WebGPU) or Wllama (WebAssembly)
- **"openai"**: Connects to OpenAI-compatible servers
- **"horde"**: Uses the distributed AI Horde network
- **"internal"**: Calls a self-hosted MiniSearch inference endpoint (requires `VITE_INTERNAL_API_ENABLED`)

To switch backends programmatically, update the global settings store:

```typescript
import { setSettings } from "./client/modules/pubSub";

setSettings(prev => ({
  ...prev,
  inferenceType: "openai",
  openAiApiBaseUrl: "https://api.openai.com/v1",
  openAiApiKey: "sk-your-key-here",
  openAiApiModel: "gpt-4o-mini"
}));

```

In [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts), the `searchAndRespond()` function evaluates `settings.inferenceType` and dynamically imports the appropriate generator module. When `inferenceType` is `"browser"`, the code checks WebGPU availability via `isWebGPUAvailable` to decide between WebLLM and Wllama.

## Browser-Based Local Models

Setting `inferenceType` to `"browser"` enables private, local inference. MiniSearch automatically selects the optimal engine based on hardware capabilities and the `enableWebGpu` flag.

### Configuring WebLLM (MLC AI)

**WebLLM** provides GPU-accelerated inference using WebGPU. The configuration is stored in [`client/modules/textGenerationWithWebLlm.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithWebLlm.ts).

Required settings:
- `settings.enableWebGpu`: Must be `true` (default when WebGPU is available)
- `settings.webLlmModelId`: Model identifier, defaulting to the value specified in `VITE_WEBLLM_DEFAULT_*_MODEL_ID`

Code path executed:

```typescript
// From client/modules/textGeneration.ts
const engine = await initializeWebLlmEngine();
const completion = await engine.chat.completions.create({
  ...getDefaultChatCompletionCreateParamsStreaming(),
  messages: getDefaultChatMessages(getFormattedSearchResults(true))
});
await handleStreamingResponse(completion, updateResponse, { 
  shouldUpdateGeneratingState: true 
});
engine.unload(); // Free WASM memory after generation

```

Example configuration:

```typescript
setSettings(s => ({
  ...s,
  inferenceType: "browser",
  enableWebGpu: true,
  webLlmModelId: "Llama-2-7b-chat-hf-q4f32_1"
}));

```

### Configuring Wllama (WebAssembly)

**Wllama** serves as the fallback for devices without WebGPU support, utilizing WebAssembly for CPU inference. The implementation resides in [`client/modules/textGenerationWithWllama.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithWllama.ts).

Required settings:
- `settings.enableWebGpu`: Set to `false` to force Wllama selection
- `settings.wllamaModelId`: Hugging Face model ID, defaulting from `VITE_WLLAMA_DEFAULT_MODEL_ID`

The `initializeWllamaInstance()` function in [`client/modules/wllama.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/wllama.ts) handles model downloading from Hugging Face and warm-up:

```typescript
const { wllama, model } = await initializeWllamaInstance(progressCallback);
const response = await generateWithWllama({ wllama, model, ... });

```

Example configuration:

```typescript
setSettings(s => ({
  ...s,
  inferenceType: "browser",
  enableWebGpu: false,
  wllamaModelId: "Meta-Llama-3-8B-Instruct"
}));

```

## Remote Server Backends

For scenarios requiring larger models or faster inference, MiniSearch supports three remote backends that stream responses via HTTP through identical PubSub-driven flows.

### OpenAI-Compatible APIs

The OpenAI integration in [`client/modules/textGenerationWithOpenAi.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithOpenAi.ts) uses the Vercel AI SDK to stream completions from any OpenAI-compatible endpoint.

Configuration keys:
- `settings.openAiApiBaseUrl`: Endpoint URL (e.g., `"https://api.openai.com/v1"`)
- `settings.openAiApiKey`: Authentication token (never persisted to disk)
- `settings.openAiApiModel`: Model name; if empty, the system falls back to a random available model per `listOpenAiCompatibleModels` in [`shared/openaiModels.ts`](https://github.com/felladrin/minisearch/blob/main/shared/openaiModels.ts)

Generation flow:

```typescript
const openaiProvider = createOpenAICompatible({
  name: settings.openAiApiBaseUrl,
  baseURL: settings.openAiApiBaseUrl,
  apiKey: settings.openAiApiKey,
});
const stream = streamText({
  model: openaiProvider.chatModel(effectiveModel),
  messages,
  // temperature, top_p, etc.
});
for await (const part of stream.fullStream) {
  // Streaming deltas forwarded via onUpdate()
}

```

### AI Horde Distributed Inference

**AI Horde** provides access to a distributed network of volunteer GPUs. The implementation in [`client/modules/textGenerationWithHorde.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithHorde.ts) includes unique parallel-generation logic that races two simultaneous requests and returns the faster result.

Configuration:
- `settings.hordeApiKey`: Anonymous key `"0000000000"` works with limited priority, or use a personal key for higher speeds
- `settings.hordeModel`: Specific model name, or empty string for auto-selection

The `executeHordeGeneration()` function handles status polling and cancellation of the slower parallel request:

```typescript
const messages = getDefaultChatMessages(getFormattedSearchResults(true));
await executeHordeGeneration(messages, text => updateResponse(text));

```

### Internal Self-Hosted API

The **Internal API** connects to a MiniSearch backend server on the same origin, useful for private deployments with centralized GPU resources. The client module [`client/modules/textGenerationWithInternalApi.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithInternalApi.ts) posts to the `/inference` endpoint.

Security mechanism:
- Requests include a per-session token via `getSearchTokenHash()` in the query string
- No API keys exposed in client bundles

Request structure:

```typescript
const inferenceUrl = new URL("/inference", self.location.origin);
inferenceUrl.searchParams.set("token", await getSearchTokenHash());

const response = await fetch(inferenceUrl, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    ...getDefaultChatCompletionCreateParamsStreaming(),
    messages,
  })
});
// NDJSON streaming response processed via onChunk()

```

## Programmatic Backend Switching

You can wrap the settings logic in a reusable hook to switch backends dynamically based on user selection:

```typescript
import { useEffect } from "react";
import { getSettings, setSettings } from "./client/modules/pubSub";
import { searchAndRespond } from "./client/modules/textGeneration";

export function useAiBackend(backend: "openai" | "horde" | "web-llm" | "wllama" | "internal") {
  useEffect(() => {
    const map = {
      "openai": "openai",
      "horde": "horde",
      "web-llm": "browser",
      "wllama": "browser",
      "internal": "internal"
    };
    
    setSettings(prev => ({
      ...prev,
      inferenceType: map[backend],
      ...(backend === "web-llm" && { enableWebGpu: true, webLlmModelId: "Llama-2-7b-chat-hf-q4f32_1" }),
      ...(backend === "wllama" && { enableWebGpu: false, wllamaModelId: "Meta-Llama-3-8B-Instruct" })
    }));
  }, [backend]);

  useEffect(() => {
    if (getSettings().enableAiResponse) {
      searchAndRespond(); // Automatically routes to selected generator
    }
  }, [getSettings().inferenceType]);
}

```

## Summary

- **Backend selection** is controlled by `settings.inferenceType` in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts), supporting `"browser"`, `"openai"`, `"horde"`, and `"internal"` values defined in `inferenceTypes`.
- **Browser inference** automatically chooses **WebLLM** when `enableWebGpu` is true and WebGPU is available, otherwise falling back to **Wllama** for CPU execution via WebAssembly.
- **Remote APIs** require specific credential settings (`openAiApiKey`, `hordeApiKey`) but follow identical streaming contracts through the `updateResponse` callback.
- **Security**: The OpenAI key remains in client memory only and is never logged; the Internal API uses ephemeral session tokens via `getSearchTokenHash()`.
- **Dynamic switching**: Change `inferenceType` at any time via `setSettings()`; subsequent calls to `searchAndRespond()` automatically import and execute the correct generator module from `client/modules/textGenerationWith[Backend].ts`.

## Frequently Asked Questions

### How does MiniSearch choose between WebLLM and Wllama?

When `inferenceType` is set to `"browser"`, the `searchAndRespond()` function in [`client/modules/textGeneration.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGeneration.ts) checks `isWebGPUAvailable` and `settings.enableWebGpu`. If both are true, it dynamically imports [`textGenerationWithWebLlm.ts`](https://github.com/felladrin/minisearch/blob/main/textGenerationWithWebLlm.ts) for GPU acceleration; otherwise, it loads [`textGenerationWithWllama.ts`](https://github.com/felladrin/minisearch/blob/main/textGenerationWithWllama.ts) for WebAssembly-based CPU inference.

### Is the OpenAI API key stored securely?

**Yes.** According to the MiniSearch source code in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts), the `openAiApiKey` value resides only in the client's memory within the PubSub store. It is never persisted to local storage, IndexedDB, or server logs, ensuring the key remains secret between sessions.

### Can I use a custom OpenAI-compatible server instead of OpenAI's official API?

**Absolutely.** The OpenAI backend in [`client/modules/textGenerationWithOpenAi.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithOpenAi.ts) uses `createOpenAICompatible` from the Vercel AI SDK, which accepts any `baseURL`. Simply set `openAiApiBaseUrl` to your local LLM server (e.g., LM Studio, Ollama, or vLLM endpoint) and provide the appropriate model name in `openAiApiModel`. The system will automatically handle streaming from your custom endpoint.

### What happens if AI Horde has no available workers?

The `executeHordeGeneration()` function in [`client/modules/textGenerationWithHorde.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithHorde.ts) includes polling logic that waits for available workers. If you specify a specific `hordeModel` and no workers are available, the request will queue until capacity becomes available. For faster responses, leave `hordeModel` empty to allow Horde to select any available model, or leverage the parallel generation feature which races two independent requests and cancels the slower one.