How to Configure and Use AI Text Generation Backends in MiniSearch

MiniSearch routes text generation through a plugin-style architecture controlled by the inferenceType setting, supporting local browser models (WebLLM, Wllama) and remote APIs (OpenAI, AI Horde, Internal) via dynamic module imports.

MiniSearch implements a flexible backend system that lets you switch between privacy-focused browser inference and powerful remote APIs without modifying core logic. The central dispatcher in client/modules/textGeneration.ts reads settings.inferenceType to determine whether to load WebLLM, Wllama, OpenAI, AI Horde, or the internal API generator. All backends share a common contract: they accept an array of ChatMessage objects and stream partial results to the UI through the onUpdate callback managed by the PubSub store.

Global Settings – Selecting the Inference Backend

The backend selection originates in client/modules/settings.ts, which defines the available inferenceTypes and their default configurations.

  • "browser": Runs models locally using WebLLM (WebGPU) or Wllama (WebAssembly)
  • "openai": Connects to OpenAI-compatible servers
  • "horde": Uses the distributed AI Horde network
  • "internal": Calls a self-hosted MiniSearch inference endpoint (requires VITE_INTERNAL_API_ENABLED)

To switch backends programmatically, update the global settings store:

import { setSettings } from "./client/modules/pubSub";

setSettings(prev => ({
  ...prev,
  inferenceType: "openai",
  openAiApiBaseUrl: "https://api.openai.com/v1",
  openAiApiKey: "sk-your-key-here",
  openAiApiModel: "gpt-4o-mini"
}));

In client/modules/textGeneration.ts, the searchAndRespond() function evaluates settings.inferenceType and dynamically imports the appropriate generator module. When inferenceType is "browser", the code checks WebGPU availability via isWebGPUAvailable to decide between WebLLM and Wllama.

Browser-Based Local Models

Setting inferenceType to "browser" enables private, local inference. MiniSearch automatically selects the optimal engine based on hardware capabilities and the enableWebGpu flag.

Configuring WebLLM (MLC AI)

WebLLM provides GPU-accelerated inference using WebGPU. The configuration is stored in client/modules/textGenerationWithWebLlm.ts.

Required settings:

  • settings.enableWebGpu: Must be true (default when WebGPU is available)
  • settings.webLlmModelId: Model identifier, defaulting to the value specified in VITE_WEBLLM_DEFAULT_*_MODEL_ID

Code path executed:

// From client/modules/textGeneration.ts
const engine = await initializeWebLlmEngine();
const completion = await engine.chat.completions.create({
  ...getDefaultChatCompletionCreateParamsStreaming(),
  messages: getDefaultChatMessages(getFormattedSearchResults(true))
});
await handleStreamingResponse(completion, updateResponse, { 
  shouldUpdateGeneratingState: true 
});
engine.unload(); // Free WASM memory after generation

Example configuration:

setSettings(s => ({
  ...s,
  inferenceType: "browser",
  enableWebGpu: true,
  webLlmModelId: "Llama-2-7b-chat-hf-q4f32_1"
}));

Configuring Wllama (WebAssembly)

Wllama serves as the fallback for devices without WebGPU support, utilizing WebAssembly for CPU inference. The implementation resides in client/modules/textGenerationWithWllama.ts.

Required settings:

  • settings.enableWebGpu: Set to false to force Wllama selection
  • settings.wllamaModelId: Hugging Face model ID, defaulting from VITE_WLLAMA_DEFAULT_MODEL_ID

The initializeWllamaInstance() function in client/modules/wllama.ts handles model downloading from Hugging Face and warm-up:

const { wllama, model } = await initializeWllamaInstance(progressCallback);
const response = await generateWithWllama({ wllama, model, ... });

Example configuration:

setSettings(s => ({
  ...s,
  inferenceType: "browser",
  enableWebGpu: false,
  wllamaModelId: "Meta-Llama-3-8B-Instruct"
}));

Remote Server Backends

For scenarios requiring larger models or faster inference, MiniSearch supports three remote backends that stream responses via HTTP through identical PubSub-driven flows.

OpenAI-Compatible APIs

The OpenAI integration in client/modules/textGenerationWithOpenAi.ts uses the Vercel AI SDK to stream completions from any OpenAI-compatible endpoint.

Configuration keys:

  • settings.openAiApiBaseUrl: Endpoint URL (e.g., "https://api.openai.com/v1")
  • settings.openAiApiKey: Authentication token (never persisted to disk)
  • settings.openAiApiModel: Model name; if empty, the system falls back to a random available model per listOpenAiCompatibleModels in shared/openaiModels.ts

Generation flow:

const openaiProvider = createOpenAICompatible({
  name: settings.openAiApiBaseUrl,
  baseURL: settings.openAiApiBaseUrl,
  apiKey: settings.openAiApiKey,
});
const stream = streamText({
  model: openaiProvider.chatModel(effectiveModel),
  messages,
  // temperature, top_p, etc.
});
for await (const part of stream.fullStream) {
  // Streaming deltas forwarded via onUpdate()
}

AI Horde Distributed Inference

AI Horde provides access to a distributed network of volunteer GPUs. The implementation in client/modules/textGenerationWithHorde.ts includes unique parallel-generation logic that races two simultaneous requests and returns the faster result.

Configuration:

  • settings.hordeApiKey: Anonymous key "0000000000" works with limited priority, or use a personal key for higher speeds
  • settings.hordeModel: Specific model name, or empty string for auto-selection

The executeHordeGeneration() function handles status polling and cancellation of the slower parallel request:

const messages = getDefaultChatMessages(getFormattedSearchResults(true));
await executeHordeGeneration(messages, text => updateResponse(text));

Internal Self-Hosted API

The Internal API connects to a MiniSearch backend server on the same origin, useful for private deployments with centralized GPU resources. The client module client/modules/textGenerationWithInternalApi.ts posts to the /inference endpoint.

Security mechanism:

  • Requests include a per-session token via getSearchTokenHash() in the query string
  • No API keys exposed in client bundles

Request structure:

const inferenceUrl = new URL("/inference", self.location.origin);
inferenceUrl.searchParams.set("token", await getSearchTokenHash());

const response = await fetch(inferenceUrl, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    ...getDefaultChatCompletionCreateParamsStreaming(),
    messages,
  })
});
// NDJSON streaming response processed via onChunk()

Programmatic Backend Switching

You can wrap the settings logic in a reusable hook to switch backends dynamically based on user selection:

import { useEffect } from "react";
import { getSettings, setSettings } from "./client/modules/pubSub";
import { searchAndRespond } from "./client/modules/textGeneration";

export function useAiBackend(backend: "openai" | "horde" | "web-llm" | "wllama" | "internal") {
  useEffect(() => {
    const map = {
      "openai": "openai",
      "horde": "horde",
      "web-llm": "browser",
      "wllama": "browser",
      "internal": "internal"
    };
    
    setSettings(prev => ({
      ...prev,
      inferenceType: map[backend],
      ...(backend === "web-llm" && { enableWebGpu: true, webLlmModelId: "Llama-2-7b-chat-hf-q4f32_1" }),
      ...(backend === "wllama" && { enableWebGpu: false, wllamaModelId: "Meta-Llama-3-8B-Instruct" })
    }));
  }, [backend]);

  useEffect(() => {
    if (getSettings().enableAiResponse) {
      searchAndRespond(); // Automatically routes to selected generator
    }
  }, [getSettings().inferenceType]);
}

Summary

  • Backend selection is controlled by settings.inferenceType in client/modules/settings.ts, supporting "browser", "openai", "horde", and "internal" values defined in inferenceTypes.
  • Browser inference automatically chooses WebLLM when enableWebGpu is true and WebGPU is available, otherwise falling back to Wllama for CPU execution via WebAssembly.
  • Remote APIs require specific credential settings (openAiApiKey, hordeApiKey) but follow identical streaming contracts through the updateResponse callback.
  • Security: The OpenAI key remains in client memory only and is never logged; the Internal API uses ephemeral session tokens via getSearchTokenHash().
  • Dynamic switching: Change inferenceType at any time via setSettings(); subsequent calls to searchAndRespond() automatically import and execute the correct generator module from client/modules/textGenerationWith[Backend].ts.

Frequently Asked Questions

How does MiniSearch choose between WebLLM and Wllama?

When inferenceType is set to "browser", the searchAndRespond() function in client/modules/textGeneration.ts checks isWebGPUAvailable and settings.enableWebGpu. If both are true, it dynamically imports textGenerationWithWebLlm.ts for GPU acceleration; otherwise, it loads textGenerationWithWllama.ts for WebAssembly-based CPU inference.

Is the OpenAI API key stored securely?

Yes. According to the MiniSearch source code in client/modules/settings.ts, the openAiApiKey value resides only in the client's memory within the PubSub store. It is never persisted to local storage, IndexedDB, or server logs, ensuring the key remains secret between sessions.

Can I use a custom OpenAI-compatible server instead of OpenAI's official API?

Absolutely. The OpenAI backend in client/modules/textGenerationWithOpenAi.ts uses createOpenAICompatible from the Vercel AI SDK, which accepts any baseURL. Simply set openAiApiBaseUrl to your local LLM server (e.g., LM Studio, Ollama, or vLLM endpoint) and provide the appropriate model name in openAiApiModel. The system will automatically handle streaming from your custom endpoint.

What happens if AI Horde has no available workers?

The executeHordeGeneration() function in client/modules/textGenerationWithHorde.ts includes polling logic that waits for available workers. If you specify a specific hordeModel and no workers are available, the request will queue until capacity becomes available. For faster responses, leave hordeModel empty to allow Horde to select any available model, or leverage the parallel generation feature which races two independent requests and cancels the slower one.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →