How to Add Custom Text Generation Backends to MiniSearch: A Complete Developer Guide

To add a custom text generation backend to MiniSearch, declare a new inference type in client/modules/settings.ts, extend the model name resolver and runtime dispatch logic in client/modules/textGeneration.ts, and implement a new module exporting generateTextWith* and generateChatWith* functions that conform to the existing API contract.

MiniSearch, the privacy-focused search interface developed by felladrin, delegates every LLM request to a swappable backend implementation keyed by settings.inferenceType. This modular architecture allows you to integrate proprietary models, internal corporate APIs, or experimental inference services without touching the core search logic or UI components. This guide provides the exact file paths, function signatures, and code patterns required to extend the architecture.

Understanding the Backend Architecture

MiniSearch selects backends based on the inferenceType value stored in the global settings state. The delegation logic resides primarily in client/modules/textGeneration.ts and operates through two distinct mechanisms that you must extend.

Model Name Resolution

The getCurrentModelName() function (lines 43-58) maps inference types to human-readable identifiers for UI labels and telemetry. This switch statement returns a string based on the active backend:

function getCurrentModelName(): string {
  const settings = getSettings();
  switch (settings.inferenceType) {
    case "openai":   return settings.openAiApiModel || "";
    case "horde":   return "AI Horde";
    case "internal":return "Internal API";
    case "browser": return settings.enableWebGpu
                      ? settings.webLlmModelId || "WebLLM"
                      : settings.wllamaModelId || "Wllama";
    default:        return "Unknown";
  }
}

Runtime Dispatch with Dynamic Imports

Both the search-and-respond flow and the chat flow use conditional dynamic imports to load backend-specific generators only when needed. The chat path (around line 207) demonstrates this pattern:

if (settings.inferenceType === "openai") {
  const { generateChatWithOpenAi } = await import("./textGenerationWithOpenAi");
  response = await generateChatWithOpenAi(lastMessages, onUpdate);
} else if (settings.inferenceType === "internal") {
  const { generateChatWithInternalApi } = await import("./textGenerationWithInternalApi");
  response = await generateChatWithInternalApi(lastMessages, onUpdate);
} else if (settings.inferenceType === "horde") {
  const { generateChatWithHorde } = await import("./textGenerationWithHorde");
  response = await generateChatWithHorde(lastMessages, onUpdate);
}

The same pattern appears in the standalone text generation path (around line 207-363), ensuring that backend code is loaded on-demand to minimize bundle size.

Step-by-Step Implementation Guide

To plug a custom backend (e.g., MyLLM) into felladrin/minisearch, complete these five concrete steps:

1. Declare the New Inference Type

Add your backend identifier to the inferenceTypes array in client/modules/settings.ts. This registration automatically exposes the option in the Settings UI dropdown without requiring React component modifications.

// client/modules/settings.ts
export const inferenceTypes = [
  { value: "browser", label: "In the browser (Private)" },
  { value: "openai",  label: "Remote server (API)" },
  { value: "horde",   label: "AI Horde (Pre-configured)" },
  { value: "myBackend", label: "My LLM (Custom)" },   // ← new entry
  ...(VITE_INTERNAL_API_ENABLED
    ? [{ value: "internal", label: VITE_INTERNAL_API_NAME }]
    : []),
];

2. Extend the Model Name Resolver

Add a case to the switch statement in getCurrentModelName() within client/modules/textGeneration.ts to return a readable identifier when your backend is active:

// client/modules/textGeneration.ts (add after existing cases)
case "myBackend":
  return "My LLM";

3. Implement the Backend Module

Create a new file (e.g., client/modules/textGenerationWithMyBackend.ts) that exports two functions: generateTextWithMyBackend for standalone search responses and generateChatWithMyBackend for conversational streaming. Both must use the global pub/sub utilities (updateResponse, updateTextGenerationState) to communicate with the UI.

// client/modules/textGenerationWithMyBackend.ts
import type { ChatMessage } from "./types";
import {
  getSettings,
  updateResponse,
  updateTextGenerationState,
  getTextGenerationState,
} from "./pubSub";
import { ChatGenerationError } from "./textGenerationUtilities";

export async function generateTextWithMyBackend(): Promise<void> {
  await canStartResponding();
  updateTextGenerationState("preparingToGenerate");

  const settings = getSettings();
  const prompt = getSystemPrompt(getFormattedSearchResults(true));

  const resp = await fetch(settings.myBackendApiUrl, {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      "Authorization": `Bearer ${settings.myBackendApiKey}`,
    },
    body: JSON.stringify({ 
      prompt, 
      maxTokens: 1024, 
      temperature: settings.inferenceTemperature 
    })
  });

  if (!resp.ok) throw new ChatGenerationError(`MyBackend error ${resp.status}`);

  const { text } = await resp.json();
  updateResponse(text);
}

export async function generateChatWithMyBackend(
  messages: ChatMessage[],
  onUpdate: (partial: string) => void,
): Promise<string> {
  const settings = getSettings();

  const resp = await fetch(settings.myBackendApiUrl, {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      "Authorization": `Bearer ${settings.myBackendApiKey}`,
    },
    body: JSON.stringify({ messages })
  });

  if (!resp.ok) throw new ChatGenerationError(`MyBackend chat error ${resp.status}`);

  const reader = resp.body!.getReader();
  const decoder = new TextDecoder();
  let full = "";
  
  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    const chunk = decoder.decode(value);
    
    for (const line of chunk.split("\n")) {
      if (line.startsWith("data:")) {
        const payload = JSON.parse(line.slice(5));
        full += payload.text;
        onUpdate(full);
      }
    }
    
    if (getTextGenerationState() === "interrupted") {
      throw new ChatGenerationError("Chat generation interrupted");
    }
  }
  return full;
}

4. Hook the Dispatch Logic

Extend the conditional chains in client/modules/textGeneration.ts to import and invoke your functions when settings.inferenceType matches your new type. Add entries in both the chat generation block and the searchAndRespond function:

// Inside generateChatResponse (chat flow)
if (settings.inferenceType === "myBackend") {
  const { generateChatWithMyBackend } = await import("./textGenerationWithMyBackend");
  response = await generateChatWithMyBackend(lastMessages, onUpdate);
}
// Inside searchAndRespond (search flow)
else if (settings.inferenceType === "myBackend") {
  const { generateTextWithMyBackend } = await import("./textGenerationWithMyBackend");
  await generateTextWithMyBackend();
}

5. Verify UI Integration

The Settings UI consumes the inferenceTypes array directly from client/modules/settings.ts, so your new backend appears automatically in the dropdown selector. If you need custom configuration fields (e.g., API URL inputs), add them to the AISettings components in client/components/Pages/Main/Menu/AISettings/.

Key Architectural Benefits

  • Dynamic imports keep the bundle size minimal—a backend module is only loaded when the user selects it.
  • Centralized settings (settings.inferenceType) provide a single source of truth for both the UI and generation engine.
  • Uniform API contract requires every backend to return plain strings (or streamed chunks) and update the global response store via updateResponse, keeping MiniSearch agnostic to the underlying provider.

Summary

Frequently Asked Questions

What files must I modify to add a custom LLM provider to MiniSearch?

You must edit client/modules/settings.ts to register the inference type, client/modules/textGeneration.ts to add dispatch logic and model name resolution, and create a new implementation file (e.g., textGenerationWithMyBackend.ts). The Settings UI updates automatically when you add entries to the inferenceTypes array.

Does MiniSearch support streaming responses from custom backends?

Yes. The architecture expects your generateChatWith* function to accept an onUpdate callback that receives partial text chunks. Implement streaming by reading the response body via getReader() and invoking onUpdate with decoded chunks, following the pattern used in the WebLLM and AI Horde implementations.

How does MiniSearch handle authentication for custom APIs?

Store sensitive credentials in the settings store by adding fields to defaultSettings in client/modules/settings.ts, then access them via getSettings() inside your backend module. Never hardcode API keys; use the settings pub/sub pattern to retrieve user-configured values at runtime.

Can I add multiple custom backends simultaneously?

Yes. Each backend requires a unique inferenceType string and a separate implementation module. Users switch between them via the Settings UI dropdown, and MiniSearch dynamically imports only the active backend's code, ensuring the client bundle remains efficient regardless of how many providers you define.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →