# How to Integrate MiniSearch with External OpenAI-Compatible APIs: Base URL, API Key, and Model Configuration

> Integrate MiniSearch with external OpenAI-compatible APIs. Easily configure Base URL, API Key, and Model for seamless inference using simple environment variables.

- Repository: [Victor Nogueira/minisearch](https://github.com/felladrin/minisearch)
- Tags: how-to-guide
- Published: 2026-03-01

---

**MiniSearch integrates with any OpenAI-compatible service by setting three environment variables (base URL, API key, and optional model) and selecting "Internal" as the inference type, which routes requests through a server-side proxy that handles authentication, model selection, and streaming.**

MiniSearch is an open-source AI search interface that supports multiple inference backends. To integrate MiniSearch with external OpenAI-compatible APIs, you configure environment variables that point to your provider's base URL, API key, and preferred model. This setup enables MiniSearch to route generation requests through its internal proxy or connect directly from the client, supporting providers like OpenAI, Azure OpenAI, vLLM, and Ollama.

## Configuration Overview

### Required Environment Variables

MiniSearch requires three environment variables to connect to an external OpenAI-compatible API:

- `INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL` – The base URL of your provider's API endpoint (e.g., `https://llm.mycompany.com/v1` or `https://api.openai.com/v1`).
- `INTERNAL_OPENAI_COMPATIBLE_API_KEY` – Your API authentication key (e.g., `sk-xxxx`).
- `INTERNAL_OPENAI_COMPATIBLE_API_MODEL` – (Optional) The specific model identifier to use. If omitted, MiniSearch automatically fetches available models from the `/v1/models` endpoint and selects one randomly.

### Optional Access Control

To secure the internal inference endpoint, set the `ACCESS_KEYS` environment variable with a comma-separated list of valid tokens:

```bash
ACCESS_KEYS="demo-key-1,demo-key-2"

```

MiniSearch validates the `token` query parameter against this list before processing requests, as implemented in [`server/verifyTokenAndRateLimit.ts`](https://github.com/felladrin/minisearch/blob/main/server/verifyTokenAndRateLimit.ts) and [`server/handleTokenVerification.ts`](https://github.com/felladrin/minisearch/blob/main/server/handleTokenVerification.ts).

## Server-Side Integration Architecture

### Token Validation and Security

When you select "Internal" as the inference type, MiniSearch routes requests through its server-side proxy. The server first validates the access token using [`server/handleTokenVerification.ts`](https://github.com/felladrin/minisearch/blob/main/server/handleTokenVerification.ts), which checks the `token` query parameter against the `ACCESS_KEYS` environment variable. This ensures only authorized clients can consume API resources.

### Provider Initialization

After validation, the server creates an OpenAI-compatible provider using the `@ai-sdk/openai-compatible` package. In [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts) (lines 20-24), the provider is instantiated with:

```typescript
const provider = createOpenAICompatible({
  baseURL: process.env.INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL,
  apiKey: process.env.INTERNAL_OPENAI_COMPATIBLE_API_KEY,
  name: "internal-openai-compatible",
});

```

### Model Selection Logic

MiniSearch handles model selection through [`shared/openaiModels.ts`](https://github.com/felladrin/minisearch/blob/main/shared/openaiModels.ts). If `INTERNAL_OPENAI_COMPATIBLE_API_MODEL` is set, it uses that specific model. Otherwise, it calls `listOpenAiCompatibleModels()` to fetch available models from the `/v1/models` endpoint and uses `selectRandomModel()` to choose one randomly. This ensures resilience if specific models are unavailable.

### Streaming and Retry Mechanisms

The server streams responses using Server-Sent Events (SSE) via the `ai` SDK's `streamText` function. In [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts) (lines 71-86), the implementation:

1. Streams tokens from the external API to the client in real-time.
2. Implements retry logic (lines 56-90) with up to 5 attempts and exponential back-off.
3. Falls back to alternative models if the initial selection fails.

This architecture ensures high availability even when individual models or endpoints experience issues.

## Client-Side Configuration

### Internal Inference Type (Server Proxy)

To use the server-side proxy, set the inference type to `"internal"` in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts). This enum value routes all generation requests to the MiniSearch server's `/inference` endpoint, which then forwards them to the external API configured via environment variables. This approach keeps your API key secure on the server and enables centralized rate limiting and token validation.

### Direct OpenAI Connection (Client-Side)

For direct client-to-API communication, set the inference type to `"openai"`. In this mode, the client uses [`client/modules/textGenerationWithOpenAi.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithOpenAi.ts) to create the provider directly in the browser. You must configure the base URL, API key, and model in the Settings UI. This mode bypasses the server proxy but requires exposing your API key to the client.

## Implementation Examples

### cURL Request to Internal Endpoint

Test your server-side integration using cURL:

```bash
curl -X POST "http://localhost:7861/inference?token=my-demo-key" \
  -H "Content-Type: application/json" \
  -d '{
        "messages": [{ "role": "user", "content": "Summarize the latest news about AI." }],
        "temperature": 0.7,
        "max_tokens": 512
      }' \
  -N

```

The `-N` flag ensures cURL streams SSE chunks as they arrive from the server.

### TypeScript Client Implementation

For custom client implementations, use the `ai` SDK to connect directly:

```typescript
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { streamText } from "ai";

async function generateWithOpenAI(
  messages: { role: "user" | "assistant"; content: string }[],
  baseUrl: string,
  apiKey: string,
  model?: string,
) {
  const provider = createOpenAICompatible({
    baseURL: baseUrl,
    apiKey,
    name: "openai",
  });

  const stream = streamText({
    model: provider.chatModel(model ?? "gpt-4"),
    messages,
    maxOutputTokens: 1024,
    temperature: 0.7,
    maxRetries: 0,
  });

  for await (const part of stream.fullStream) {
    if (part.type === "text-delta") process.stdout.write(part.text);
  }
}

```

### Settings Configuration JSON

When configuring MiniSearch via its settings API or local storage, use this structure:

```json
{
  "inferenceType": "internal",
  "openAiApiBaseUrl": "",
  "openAiApiKey": "",
  "openAiApiModel": "",
  "internalApiBaseUrl": "http://localhost:7861",
  "internalApiModel": "",
  "accessToken": "demo-key-1"
}

```

Set `inferenceType` to `"openai"` for direct client connections, or `"internal"` for server-proxied requests.

## Summary

- **Environment variables drive configuration**: Set `INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL`, `INTERNAL_OPENAI_COMPATIBLE_API_KEY`, and optionally `INTERNAL_OPENAI_COMPATIBLE_API_MODEL` to connect to any OpenAI-compatible provider.

- **Two integration modes exist**: Use **"Internal"** inference type to route requests through MiniSearch's server-side proxy ([`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts)), which validates tokens and handles retries. Use **"OpenAI"** inference type for direct client-to-API communication.

- **Robust error handling**: The server implementation includes automatic model selection from `/v1/models` if no model is specified, and implements up to 5 retry attempts with exponential back-off when models fail.

- **Security through access keys**: The `ACCESS_KEYS` environment variable and [`server/verifyTokenAndRateLimit.ts`](https://github.com/felladrin/minisearch/blob/main/server/verifyTokenAndRateLimit.ts) implementation ensure only authorized users can consume API resources through the internal endpoint.

## Frequently Asked Questions

### What OpenAI-compatible providers work with MiniSearch?

MiniSearch works with any provider implementing the OpenAI API specification, including OpenAI itself, Azure OpenAI, vLLM, Ollama, LocalAI, and custom deployments. The integration uses the `@ai-sdk/openai-compatible` package in [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts) to create a provider instance that communicates with your specified base URL.

### How does MiniSearch handle model selection if I don't specify a model?

When `INTERNAL_OPENAI_COMPATIBLE_API_MODEL` is not set, MiniSearch automatically calls the `/v1/models` endpoint using `listOpenAiCompatibleModels()` from [`shared/openaiModels.ts`](https://github.com/felladrin/minisearch/blob/main/shared/openaiModels.ts) to retrieve available models. It then uses `selectRandomModel()` to choose one randomly, ensuring resilience if specific models are unavailable or rate-limited.

### Can I use MiniSearch without the server-side proxy?

Yes, by setting the `inferenceType` to `"openai"` in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts), MiniSearch bypasses the internal proxy and connects directly from the browser to your API. In this mode, the client uses [`client/modules/textGenerationWithOpenAi.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithOpenAi.ts) to create the provider and stream responses, though this exposes your API key to the client.

### What happens if the external API request fails?

The server implementation in [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts) includes robust error handling with up to 5 retry attempts and exponential back-off. If a specific model fails, the system automatically selects an alternative model from the available list and retries the request, ensuring high availability even during partial outages.