# How to Set Up and Configure the Internal OpenAI-Compatible API in MiniSearch

> Learn how to set up and configure the internal OpenAI-compatible API in MiniSearch. Easily proxy requests to your self-hosted LLM via the inference endpoint.

- Repository: [Victor Nogueira/minisearch](https://github.com/felladrin/minisearch)
- Tags: how-to-guide
- Published: 2026-03-01

---

**To set up the internal OpenAI-compatible API in MiniSearch, configure the `INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL` and related environment variables, which enables the `/inference` endpoint to proxy requests to your self-hosted LLM.**

MiniSearch supports forwarding chat-completion requests to a self-hosted OpenAI-compatible endpoint through its internal API feature. This allows you to serve users with a private LLM server while maintaining the same streaming interface as external providers. The implementation spans environment configuration, server-side proxy logic in [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts), and client-side handling via [`client/modules/textGenerationWithInternalApi.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithInternalApi.ts).

## Environment Variable Configuration

All settings for the internal API reside in a `.env` file at the project root or via Docker environment variables. These variables define the connection to your OpenAI-compatible backend and control UI visibility.

Configure these five key variables:

- **`INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL`** — The base URL of your OpenAI-compatible service (e.g., `https://llm.internal.company.com/v1`)
- **`INTERNAL_OPENAI_COMPATIBLE_API_KEY`** — Authentication key for the internal API (e.g., `sk-internal-xxx`)
- **`INTERNAL_OPENAI_COMPATIBLE_API_MODEL`** — (Optional) Model identifier such as `llama-3.1-8b`; falls back to auto-detection if omitted
- **`INTERNAL_OPENAI_COMPATIBLE_API_NAME`** — Human-readable name displayed in the UI dropdown (e.g., `Company LLM`)
- **`VITE_INTERNAL_API_ENABLED`** — Boolean flag injected by Vite that enables the "internal" option in the front-end settings

Example `.env` configuration:

```bash
INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL="https://llm.internal.company.com/v1"
INTERNAL_OPENAI_COMPATIBLE_API_KEY="sk-internal-xxx"
INTERNAL_OPENAI_COMPATIBLE_API_MODEL="llama-3.1-8b"
INTERNAL_OPENAI_COMPATIBLE_API_NAME="Company LLM"

```

For Docker deployments, expose these variables in [`docker-compose.yml`](https://github.com/felladrin/minisearch/blob/main/docker-compose.yml):

```yaml
environment:
  - INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL=${INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL}
  - INTERNAL_OPENAI_COMPATIBLE_API_KEY=${INTERNAL_OPENAI_COMPATIBLE_API_KEY}
  - INTERNAL_OPENAI_COMPATIBLE_API_MODEL=${INTERNAL_OPENAI_COMPATIBLE_API_MODEL}
  - INTERNAL_OPENAI_COMPATIBLE_API_NAME=${INTERNAL_OPENAI_COMPATIBLE_API_NAME}

```

## Server-Side Proxy Implementation

The server hook **`internalApiEndpointServerHook`** registers middleware that handles the **`/inference`** route. This endpoint validates requests, forwards them to your configured OpenAI-compatible API, and streams responses back via Server-Sent Events (SSE).

The hook performs these operations as implemented in [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts):

1. **Matches** the request path `/inference` and ensures it is a `POST` with `application/json` content type (lines 42-70)
2. **Verifies** the access token via `handleTokenVerification`, which internally calls `verifyTokenAndRateLimit` to enforce rate limiting
3. **Creates** an OpenAI-compatible client using `@ai-sdk/openai-compatible` with `baseURL` and `apiKey` read from `process.env` (lines 72-80)
4. **Streams** the response as SSE using helper functions `sendSseData`, `sendSseDone`, and `sendSseError` (lines 100-125), constructing `ChatCompletionChunk` objects for each token

The server respects the same token-rate-limit logic as other providers, ensuring consistent access control across all inference types.

## Client-Side Integration

When users select **"internal"** in the Settings UI, the front-end sets `settings.inferenceType = "internal"` as defined in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts). The text generation module then routes requests through **[`textGenerationWithInternalApi.ts`](https://github.com/felladrin/minisearch/blob/main/textGenerationWithInternalApi.ts)**.

This client module:

- Builds the request body containing the system prompt and user query
- **POSTs** to `"/inference"` on the same origin, attaching the search token as a query parameter (`token`) (lines 61-74)
- Reads the response as a stream, parsing each `data:` line to extract `choices[0].delta.content` from the SSE payload (lines 84-102)
- Invokes UI callbacks to progressively display the streamed answer

Because the endpoint lives on the **same origin**, the browser handles CORS and SSE headers (`Content-Type: text/event-stream`) automatically without additional configuration.

## Enabling the UI Selection Option

Vite injects the boolean flag **`VITE_INTERNAL_API_ENABLED`** based on the presence of `INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL`. This flag controls whether the "internal" entry appears in the inference-type dropdown.

The injection occurs in [`vite.config.ts`](https://github.com/felladrin/minisearch/blob/main/vite.config.ts) (around line 48), and the front-end reads it through the generated [`client/types.d.ts`](https://github.com/felladrin/minisearch/blob/main/client/types.d.ts) declaration:

```typescript
declare const VITE_INTERNAL_API_ENABLED: boolean;

```

To force the option visible regardless of environment variables, manually set `VITE_INTERNAL_API_ENABLED` in [`vite.config.ts`](https://github.com/felladrin/minisearch/blob/main/vite.config.ts). Otherwise, the UI option only appears when the base URL variable is properly configured.

## Complete Configuration Example

Below is a production-ready configuration connecting MiniSearch to a private Llama 3.1 instance.

**`.env` file:**

```bash
ACCESS_KEYS="my-team-key"
DEFAULT_INFERENCE_TYPE="internal"
INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL="https://llm.internal.company.com/v1"
INTERNAL_OPENAI_COMPATIBLE_API_KEY="sk-internal-xyz"
INTERNAL_OPENAI_COMPATIBLE_API_MODEL="llama-3.1-70b"
INTERNAL_OPENAI_COMPATIBLE_API_NAME="Company LLM"

```

**[`docker-compose.yml`](https://github.com/felladrin/minisearch/blob/main/docker-compose.yml) (development):**

```yaml
services:
  development-server:
    build: .
    ports:
      - "7861:7860"
    environment:
      - ACCESS_KEYS=${ACCESS_KEYS}
      - DEFAULT_INFERENCE_TYPE=${DEFAULT_INFERENCE_TYPE}
      - INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL=${INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL}
      - INTERNAL_OPENAI_COMPATIBLE_API_KEY=${INTERNAL_OPENAI_COMPATIBLE_API_KEY}
      - INTERNAL_OPENAI_COMPATIBLE_API_MODEL=${INTERNAL_OPENAI_COMPATIBLE_API_MODEL}
      - INTERNAL_OPENAI_COMPATIBLE_API_NAME=${INTERNAL_OPENAI_COMPATIBLE_API_NAME}

```

**Front-end usage requires no code changes:**

1. Open the MiniSearch UI and navigate to **Settings → AI Provider**
2. Select the entry named *Company LLM* (the value of `INTERNAL_OPENAI_COMPATIBLE_API_NAME`)
3. Perform a search; the answer streams from your private LLM via the internal proxy at `/inference`

## Summary

- Configure `INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL`, `INTERNAL_OPENAI_COMPATIBLE_API_KEY`, and related variables in your `.env` file to define the connection to your self-hosted LLM.
- The server-side proxy in [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts) handles the `/inference` route, validates tokens via `handleTokenVerification`, and streams SSE responses using `sendSseData`.
- Client-side code in [`client/modules/textGenerationWithInternalApi.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithInternalApi.ts) posts to `/inference` and parses the streaming response to update the UI progressively.
- Enable the UI option by ensuring `INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL` is set, which triggers `VITE_INTERNAL_API_ENABLED` injection in [`vite.config.ts`](https://github.com/felladrin/minisearch/blob/main/vite.config.ts).
- The internal API respects the same token-rate-limiting logic as external providers, ensuring consistent security across all inference types.

## Frequently Asked Questions

### How do I secure the internal API endpoint from unauthorized access?

The `/inference` endpoint uses `handleTokenVerification` (which calls `verifyTokenAndRateLimit`) to validate access tokens before proxying requests. Ensure you define `ACCESS_KEYS` in your environment variables, and the client automatically appends the token as a query parameter when posting to `/inference`.

### Can I use the internal API without enabling it in the Vite configuration?

No. The front-end requires `VITE_INTERNAL_API_ENABLED` to display the "internal" option in the inference-type dropdown. This boolean is automatically injected by [`vite.config.ts`](https://github.com/felladrin/minisearch/blob/main/vite.config.ts) when `INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL` is present, or you can manually set it in the Vite configuration to force visibility.

### What streaming format does the internal API use?

The internal API uses Server-Sent Events (SSE) with the standard OpenAI `ChatCompletionChunk` format. The server sends partial tokens via `sendSseData`, termination signals via `sendSseDone`, and error messages via `sendSseError`, all defined in [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts).

### Does the internal API support model auto-detection?

Yes. While you can specify a model via `INTERNAL_OPENAI_COMPATIBLE_API_MODEL`, the system falls back to auto-detection if this variable is omitted. The client module [`textGenerationWithInternalApi.ts`](https://github.com/felladrin/minisearch/blob/main/textGenerationWithInternalApi.ts) handles the request construction regardless of whether the model is explicitly defined.