How to Set Up and Configure the Internal OpenAI-Compatible API in MiniSearch
To set up the internal OpenAI-compatible API in MiniSearch, configure the INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL and related environment variables, which enables the /inference endpoint to proxy requests to your self-hosted LLM.
MiniSearch supports forwarding chat-completion requests to a self-hosted OpenAI-compatible endpoint through its internal API feature. This allows you to serve users with a private LLM server while maintaining the same streaming interface as external providers. The implementation spans environment configuration, server-side proxy logic in server/internalApiEndpointServerHook.ts, and client-side handling via client/modules/textGenerationWithInternalApi.ts.
Environment Variable Configuration
All settings for the internal API reside in a .env file at the project root or via Docker environment variables. These variables define the connection to your OpenAI-compatible backend and control UI visibility.
Configure these five key variables:
INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL— The base URL of your OpenAI-compatible service (e.g.,https://llm.internal.company.com/v1)INTERNAL_OPENAI_COMPATIBLE_API_KEY— Authentication key for the internal API (e.g.,sk-internal-xxx)INTERNAL_OPENAI_COMPATIBLE_API_MODEL— (Optional) Model identifier such asllama-3.1-8b; falls back to auto-detection if omittedINTERNAL_OPENAI_COMPATIBLE_API_NAME— Human-readable name displayed in the UI dropdown (e.g.,Company LLM)VITE_INTERNAL_API_ENABLED— Boolean flag injected by Vite that enables the "internal" option in the front-end settings
Example .env configuration:
INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL="https://llm.internal.company.com/v1"
INTERNAL_OPENAI_COMPATIBLE_API_KEY="sk-internal-xxx"
INTERNAL_OPENAI_COMPATIBLE_API_MODEL="llama-3.1-8b"
INTERNAL_OPENAI_COMPATIBLE_API_NAME="Company LLM"
For Docker deployments, expose these variables in docker-compose.yml:
environment:
- INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL=${INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL}
- INTERNAL_OPENAI_COMPATIBLE_API_KEY=${INTERNAL_OPENAI_COMPATIBLE_API_KEY}
- INTERNAL_OPENAI_COMPATIBLE_API_MODEL=${INTERNAL_OPENAI_COMPATIBLE_API_MODEL}
- INTERNAL_OPENAI_COMPATIBLE_API_NAME=${INTERNAL_OPENAI_COMPATIBLE_API_NAME}
Server-Side Proxy Implementation
The server hook internalApiEndpointServerHook registers middleware that handles the /inference route. This endpoint validates requests, forwards them to your configured OpenAI-compatible API, and streams responses back via Server-Sent Events (SSE).
The hook performs these operations as implemented in server/internalApiEndpointServerHook.ts:
- Matches the request path
/inferenceand ensures it is aPOSTwithapplication/jsoncontent type (lines 42-70) - Verifies the access token via
handleTokenVerification, which internally callsverifyTokenAndRateLimitto enforce rate limiting - Creates an OpenAI-compatible client using
@ai-sdk/openai-compatiblewithbaseURLandapiKeyread fromprocess.env(lines 72-80) - Streams the response as SSE using helper functions
sendSseData,sendSseDone, andsendSseError(lines 100-125), constructingChatCompletionChunkobjects for each token
The server respects the same token-rate-limit logic as other providers, ensuring consistent access control across all inference types.
Client-Side Integration
When users select "internal" in the Settings UI, the front-end sets settings.inferenceType = "internal" as defined in client/modules/settings.ts. The text generation module then routes requests through textGenerationWithInternalApi.ts.
This client module:
- Builds the request body containing the system prompt and user query
- POSTs to
"/inference"on the same origin, attaching the search token as a query parameter (token) (lines 61-74) - Reads the response as a stream, parsing each
data:line to extractchoices[0].delta.contentfrom the SSE payload (lines 84-102) - Invokes UI callbacks to progressively display the streamed answer
Because the endpoint lives on the same origin, the browser handles CORS and SSE headers (Content-Type: text/event-stream) automatically without additional configuration.
Enabling the UI Selection Option
Vite injects the boolean flag VITE_INTERNAL_API_ENABLED based on the presence of INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL. This flag controls whether the "internal" entry appears in the inference-type dropdown.
The injection occurs in vite.config.ts (around line 48), and the front-end reads it through the generated client/types.d.ts declaration:
declare const VITE_INTERNAL_API_ENABLED: boolean;
To force the option visible regardless of environment variables, manually set VITE_INTERNAL_API_ENABLED in vite.config.ts. Otherwise, the UI option only appears when the base URL variable is properly configured.
Complete Configuration Example
Below is a production-ready configuration connecting MiniSearch to a private Llama 3.1 instance.
.env file:
ACCESS_KEYS="my-team-key"
DEFAULT_INFERENCE_TYPE="internal"
INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL="https://llm.internal.company.com/v1"
INTERNAL_OPENAI_COMPATIBLE_API_KEY="sk-internal-xyz"
INTERNAL_OPENAI_COMPATIBLE_API_MODEL="llama-3.1-70b"
INTERNAL_OPENAI_COMPATIBLE_API_NAME="Company LLM"
docker-compose.yml (development):
services:
development-server:
build: .
ports:
- "7861:7860"
environment:
- ACCESS_KEYS=${ACCESS_KEYS}
- DEFAULT_INFERENCE_TYPE=${DEFAULT_INFERENCE_TYPE}
- INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL=${INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL}
- INTERNAL_OPENAI_COMPATIBLE_API_KEY=${INTERNAL_OPENAI_COMPATIBLE_API_KEY}
- INTERNAL_OPENAI_COMPATIBLE_API_MODEL=${INTERNAL_OPENAI_COMPATIBLE_API_MODEL}
- INTERNAL_OPENAI_COMPATIBLE_API_NAME=${INTERNAL_OPENAI_COMPATIBLE_API_NAME}
Front-end usage requires no code changes:
- Open the MiniSearch UI and navigate to Settings → AI Provider
- Select the entry named Company LLM (the value of
INTERNAL_OPENAI_COMPATIBLE_API_NAME) - Perform a search; the answer streams from your private LLM via the internal proxy at
/inference
Summary
- Configure
INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL,INTERNAL_OPENAI_COMPATIBLE_API_KEY, and related variables in your.envfile to define the connection to your self-hosted LLM. - The server-side proxy in
server/internalApiEndpointServerHook.tshandles the/inferenceroute, validates tokens viahandleTokenVerification, and streams SSE responses usingsendSseData. - Client-side code in
client/modules/textGenerationWithInternalApi.tsposts to/inferenceand parses the streaming response to update the UI progressively. - Enable the UI option by ensuring
INTERNAL_OPENAI_COMPATIBLE_API_BASE_URLis set, which triggersVITE_INTERNAL_API_ENABLEDinjection invite.config.ts. - The internal API respects the same token-rate-limiting logic as external providers, ensuring consistent security across all inference types.
Frequently Asked Questions
How do I secure the internal API endpoint from unauthorized access?
The /inference endpoint uses handleTokenVerification (which calls verifyTokenAndRateLimit) to validate access tokens before proxying requests. Ensure you define ACCESS_KEYS in your environment variables, and the client automatically appends the token as a query parameter when posting to /inference.
Can I use the internal API without enabling it in the Vite configuration?
No. The front-end requires VITE_INTERNAL_API_ENABLED to display the "internal" option in the inference-type dropdown. This boolean is automatically injected by vite.config.ts when INTERNAL_OPENAI_COMPATIBLE_API_BASE_URL is present, or you can manually set it in the Vite configuration to force visibility.
What streaming format does the internal API use?
The internal API uses Server-Sent Events (SSE) with the standard OpenAI ChatCompletionChunk format. The server sends partial tokens via sendSseData, termination signals via sendSseDone, and error messages via sendSseError, all defined in server/internalApiEndpointServerHook.ts.
Does the internal API support model auto-detection?
Yes. While you can specify a model via INTERNAL_OPENAI_COMPATIBLE_API_MODEL, the system falls back to auto-detection if this variable is omitted. The client module textGenerationWithInternalApi.ts handles the request construction regardless of whether the model is explicitly defined.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →