How to Configure Inference Parameters (Temperature, Top‑P, Min‑P, and Penalties) in MiniSearch
MiniSearch exposes inference parameters through client-side settings in client/modules/settings.ts and server-side environment variables in server/config/modelConfig.ts, allowing you to control temperature, top‑P, min‑P, frequency penalty, and presence penalty via the UI, API, or Docker configuration.
MiniSearch is a lightweight search interface that leverages Large Language Models for text generation. Controlling how the model generates text requires tuning inference parameters that affect randomness, diversity, and repetition. This guide explains how to configure temperature, top‑P, min‑P, and penalty values across the MiniSearch client and server architecture.
Client‑Side Configuration
The primary interface for inference tuning lives in client/modules/settings.ts. This file exports defaultSettings, an object that stores user-editable values for all major generation parameters.
Default Parameter Values
The following defaults are defined in defaultSettings (lines 29‑33):
inferenceTemperature: 0.7 — Controls randomness; higher values increase creativity.inferenceTopP: 0.9 — Nucleus sampling cutoff; limits token pool to cumulative probability mass.minP: 0.1 — Minimum probability threshold for token consideration.inferenceFrequencyPenalty: 0 — Penalizes repeated tokens based on frequency.inferencePresencePenalty: 0 — Penalizes tokens already present in the text.
How Parameters Flow to the Generation Request
When a user initiates a chat, getDefaultChatCompletionCreateParamsStreaming() in client/modules/textGenerationUtilities.ts (lines 63‑73) assembles the request payload:
export function getDefaultChatCompletionCreateParamsStreaming() {
const settings = getSettings();
return {
stream: true,
max_tokens: settings.openAiContextLength ?? defaultContextSize,
temperature: settings.inferenceTemperature,
top_p: settings.inferenceTopP,
min_p: settings.minP,
frequency_penalty: settings.inferenceFrequencyPenalty,
presence_penalty: settings.inferencePresencePenalty,
} as const;
}
These parameters are then spread into the OpenAI‑compatible SDK call inside client/modules/textGenerationWithOpenAi.ts (lines 86‑95):
const stream = streamText({
model: openaiProvider.chatModel(effectiveModel),
messages,
maxOutputTokens: params.max_tokens,
temperature: params.temperature,
topP: params.top_p,
frequencyPenalty: params.frequency_penalty,
presencePenalty: params.presence_penalty,
// …
});
Any change made through the Settings UI updates defaultSettings via the Pub/Sub system, and the next generation request automatically uses the new values.
Server‑Side Configuration
For deployments using the internal inference endpoint (/inference), MiniSearch provides server‑side overrides via environment variables defined in server/config/modelConfig.ts.
Environment Variable Overrides
The modelConfig object (lines 21‑24) defines the following defaults:
| Variable | Description | Default |
|---|---|---|
MODEL_TEMPERATURE |
Sampling temperature | 0.7 |
MODEL_TOP_P |
Nucleus sampling parameter | 0.9 |
MODEL_FREQUENCY_PENALTY |
Token repetition penalty | 0 |
MODEL_PRESENCE_PENALTY |
Token presence penalty | 0 |
The getModelConfig() function (lines 51‑62) reads these from process.env, allowing runtime overrides without code changes.
Internal Inference Endpoint
When calling the internal /inference endpoint directly, you can pass parameters in the request body. The handler in server/internalApiEndpointServerHook.ts (lines 24‑31) forwards these to the AI SDK:
const stream = streamText({
model: openaiProvider.chatModel(model),
messages: requestBody.messages,
temperature: requestBody.temperature,
topP: requestBody.top_p,
frequencyPenalty: requestBody.frequency_penalty,
presencePenalty: requestBody.presence_penalty,
maxOutputTokens: requestBody.max_tokens,
// …
});
If the request body omits these fields, the server falls back to the environment‑driven modelConfig values.
Practical Examples
Updating Client Settings Programmatically
To adjust inference behavior from client code:
import { defaultSettings } from '@/modules/settings';
// Enable more creative responses
defaultSettings.inferenceTemperature = 1.2;
defaultSettings.inferenceTopP = 0.95;
defaultSettings.minP = 0.05;
// Reduce repetition
defaultSettings.inferenceFrequencyPenalty = 0.2;
defaultSettings.inferencePresencePenalty = 0.1;
Calling the Internal Endpoint with Custom Parameters
curl -X POST "http://localhost:3000/inference?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role":"user","content":"Explain quantum tunnelling"}],
"temperature": 1.0,
"top_p": 0.8,
"frequency_penalty": 0.1,
"presence_penalty": 0,
"max_tokens": 512
}'
Docker Deployment with Custom Defaults
# docker-compose.yml
services:
minisearch:
image: felladrin/minisearch
environment:
- MODEL_TEMPERATURE=0.5
- MODEL_TOP_P=0.7
- MODEL_FREQUENCY_PENALTY=0.3
- MODEL_PRESENCE_PENALTY=0.0
Summary
- Client settings in
client/modules/settings.tsprovide user‑editable defaults forinferenceTemperature,inferenceTopP,minP,inferenceFrequencyPenalty, andinferencePresencePenalty. - Parameter propagation occurs through
getDefaultChatCompletionCreateParamsStreaming()inclient/modules/textGenerationUtilities.ts, which feeds intotextGenerationWithOpenAi.ts. - Server overrides are available via environment variables (
MODEL_TEMPERATURE,MODEL_TOP_P,MODEL_FREQUENCY_PENALTY,MODEL_PRESENCE_PENALTY) defined inserver/config/modelConfig.ts. - Direct API control is possible through the
/inferenceendpoint handled inserver/internalApiEndpointServerHook.ts, accepting explicittemperature,top_p, and penalty fields.
Frequently Asked Questions
What is the default temperature in MiniSearch?
The default temperature is 0.7, defined in client/modules/settings.ts as inferenceTemperature and mirrored by the server‑side MODEL_TEMPERATURE environment variable in server/config/modelConfig.ts.
How do I make MiniSearch responses more deterministic?
Lower the temperature toward 0.0 (e.g., 0.1 or 0.2) and reduce top‑P to 0.5 or lower via the Settings UI, or by setting MODEL_TEMPERATURE=0.1 and MODEL_TOP_P=0.5 as environment variables.
Can I set different penalties for frequency and presence?
Yes. The client exposes inferenceFrequencyPenalty and inferencePresencePenalty in defaultSettings, both defaulting to 0. You can adjust these in the UI or programmatically, and the server respects MODEL_FREQUENCY_PENALTY and MODEL_PRESENCE_PENALTY for internal endpoint calls.
Does the min‑P parameter work with all providers?
The min‑P parameter (minP defaulting to 0.1) is passed through getDefaultChatCompletionCreateParamsStreaming() to the SDK, but actual support depends on the underlying LLM provider. If the provider does not recognize min_p, it will be ignored during inference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →