How to Configure Inference Parameters (Temperature, Top‑P, Min‑P, and Penalties) in MiniSearch

MiniSearch exposes inference parameters through client-side settings in client/modules/settings.ts and server-side environment variables in server/config/modelConfig.ts, allowing you to control temperature, top‑P, min‑P, frequency penalty, and presence penalty via the UI, API, or Docker configuration.

MiniSearch is a lightweight search interface that leverages Large Language Models for text generation. Controlling how the model generates text requires tuning inference parameters that affect randomness, diversity, and repetition. This guide explains how to configure temperature, top‑P, min‑P, and penalty values across the MiniSearch client and server architecture.

Client‑Side Configuration

The primary interface for inference tuning lives in client/modules/settings.ts. This file exports defaultSettings, an object that stores user-editable values for all major generation parameters.

Default Parameter Values

The following defaults are defined in defaultSettings (lines 29‑33):

  • inferenceTemperature: 0.7 — Controls randomness; higher values increase creativity.
  • inferenceTopP: 0.9 — Nucleus sampling cutoff; limits token pool to cumulative probability mass.
  • minP: 0.1 — Minimum probability threshold for token consideration.
  • inferenceFrequencyPenalty: 0 — Penalizes repeated tokens based on frequency.
  • inferencePresencePenalty: 0 — Penalizes tokens already present in the text.

How Parameters Flow to the Generation Request

When a user initiates a chat, getDefaultChatCompletionCreateParamsStreaming() in client/modules/textGenerationUtilities.ts (lines 63‑73) assembles the request payload:

export function getDefaultChatCompletionCreateParamsStreaming() {
  const settings = getSettings();
  return {
    stream: true,
    max_tokens: settings.openAiContextLength ?? defaultContextSize,
    temperature: settings.inferenceTemperature,
    top_p: settings.inferenceTopP,
    min_p: settings.minP,
    frequency_penalty: settings.inferenceFrequencyPenalty,
    presence_penalty: settings.inferencePresencePenalty,
  } as const;
}

These parameters are then spread into the OpenAI‑compatible SDK call inside client/modules/textGenerationWithOpenAi.ts (lines 86‑95):

const stream = streamText({
  model: openaiProvider.chatModel(effectiveModel),
  messages,
  maxOutputTokens: params.max_tokens,
  temperature: params.temperature,
  topP: params.top_p,
  frequencyPenalty: params.frequency_penalty,
  presencePenalty: params.presence_penalty,
  // …
});

Any change made through the Settings UI updates defaultSettings via the Pub/Sub system, and the next generation request automatically uses the new values.

Server‑Side Configuration

For deployments using the internal inference endpoint (/inference), MiniSearch provides server‑side overrides via environment variables defined in server/config/modelConfig.ts.

Environment Variable Overrides

The modelConfig object (lines 21‑24) defines the following defaults:

Variable Description Default
MODEL_TEMPERATURE Sampling temperature 0.7
MODEL_TOP_P Nucleus sampling parameter 0.9
MODEL_FREQUENCY_PENALTY Token repetition penalty 0
MODEL_PRESENCE_PENALTY Token presence penalty 0

The getModelConfig() function (lines 51‑62) reads these from process.env, allowing runtime overrides without code changes.

Internal Inference Endpoint

When calling the internal /inference endpoint directly, you can pass parameters in the request body. The handler in server/internalApiEndpointServerHook.ts (lines 24‑31) forwards these to the AI SDK:

const stream = streamText({
  model: openaiProvider.chatModel(model),
  messages: requestBody.messages,
  temperature: requestBody.temperature,
  topP: requestBody.top_p,
  frequencyPenalty: requestBody.frequency_penalty,
  presencePenalty: requestBody.presence_penalty,
  maxOutputTokens: requestBody.max_tokens,
  // …
});

If the request body omits these fields, the server falls back to the environment‑driven modelConfig values.

Practical Examples

Updating Client Settings Programmatically

To adjust inference behavior from client code:

import { defaultSettings } from '@/modules/settings';

// Enable more creative responses
defaultSettings.inferenceTemperature = 1.2;
defaultSettings.inferenceTopP = 0.95;
defaultSettings.minP = 0.05;

// Reduce repetition
defaultSettings.inferenceFrequencyPenalty = 0.2;
defaultSettings.inferencePresencePenalty = 0.1;

Calling the Internal Endpoint with Custom Parameters

curl -X POST "http://localhost:3000/inference?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "messages": [{"role":"user","content":"Explain quantum tunnelling"}],
        "temperature": 1.0,
        "top_p": 0.8,
        "frequency_penalty": 0.1,
        "presence_penalty": 0,
        "max_tokens": 512
      }'

Docker Deployment with Custom Defaults


# docker-compose.yml

services:
  minisearch:
    image: felladrin/minisearch
    environment:
      - MODEL_TEMPERATURE=0.5
      - MODEL_TOP_P=0.7
      - MODEL_FREQUENCY_PENALTY=0.3
      - MODEL_PRESENCE_PENALTY=0.0

Summary

Frequently Asked Questions

What is the default temperature in MiniSearch?

The default temperature is 0.7, defined in client/modules/settings.ts as inferenceTemperature and mirrored by the server‑side MODEL_TEMPERATURE environment variable in server/config/modelConfig.ts.

How do I make MiniSearch responses more deterministic?

Lower the temperature toward 0.0 (e.g., 0.1 or 0.2) and reduce top‑P to 0.5 or lower via the Settings UI, or by setting MODEL_TEMPERATURE=0.1 and MODEL_TOP_P=0.5 as environment variables.

Can I set different penalties for frequency and presence?

Yes. The client exposes inferenceFrequencyPenalty and inferencePresencePenalty in defaultSettings, both defaulting to 0. You can adjust these in the UI or programmatically, and the server respects MODEL_FREQUENCY_PENALTY and MODEL_PRESENCE_PENALTY for internal endpoint calls.

Does the min‑P parameter work with all providers?

The min‑P parameter (minP defaulting to 0.1) is passed through getDefaultChatCompletionCreateParamsStreaming() to the SDK, but actual support depends on the underlying LLM provider. If the provider does not recognize min_p, it will be ignored during inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →