# How to Configure Inference Parameters (Temperature, Top‑P, Min‑P, and Penalties) in MiniSearch

> Learn to configure MiniSearch inference parameters like temperature, topP, minP, and penalties. Control generation settings via UI, API, or Docker for optimal results.

- Repository: [Victor Nogueira/minisearch](https://github.com/felladrin/minisearch)
- Tags: how-to-guide
- Published: 2026-03-01

---

**MiniSearch exposes inference parameters through client-side settings in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts) and server-side environment variables in [`server/config/modelConfig.ts`](https://github.com/felladrin/minisearch/blob/main/server/config/modelConfig.ts), allowing you to control temperature, top‑P, min‑P, frequency penalty, and presence penalty via the UI, API, or Docker configuration.**

MiniSearch is a lightweight search interface that leverages Large Language Models for text generation. Controlling how the model generates text requires tuning inference parameters that affect randomness, diversity, and repetition. This guide explains how to configure temperature, top‑P, min‑P, and penalty values across the MiniSearch client and server architecture.

## Client‑Side Configuration

The primary interface for inference tuning lives in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts). This file exports `defaultSettings`, an object that stores user-editable values for all major generation parameters.

### Default Parameter Values

The following defaults are defined in `defaultSettings` (lines 29‑33):

- `inferenceTemperature`: **0.7** — Controls randomness; higher values increase creativity.
- `inferenceTopP`: **0.9** — Nucleus sampling cutoff; limits token pool to cumulative probability mass.
- `minP`: **0.1** — Minimum probability threshold for token consideration.
- `inferenceFrequencyPenalty`: **0** — Penalizes repeated tokens based on frequency.
- `inferencePresencePenalty`: **0** — Penalizes tokens already present in the text.

### How Parameters Flow to the Generation Request

When a user initiates a chat, `getDefaultChatCompletionCreateParamsStreaming()` in [`client/modules/textGenerationUtilities.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationUtilities.ts) (lines 63‑73) assembles the request payload:

```typescript
export function getDefaultChatCompletionCreateParamsStreaming() {
  const settings = getSettings();
  return {
    stream: true,
    max_tokens: settings.openAiContextLength ?? defaultContextSize,
    temperature: settings.inferenceTemperature,
    top_p: settings.inferenceTopP,
    min_p: settings.minP,
    frequency_penalty: settings.inferenceFrequencyPenalty,
    presence_penalty: settings.inferencePresencePenalty,
  } as const;
}

```

These parameters are then spread into the OpenAI‑compatible SDK call inside [`client/modules/textGenerationWithOpenAi.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationWithOpenAi.ts) (lines 86‑95):

```typescript
const stream = streamText({
  model: openaiProvider.chatModel(effectiveModel),
  messages,
  maxOutputTokens: params.max_tokens,
  temperature: params.temperature,
  topP: params.top_p,
  frequencyPenalty: params.frequency_penalty,
  presencePenalty: params.presence_penalty,
  // …
});

```

Any change made through the Settings UI updates `defaultSettings` via the Pub/Sub system, and the next generation request automatically uses the new values.

## Server‑Side Configuration

For deployments using the internal inference endpoint (`/inference`), MiniSearch provides server‑side overrides via environment variables defined in [`server/config/modelConfig.ts`](https://github.com/felladrin/minisearch/blob/main/server/config/modelConfig.ts).

### Environment Variable Overrides

The `modelConfig` object (lines 21‑24) defines the following defaults:

| Variable | Description | Default |
|----------|-------------|---------|
| `MODEL_TEMPERATURE` | Sampling temperature | **0.7** |
| `MODEL_TOP_P` | Nucleus sampling parameter | **0.9** |
| `MODEL_FREQUENCY_PENALTY` | Token repetition penalty | **0** |
| `MODEL_PRESENCE_PENALTY` | Token presence penalty | **0** |

The `getModelConfig()` function (lines 51‑62) reads these from `process.env`, allowing runtime overrides without code changes.

### Internal Inference Endpoint

When calling the internal `/inference` endpoint directly, you can pass parameters in the request body. The handler in [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts) (lines 24‑31) forwards these to the AI SDK:

```typescript
const stream = streamText({
  model: openaiProvider.chatModel(model),
  messages: requestBody.messages,
  temperature: requestBody.temperature,
  topP: requestBody.top_p,
  frequencyPenalty: requestBody.frequency_penalty,
  presencePenalty: requestBody.presence_penalty,
  maxOutputTokens: requestBody.max_tokens,
  // …
});

```

If the request body omits these fields, the server falls back to the environment‑driven `modelConfig` values.

## Practical Examples

### Updating Client Settings Programmatically

To adjust inference behavior from client code:

```typescript
import { defaultSettings } from '@/modules/settings';

// Enable more creative responses
defaultSettings.inferenceTemperature = 1.2;
defaultSettings.inferenceTopP = 0.95;
defaultSettings.minP = 0.05;

// Reduce repetition
defaultSettings.inferenceFrequencyPenalty = 0.2;
defaultSettings.inferencePresencePenalty = 0.1;

```

### Calling the Internal Endpoint with Custom Parameters

```bash
curl -X POST "http://localhost:3000/inference?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "messages": [{"role":"user","content":"Explain quantum tunnelling"}],
        "temperature": 1.0,
        "top_p": 0.8,
        "frequency_penalty": 0.1,
        "presence_penalty": 0,
        "max_tokens": 512
      }'

```

### Docker Deployment with Custom Defaults

```yaml

# docker-compose.yml

services:
  minisearch:
    image: felladrin/minisearch
    environment:
      - MODEL_TEMPERATURE=0.5
      - MODEL_TOP_P=0.7
      - MODEL_FREQUENCY_PENALTY=0.3
      - MODEL_PRESENCE_PENALTY=0.0

```

## Summary

- **Client settings** in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts) provide user‑editable defaults for `inferenceTemperature`, `inferenceTopP`, `minP`, `inferenceFrequencyPenalty`, and `inferencePresencePenalty`.
- **Parameter propagation** occurs through `getDefaultChatCompletionCreateParamsStreaming()` in [`client/modules/textGenerationUtilities.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/textGenerationUtilities.ts), which feeds into [`textGenerationWithOpenAi.ts`](https://github.com/felladrin/minisearch/blob/main/textGenerationWithOpenAi.ts).
- **Server overrides** are available via environment variables (`MODEL_TEMPERATURE`, `MODEL_TOP_P`, `MODEL_FREQUENCY_PENALTY`, `MODEL_PRESENCE_PENALTY`) defined in [`server/config/modelConfig.ts`](https://github.com/felladrin/minisearch/blob/main/server/config/modelConfig.ts).
- **Direct API control** is possible through the `/inference` endpoint handled in [`server/internalApiEndpointServerHook.ts`](https://github.com/felladrin/minisearch/blob/main/server/internalApiEndpointServerHook.ts), accepting explicit `temperature`, `top_p`, and penalty fields.

## Frequently Asked Questions

### What is the default temperature in MiniSearch?

The default **temperature** is **0.7**, defined in [`client/modules/settings.ts`](https://github.com/felladrin/minisearch/blob/main/client/modules/settings.ts) as `inferenceTemperature` and mirrored by the server‑side `MODEL_TEMPERATURE` environment variable in [`server/config/modelConfig.ts`](https://github.com/felladrin/minisearch/blob/main/server/config/modelConfig.ts).

### How do I make MiniSearch responses more deterministic?

Lower the **temperature** toward **0.0** (e.g., 0.1 or 0.2) and reduce **top‑P** to **0.5** or lower via the Settings UI, or by setting `MODEL_TEMPERATURE=0.1` and `MODEL_TOP_P=0.5` as environment variables.

### Can I set different penalties for frequency and presence?

Yes. The client exposes `inferenceFrequencyPenalty` and `inferencePresencePenalty` in `defaultSettings`, both defaulting to **0**. You can adjust these in the UI or programmatically, and the server respects `MODEL_FREQUENCY_PENALTY` and `MODEL_PRESENCE_PENALTY` for internal endpoint calls.

### Does the min‑P parameter work with all providers?

The **min‑P** parameter (`minP` defaulting to **0.1**) is passed through `getDefaultChatCompletionCreateParamsStreaming()` to the SDK, but actual support depends on the underlying LLM provider. If the provider does not recognize `min_p`, it will be ignored during inference.