# How to Tune LLM Generation Parameters in PrivateGPT: Temperature, Top-k, Top-p, and Repeat Penalty

> Master LLM generation parameters in PrivateGPT. Learn to tune temperature, max_tokens, top_k, top_p, and repeat_penalty for optimal AI output by editing settings models.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: how-to-guide
- Published: 2026-03-06

---

**You tune LLM generation parameters in PrivateGPT by modifying the Pydantic settings models defined in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py), where `temperature` and `max_new_tokens` reside in the `LLMSettings` class while `top_k`, `top_p`, and `repeat_penalty` live in backend-specific configurations like `LlamaCPPSettings`, all of which are injected into the runtime via `LLMComponent`.**

PrivateGPT centralizes all inference configuration in structured settings models. Whether you are running LlamaCPP, Ollama, or OpenAI backends, understanding how to tune LLM generation parameters ensures you control creativity, output length, and repetition.

## Where Generation Parameters Are Defined

All tunable knobs have default values defined in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) and are injected into the LLM constructor in [`private_gpt/components/llm/llm_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/llm/llm_component.py).

| Parameter | Settings Class | Default | Runtime Injection |
|-----------|----------------|---------|-------------------|
| **temperature** | `LLMSettings` | `0.1` | Passed directly to the LLM implementation in `LLMComponent.__init__` |
| **max_new_tokens** | `LLMSettings` | `256` | Forwarded to the LLM constructor in `LLMComponent` |
| **top_k** | `LlamaCPPSettings` | `40` | Injected into `generate_kwargs` for LlamaCPP and Ollama backends |
| **top_p** | `LlamaCPPSettings` | `0.9` | Injected into `generate_kwargs` alongside `top_k` |
| **repeat_penalty** | `LlamaCPPSettings` | `1.1` | Added to `generate_kwargs` in the component builder |

The `LLMComponent` class merges these dictionaries when building the concrete LLM instance. This ensures that backend-agnostic parameters (`temperature`, `max_new_tokens`) apply universally, while sampling parameters (`top_k`, `top_p`, `repeat_penalty`) customize specific local inference engines.

## Method 1: Static YAML Configuration

The most common way to tune parameters is editing [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml). The file maps directly onto the Pydantic models, using nested blocks for global and backend-specific settings.

### Global Parameters (All Backends)

```yaml
llm:
  mode: llamacpp
  temperature: 0.7        # Increase from 0.1 for more creative output

  max_new_tokens: 512     # Allow longer completions than default 256

  context_window: 3900
  prompt_style: "llama3"

```

### Backend-Specific Sampling Parameters

```yaml
llamacpp:
  llm_hf_repo_id: "lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF"
  llm_hf_model_file: "Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf"
  top_k: 80               # Widen sampling pool from default 40

  top_p: 0.95             # Nucleus sampling threshold

  repeat_penalty: 1.3     # Aggressively discourage repetition loops

```

For Ollama deployments, use the `ollama` block instead:

```yaml
ollama:
  llm_model: "llama3.1"
  top_k: 120
  top_p: 0.96
  repeat_penalty: 1.0

```

After editing [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml), restart the service to load the new configuration.

## Method 2: Runtime Overrides via Environment Variables

PrivateGPT supports dynamic substitution using `${VAR:default}` syntax. This approach is ideal for Docker deployments and CI/CD pipelines.

Set your overrides in the shell:

```bash
export TEMPERATURE=0.9
export MAX_NEW_TOKENS=1024
export TOP_K=100
export TOP_P=0.97
export REPEAT_PENALTY=1.05

```

Reference these variables in [`settings.yaml`](https://github.com/zylon-ai/private-gpt/blob/main/settings.yaml):

```yaml
llm:
  temperature: ${TEMPERATURE:0.1}
  max_new_tokens: ${MAX_NEW_TOKENS:256}

llamacpp:
  top_k: ${TOP_K:40}
  top_p: ${TOP_P:0.9}
  repeat_penalty: ${REPEAT_PENALTY:1.1}

```

When the application initializes, the `Settings` model parses these placeholders and substitutes the runtime values.

## Method 3: Programmatic Tuning

For testing or custom applications, instantiate the `Settings` object directly and inject it into `LLMComponent`.

```python
from private_gpt.settings.settings import Settings, LLMSettings, LlamaCPPSettings
from private_gpt.components.llm.llm_component import LLMComponent

# Construct custom configuration

custom_settings = Settings(
    llm=LLMSettings(
        mode="llamacpp",
        temperature=0.6,
        max_new_tokens=400,
        context_window=3900,
        prompt_style="default",
    ),
    llamacpp=LlamaCPPSettings(
        llm_hf_repo_id="lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF",
        llm_hf_model_file="Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf",
        top_k=60,
        top_p=0.92,
        repeat_penalty=1.2,
    ),
)

# Initialize component with custom settings

llm_component = LLMComponent(custom_settings)
response = llm_component.llm.complete("Explain quantum entanglement in two sentences.")

```

This bypasses YAML parsing entirely, allowing you to tune parameters dynamically based on runtime context.

## Parameter Behavior Reference

Understanding the impact of each parameter helps you tune effectively:

- **temperature**: Controls randomness in token selection. Lower values (0.1) produce deterministic output; higher values (0.9+) increase creativity.
- **max_new_tokens**: Hard limit on the number of tokens generated in the response. Distinct from `context_window`, which controls the input history size.
- **top_k**: Limits the candidate pool to the K most likely next tokens. Higher values allow more diverse vocabulary.
- **top_p**: Enables nucleus sampling, dynamically selecting from the smallest set of tokens whose cumulative probability exceeds P.
- **repeat_penalty**: Penalizes tokens that have already appeared in the generated text. Values above 1.0 discourage repetition; 1.0 disables the feature.

## Summary

- **Centralized Configuration**: All LLM generation parameters are defined in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py) and consumed by [`private_gpt/components/llm/llm_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/llm/llm_component.py).
- **Global vs. Backend-Specific**: Set `temperature` and `max_new_tokens` in the `llm` block; configure `top_k`, `top_p`, and `repeat_penalty` in the `llamacpp` or `ollama` blocks.
- **Flexible Deployment**: Tune parameters via static YAML files, environment variable substitution, or programmatic Python instantiation.
- **Immediate Effect**: Changes require service restart unless using programmatic injection, as the `LLMComponent` builds the LLM instance once during initialization.

## Frequently Asked Questions

### What is the difference between max_new_tokens and context_window in PrivateGPT?

**`max_new_tokens`** limits how many tokens the model generates in a single response (default 256), while **`context_window`** defines the total token budget for the conversation history (typically 3900 or higher). The context window includes system prompts, user messages, and previous assistant responses, whereas max_new_tokens only constrains the new content being generated.

### Why don't top_k, top_p, and repeat_penalty appear in the main llm block?

These parameters are backend-specific sampling strategies. According to the source code in [`private_gpt/settings/settings.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/settings/settings.py), they reside in `LlamaCPPSettings` because they control low-level sampling algorithms specific to local inference engines like LlamaCPP and Ollama. Remote APIs like OpenAI or Azure OpenAI handle sampling server-side, so PrivateGPT only exposes these fields for local backends.

### How do I verify my parameter changes are actually applied?

Check the application logs during startup. The `LLMComponent` in [`private_gpt/components/llm/llm_component.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/components/llm/llm_component.py) logs the construction of the LLM instance. You can also add debug logging to print `settings.llm` and `settings.llamacpp` after loading the configuration. If using environment variables, ensure the `${VAR:default}` syntax matches exactly, including the colon separator.

### Can I tune these parameters per-request instead of globally?

PrivateGPT currently loads generation parameters once during the `LLMComponent` initialization. Per-request tuning would require modifying the component to accept override dictionaries in the `complete()` or `chat()` methods. For now, use programmatic instantiation of `LLMComponent` with different `Settings` objects if you need distinct parameter sets for different request types.