How to Tune LLM Generation Parameters in PrivateGPT: Temperature, Top-k, Top-p, and Repeat Penalty

You tune LLM generation parameters in PrivateGPT by modifying the Pydantic settings models defined in private_gpt/settings/settings.py, where temperature and max_new_tokens reside in the LLMSettings class while top_k, top_p, and repeat_penalty live in backend-specific configurations like LlamaCPPSettings, all of which are injected into the runtime via LLMComponent.

PrivateGPT centralizes all inference configuration in structured settings models. Whether you are running LlamaCPP, Ollama, or OpenAI backends, understanding how to tune LLM generation parameters ensures you control creativity, output length, and repetition.

Where Generation Parameters Are Defined

All tunable knobs have default values defined in private_gpt/settings/settings.py and are injected into the LLM constructor in private_gpt/components/llm/llm_component.py.

Parameter Settings Class Default Runtime Injection
temperature LLMSettings 0.1 Passed directly to the LLM implementation in LLMComponent.__init__
max_new_tokens LLMSettings 256 Forwarded to the LLM constructor in LLMComponent
top_k LlamaCPPSettings 40 Injected into generate_kwargs for LlamaCPP and Ollama backends
top_p LlamaCPPSettings 0.9 Injected into generate_kwargs alongside top_k
repeat_penalty LlamaCPPSettings 1.1 Added to generate_kwargs in the component builder

The LLMComponent class merges these dictionaries when building the concrete LLM instance. This ensures that backend-agnostic parameters (temperature, max_new_tokens) apply universally, while sampling parameters (top_k, top_p, repeat_penalty) customize specific local inference engines.

Method 1: Static YAML Configuration

The most common way to tune parameters is editing settings.yaml. The file maps directly onto the Pydantic models, using nested blocks for global and backend-specific settings.

Global Parameters (All Backends)

llm:
  mode: llamacpp
  temperature: 0.7        # Increase from 0.1 for more creative output

  max_new_tokens: 512     # Allow longer completions than default 256

  context_window: 3900
  prompt_style: "llama3"

Backend-Specific Sampling Parameters

llamacpp:
  llm_hf_repo_id: "lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF"
  llm_hf_model_file: "Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf"
  top_k: 80               # Widen sampling pool from default 40

  top_p: 0.95             # Nucleus sampling threshold

  repeat_penalty: 1.3     # Aggressively discourage repetition loops

For Ollama deployments, use the ollama block instead:

ollama:
  llm_model: "llama3.1"
  top_k: 120
  top_p: 0.96
  repeat_penalty: 1.0

After editing settings.yaml, restart the service to load the new configuration.

Method 2: Runtime Overrides via Environment Variables

PrivateGPT supports dynamic substitution using ${VAR:default} syntax. This approach is ideal for Docker deployments and CI/CD pipelines.

Set your overrides in the shell:

export TEMPERATURE=0.9
export MAX_NEW_TOKENS=1024
export TOP_K=100
export TOP_P=0.97
export REPEAT_PENALTY=1.05

Reference these variables in settings.yaml:

llm:
  temperature: ${TEMPERATURE:0.1}
  max_new_tokens: ${MAX_NEW_TOKENS:256}

llamacpp:
  top_k: ${TOP_K:40}
  top_p: ${TOP_P:0.9}
  repeat_penalty: ${REPEAT_PENALTY:1.1}

When the application initializes, the Settings model parses these placeholders and substitutes the runtime values.

Method 3: Programmatic Tuning

For testing or custom applications, instantiate the Settings object directly and inject it into LLMComponent.

from private_gpt.settings.settings import Settings, LLMSettings, LlamaCPPSettings
from private_gpt.components.llm.llm_component import LLMComponent

# Construct custom configuration

custom_settings = Settings(
    llm=LLMSettings(
        mode="llamacpp",
        temperature=0.6,
        max_new_tokens=400,
        context_window=3900,
        prompt_style="default",
    ),
    llamacpp=LlamaCPPSettings(
        llm_hf_repo_id="lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF",
        llm_hf_model_file="Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf",
        top_k=60,
        top_p=0.92,
        repeat_penalty=1.2,
    ),
)

# Initialize component with custom settings

llm_component = LLMComponent(custom_settings)
response = llm_component.llm.complete("Explain quantum entanglement in two sentences.")

This bypasses YAML parsing entirely, allowing you to tune parameters dynamically based on runtime context.

Parameter Behavior Reference

Understanding the impact of each parameter helps you tune effectively:

  • temperature: Controls randomness in token selection. Lower values (0.1) produce deterministic output; higher values (0.9+) increase creativity.
  • max_new_tokens: Hard limit on the number of tokens generated in the response. Distinct from context_window, which controls the input history size.
  • top_k: Limits the candidate pool to the K most likely next tokens. Higher values allow more diverse vocabulary.
  • top_p: Enables nucleus sampling, dynamically selecting from the smallest set of tokens whose cumulative probability exceeds P.
  • repeat_penalty: Penalizes tokens that have already appeared in the generated text. Values above 1.0 discourage repetition; 1.0 disables the feature.

Summary

  • Centralized Configuration: All LLM generation parameters are defined in private_gpt/settings/settings.py and consumed by private_gpt/components/llm/llm_component.py.
  • Global vs. Backend-Specific: Set temperature and max_new_tokens in the llm block; configure top_k, top_p, and repeat_penalty in the llamacpp or ollama blocks.
  • Flexible Deployment: Tune parameters via static YAML files, environment variable substitution, or programmatic Python instantiation.
  • Immediate Effect: Changes require service restart unless using programmatic injection, as the LLMComponent builds the LLM instance once during initialization.

Frequently Asked Questions

What is the difference between max_new_tokens and context_window in PrivateGPT?

max_new_tokens limits how many tokens the model generates in a single response (default 256), while context_window defines the total token budget for the conversation history (typically 3900 or higher). The context window includes system prompts, user messages, and previous assistant responses, whereas max_new_tokens only constrains the new content being generated.

Why don't top_k, top_p, and repeat_penalty appear in the main llm block?

These parameters are backend-specific sampling strategies. According to the source code in private_gpt/settings/settings.py, they reside in LlamaCPPSettings because they control low-level sampling algorithms specific to local inference engines like LlamaCPP and Ollama. Remote APIs like OpenAI or Azure OpenAI handle sampling server-side, so PrivateGPT only exposes these fields for local backends.

How do I verify my parameter changes are actually applied?

Check the application logs during startup. The LLMComponent in private_gpt/components/llm/llm_component.py logs the construction of the LLM instance. You can also add debug logging to print settings.llm and settings.llamacpp after loading the configuration. If using environment variables, ensure the ${VAR:default} syntax matches exactly, including the colon separator.

Can I tune these parameters per-request instead of globally?

PrivateGPT currently loads generation parameters once during the LLMComponent initialization. Per-request tuning would require modifying the component to accept override dictionaries in the complete() or chat() methods. For now, use programmatic instantiation of LLMComponent with different Settings objects if you need distinct parameter sets for different request types.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →