How to Tune LLM Generation Parameters in PrivateGPT: Temperature, Top-k, Top-p, and Repeat Penalty
You tune LLM generation parameters in PrivateGPT by modifying the Pydantic settings models defined in private_gpt/settings/settings.py, where temperature and max_new_tokens reside in the LLMSettings class while top_k, top_p, and repeat_penalty live in backend-specific configurations like LlamaCPPSettings, all of which are injected into the runtime via LLMComponent.
PrivateGPT centralizes all inference configuration in structured settings models. Whether you are running LlamaCPP, Ollama, or OpenAI backends, understanding how to tune LLM generation parameters ensures you control creativity, output length, and repetition.
Where Generation Parameters Are Defined
All tunable knobs have default values defined in private_gpt/settings/settings.py and are injected into the LLM constructor in private_gpt/components/llm/llm_component.py.
| Parameter | Settings Class | Default | Runtime Injection |
|---|---|---|---|
| temperature | LLMSettings |
0.1 |
Passed directly to the LLM implementation in LLMComponent.__init__ |
| max_new_tokens | LLMSettings |
256 |
Forwarded to the LLM constructor in LLMComponent |
| top_k | LlamaCPPSettings |
40 |
Injected into generate_kwargs for LlamaCPP and Ollama backends |
| top_p | LlamaCPPSettings |
0.9 |
Injected into generate_kwargs alongside top_k |
| repeat_penalty | LlamaCPPSettings |
1.1 |
Added to generate_kwargs in the component builder |
The LLMComponent class merges these dictionaries when building the concrete LLM instance. This ensures that backend-agnostic parameters (temperature, max_new_tokens) apply universally, while sampling parameters (top_k, top_p, repeat_penalty) customize specific local inference engines.
Method 1: Static YAML Configuration
The most common way to tune parameters is editing settings.yaml. The file maps directly onto the Pydantic models, using nested blocks for global and backend-specific settings.
Global Parameters (All Backends)
llm:
mode: llamacpp
temperature: 0.7 # Increase from 0.1 for more creative output
max_new_tokens: 512 # Allow longer completions than default 256
context_window: 3900
prompt_style: "llama3"
Backend-Specific Sampling Parameters
llamacpp:
llm_hf_repo_id: "lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF"
llm_hf_model_file: "Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf"
top_k: 80 # Widen sampling pool from default 40
top_p: 0.95 # Nucleus sampling threshold
repeat_penalty: 1.3 # Aggressively discourage repetition loops
For Ollama deployments, use the ollama block instead:
ollama:
llm_model: "llama3.1"
top_k: 120
top_p: 0.96
repeat_penalty: 1.0
After editing settings.yaml, restart the service to load the new configuration.
Method 2: Runtime Overrides via Environment Variables
PrivateGPT supports dynamic substitution using ${VAR:default} syntax. This approach is ideal for Docker deployments and CI/CD pipelines.
Set your overrides in the shell:
export TEMPERATURE=0.9
export MAX_NEW_TOKENS=1024
export TOP_K=100
export TOP_P=0.97
export REPEAT_PENALTY=1.05
Reference these variables in settings.yaml:
llm:
temperature: ${TEMPERATURE:0.1}
max_new_tokens: ${MAX_NEW_TOKENS:256}
llamacpp:
top_k: ${TOP_K:40}
top_p: ${TOP_P:0.9}
repeat_penalty: ${REPEAT_PENALTY:1.1}
When the application initializes, the Settings model parses these placeholders and substitutes the runtime values.
Method 3: Programmatic Tuning
For testing or custom applications, instantiate the Settings object directly and inject it into LLMComponent.
from private_gpt.settings.settings import Settings, LLMSettings, LlamaCPPSettings
from private_gpt.components.llm.llm_component import LLMComponent
# Construct custom configuration
custom_settings = Settings(
llm=LLMSettings(
mode="llamacpp",
temperature=0.6,
max_new_tokens=400,
context_window=3900,
prompt_style="default",
),
llamacpp=LlamaCPPSettings(
llm_hf_repo_id="lmstudio-community/Meta-Llama-3.1-8B-Instruct-GGUF",
llm_hf_model_file="Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf",
top_k=60,
top_p=0.92,
repeat_penalty=1.2,
),
)
# Initialize component with custom settings
llm_component = LLMComponent(custom_settings)
response = llm_component.llm.complete("Explain quantum entanglement in two sentences.")
This bypasses YAML parsing entirely, allowing you to tune parameters dynamically based on runtime context.
Parameter Behavior Reference
Understanding the impact of each parameter helps you tune effectively:
- temperature: Controls randomness in token selection. Lower values (0.1) produce deterministic output; higher values (0.9+) increase creativity.
- max_new_tokens: Hard limit on the number of tokens generated in the response. Distinct from
context_window, which controls the input history size. - top_k: Limits the candidate pool to the K most likely next tokens. Higher values allow more diverse vocabulary.
- top_p: Enables nucleus sampling, dynamically selecting from the smallest set of tokens whose cumulative probability exceeds P.
- repeat_penalty: Penalizes tokens that have already appeared in the generated text. Values above 1.0 discourage repetition; 1.0 disables the feature.
Summary
- Centralized Configuration: All LLM generation parameters are defined in
private_gpt/settings/settings.pyand consumed byprivate_gpt/components/llm/llm_component.py. - Global vs. Backend-Specific: Set
temperatureandmax_new_tokensin thellmblock; configuretop_k,top_p, andrepeat_penaltyin thellamacpporollamablocks. - Flexible Deployment: Tune parameters via static YAML files, environment variable substitution, or programmatic Python instantiation.
- Immediate Effect: Changes require service restart unless using programmatic injection, as the
LLMComponentbuilds the LLM instance once during initialization.
Frequently Asked Questions
What is the difference between max_new_tokens and context_window in PrivateGPT?
max_new_tokens limits how many tokens the model generates in a single response (default 256), while context_window defines the total token budget for the conversation history (typically 3900 or higher). The context window includes system prompts, user messages, and previous assistant responses, whereas max_new_tokens only constrains the new content being generated.
Why don't top_k, top_p, and repeat_penalty appear in the main llm block?
These parameters are backend-specific sampling strategies. According to the source code in private_gpt/settings/settings.py, they reside in LlamaCPPSettings because they control low-level sampling algorithms specific to local inference engines like LlamaCPP and Ollama. Remote APIs like OpenAI or Azure OpenAI handle sampling server-side, so PrivateGPT only exposes these fields for local backends.
How do I verify my parameter changes are actually applied?
Check the application logs during startup. The LLMComponent in private_gpt/components/llm/llm_component.py logs the construction of the LLM instance. You can also add debug logging to print settings.llm and settings.llamacpp after loading the configuration. If using environment variables, ensure the ${VAR:default} syntax matches exactly, including the colon separator.
Can I tune these parameters per-request instead of globally?
PrivateGPT currently loads generation parameters once during the LLMComponent initialization. Per-request tuning would require modifying the component to accept override dictionaries in the complete() or chat() methods. For now, use programmatic instantiation of LLMComponent with different Settings objects if you need distinct parameter sets for different request types.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →