How to Tune Temperature and Top‑P Parameters for LLM Models in the Hiring‑Agent Repository

To tune temperature and top‑p for LLM models in the hiring‑agent repository, modify the MODEL_PARAMETERS dictionary in prompt.py or pass a custom model_params dict to the ResumeEvaluator class.

The interviewstreet/hiring-agent repository centralizes LLM generation settings to ensure consistent resume evaluation. Tuning the temperature and top‑p parameters allows you to control the randomness and diversity of model outputs, from deterministic grading to creative feedback. This guide explains how these parameters are implemented in the codebase and how to adjust them for your specific use case.

Understanding Temperature and Top‑P Parameters

What Temperature Controls

Temperature scales the probability distribution of the next token. In the hiring‑agent codebase, values range from 0.0 to 1.0:

  • 0.0 produces deterministic, repeatable outputs ideal for consistent scoring
  • Higher values (0.7‑1.0) increase variation and creativity, useful for generating diverse feedback suggestions

What Top‑P (Nucleus Sampling) Controls

Top‑p (also called nucleus sampling) limits token selection to the smallest set whose cumulative probability exceeds the threshold. As implemented in this repository:

  • Lower values (e.g., 0.4) restrict the model to high‑probability tokens, producing focused, conservative outputs
  • Higher values (≈ 0.9) allow broader token diversity, which can yield richer but potentially less predictable responses

Where Parameters Are Configured in the Codebase

The repository implements a three‑layer configuration system:

  1. prompt.py (lines 30‑44) — Defines the MODEL_PARAMETERS dictionary that stores default temperature and top_p values for each supported model (Ollama and Gemini).

  2. evaluator.py (lines 68‑72) — The ResumeEvaluator class pulls these defaults (or a custom dictionary) and injects them into the chat request via the options field.

  3. models.py (lines 42‑45) — Converts the supplied options dict into the Gemini generation_config object for Google models.

  4. pdf.py (lines 76‑78) — Demonstrates runtime overrides where the defaults are bypassed for specific calls.

How to Tune Temperature and Top‑P in Practice

Method 1: Edit Default Model Parameters in prompt.py

To change the baseline behavior for a specific model, edit the MODEL_PARAMETERS mapping in prompt.py. For example, to make the default gemma3:4b model more deterministic:


# prompt.py – lines 30-44

MODEL_PARAMETERS = {
    # ... other models ...

    "gemma3:4b": {"temperature": 0.2, "top_p": 0.8},  # tuned for consistency

}

This affects all subsequent evaluations that rely on the default configuration.

Method 2: Override at Runtime via ResumeEvaluator

For one‑off adjustments without modifying source files, pass a custom model_params dictionary when instantiating ResumeEvaluator:

from evaluator import ResumeEvaluator

# Override defaults for this specific evaluation

evaluator = ResumeEvaluator(
    model_name="gemma3:4b",
    model_params={"temperature": 0.3, "top_p": 0.7}
)

resume_text = "... raw resume content ..."
result = evaluator.evaluate_resume(resume_text)
print(result.json())

According to the source code in evaluator.py (lines 68‑72), this custom dict bypasses the defaults and feeds directly into the LLM request payload.

Method 3: Direct Low‑Level API Calls

For direct provider access without the evaluator wrapper, pass the parameters via the options field:

from llm_utils import initialize_llm_provider

provider = initialize_llm_provider("gemma3:4b")
response = provider.chat(
    model="gemma3:4b",
    messages=[
        {"role": "system", "content": "You are a resume reviewer."},
        {"role": "user", "content": "Evaluate this resume."},
    ],
    options={"temperature": 0.2, "top_p": 0.9, "stream": False},
    format=None
)
print(response["message"]["content"])

The following configurations align with common use cases in resume evaluation:

Goal Temperature Top‑P
Deterministic scoring (consistent grading criteria) ≤ 0.1 ≈ 0.9
Creative feedback (wording suggestions, improvements) 0.5 – 0.7 0.8 – 0.9
Exploratory brainstorming (generating project ideas) ≈ 0.9 0.6 – 0.8

Because the repository caps top_p at 0.9 for all models in MODEL_PARAMETERS, you will typically adjust temperature to control variability while keeping top_p fixed at or near 0.9.

Summary

  • Configuration location: Default parameters live in prompt.py inside the MODEL_PARAMETERS dictionary (lines 30‑44).
  • Runtime override: Pass a custom model_params dict to ResumeEvaluator to bypass defaults for specific evaluations.
  • Parameter effects: Lower temperature reduces randomness; lower top‑p restricts token diversity.
  • Best practice: Use temperature ≤ 0.1 for consistent scoring and temperature 0.5‑0.7 for creative feedback tasks.

Frequently Asked Questions

What is the difference between temperature and top_p in LLM generation?

Temperature scales the logits before softmax, affecting the overall randomness of the entire distribution, while top_p (nucleus sampling) truncates the distribution by considering only the highest‑probability tokens whose cumulative probability exceeds the threshold. In the hiring‑agent codebase, temperature controls global variability, whereas top_p limits the candidate token pool.

Where are the default temperature and top_p values stored in the hiring-agent repository?

The defaults are stored in the MODEL_PARAMETERS dictionary defined in prompt.py at lines 30‑44. This mapping contains specific temperature and top_p values for each supported model, such as gemma3:4b and Gemini variants.

Can I override these parameters for a single evaluation without changing the global defaults?

Yes. Instantiate ResumeEvaluator with the model_params argument containing your desired values. As shown in evaluator.py (lines 68‑72), this dictionary overrides the defaults from prompt.py for that specific instance only, without affecting other parts of the application.

For deterministic, repeatable grading, set temperature to 0.0 or 0.1 and keep top_p at approximately 0.9. This configuration minimizes variability between runs while allowing the model to select from the most probable tokens, ensuring consistent evaluation criteria across multiple resumes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →