How to Improve Resume Extraction Accuracy by Tuning LLM Model Parameters in Hiring-Agent

You can improve resume extraction accuracy in the Hiring-Agent pipeline by lowering the temperature to 0.0–0.1 and adjusting top_p between 0.4–0.9 in the MODEL_PARAMETERS dictionary defined in prompt.py, which controls the generation behavior for all LLM calls made by PDFHandler.

The Hiring-Agent repository by InterviewStreet provides a configurable resume parsing pipeline that relies on Large Language Models to convert PDF documents into structured JSON. By tuning the generation parameters in the central configuration file, you can directly influence the determinism and accuracy of the extracted data without modifying the core extraction logic.

Understanding the Resume Extraction Pipeline

The resume extraction system operates through a three-stage architecture defined in pdf.py and prompt.py.

The Three-Stage Architecture

  1. PDF Ingestion – The PDFHandler.extract_text_from_pdf method uses PyMuPDF to read input files and converts each page to markdown format using to_markdown【pdf.py†L52-L58】.

  2. Section-Wise LLM Processing – For every resume section (basics, work, education, skills), the handler constructs a system prompt, merges it with user content, and forwards the request via self.provider.chat(...)【pdf.py†L66-L100】.

  3. Parameter Injection – Before making the LLM call, the handler retrieves generation parameters from MODEL_PARAMETERS and injects them into the request options【pdf.py†L75-L98】.

Where LLM Parameters Are Configured

All model behavior is controlled through two central dictionaries in prompt.py that serve as the single source of truth for the entire pipeline.

DEFAULT_MODEL and MODEL_PARAMETERS

The DEFAULT_MODEL variable (environment-overridable) selects the active model name, such as gemma3:4b【prompt.py†L16-L22】. The MODEL_PARAMETERS dictionary maps each model to its generation knobs:


# prompt.py

MODEL_PARAMETERS = {
    "gemma3:4b": {"temperature": 0.1, "top_p": 0.9},
    # Additional model configurations...

}

During execution, PDFHandler fetches the appropriate entry using model_params = MODEL_PARAMETERS.get(DEFAULT_MODEL, {"temperature": 0.1, "top_p": 0.9}) and passes these values to the provider【pdf.py†L75-L98】【prompt.py†L27-L44】.

Key Parameters That Affect Extraction Accuracy

Three specific knobs directly influence the quality and structure of the generated JSON output.

Temperature

Temperature controls randomness in token generation. Lower values (0.0–0.1) make output more deterministic, reducing hallucinations and ensuring strict adherence to JSON schemas for sections like work history and education.

Top-p (Nucleus Sampling)

Top-p limits token sampling to the most probable nucleus. A tight nucleus (0.4) prevents the model from drifting into free-form text, while a looser value (0.9) allows richer phrasing for unstructured sections like summaries.

Model Selection

The specific model identifier in DEFAULT_MODEL determines architectural capabilities. Larger models often handle complex formatting better, though they may introduce verbosity. The repository supports Ollama-compatible models like gemma3:4b and qwen3:4b through the provider factory in llm_utils.py.

Practical Tuning Strategies

To maximize extraction fidelity, apply deterministic configurations to structured sections while allowing slight variation for narrative content.

Parameter Effect on Extraction High-Accuracy Setting
temperature Controls randomness; lower values reduce hallucinations 0.0 – 0.1
top_p Limits token sampling nucleus; tighter values enforce structure 0.4 – 0.9 (tight for structured data)

Strategy: Set temperature to 0.0 for sections requiring strict JSON schema compliance (basics, work, education). Use top_p of 0.4 to prevent the model from generating explanatory text outside the requested JSON structure.

Implementation Examples

Example 1: Basic Configuration in prompt.py

Modify the MODEL_PARAMETERS dictionary to enforce deterministic output for the default model:


# prompt.py

MODEL_PARAMETERS = {
    "gemma3:4b": {"temperature": 0.0, "top_p": 0.4},  # High accuracy, tight nucleus

    "qwen3:4b": {"temperature": 0.0, "top_p": 0.5},
}

Example 2: Runtime Override

Override model selection and parameters programmatically without editing source files:

import os
from pdf import PDFHandler
from prompt import MODEL_PARAMETERS

# Override via environment or direct assignment

os.environ["DEFAULT_MODEL"] = "qwen3:4b"
os.environ["LLM_PROVIDER"] = "ollama"

# Runtime parameter adjustment

MODEL_PARAMETERS["qwen3:4b"] = {"temperature": 0.0, "top_p": 0.5}

handler = PDFHandler()
resume = handler.extract_json_from_pdf("candidate_resume.pdf")

Example 3: Per-Section Tuning

For advanced use cases requiring different behaviors per section, temporarily modify parameters before specific extraction calls:

from pdf import PDFHandler
from prompt import MODEL_PARAMETERS, DEFAULT_MODEL

handler = PDFHandler()

# Store original parameters

orig_params = MODEL_PARAMETERS[DEFAULT_MODEL].copy()

# Increase temperature for skills section to capture varied phrasing

MODEL_PARAMETERS[DEFAULT_MODEL] = {"temperature": 0.2, "top_p": 0.9}
skills = handler.extract_skills_section(resume_text)

# Restore strict parameters for subsequent sections

MODEL_PARAMETERS[DEFAULT_MODEL] = orig_params

Summary

  • Centralized configuration: All LLM generation parameters live in prompt.py within MODEL_PARAMETERS, providing a single point of truth for the entire extraction pipeline.
  • Deterministic extraction: Setting temperature to 0.0 and top_p to 0.4 maximizes JSON schema compliance and minimizes hallucinations in structured resume sections.
  • Runtime flexibility: Parameters can be overridden via environment variables or direct dictionary manipulation without modifying the core PDFHandler logic in pdf.py.
  • Provider abstraction: The parameter injection occurs uniformly across all LLM providers (Ollama, Gemini) through the LLMProvider protocol defined in models.py.

Frequently Asked Questions

What is the default temperature setting in Hiring-Agent?

The default fallback temperature is 0.1 with a top_p of 0.9, defined in the fallback dictionary within PDFHandler when a model is not explicitly configured in MODEL_PARAMETERS【pdf.py†L75-L98】.

Can I use different parameters for different resume sections?

Yes. While the base implementation uses global parameters, you can implement per-section tuning by temporarily overriding MODEL_PARAMETERS[DEFAULT_MODEL] before calling specific extraction methods like extract_skills_section, then restoring the original values afterward.

Which file contains the LLM provider implementation?

The provider factory and model-specific implementations reside in llm_utils.py and models.py. These files handle the concrete chat interface for providers like OllamaProvider and GeminiProvider, ensuring uniform parameter consumption across different backend services.

How does the system handle unsupported models?

If DEFAULT_MODEL references a model not present in MODEL_PARAMETERS, the system falls back to {"temperature": 0.1, "top_p": 0.9} during the parameter injection phase in pdf.py【pdf.py†L75-L98】, ensuring the pipeline continues to function with conservative default values.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →