How to Improve Resume Extraction Accuracy by Tuning LLM Model Parameters in Hiring-Agent
You can improve resume extraction accuracy in the Hiring-Agent pipeline by lowering the temperature to 0.0–0.1 and adjusting top_p between 0.4–0.9 in the MODEL_PARAMETERS dictionary defined in prompt.py, which controls the generation behavior for all LLM calls made by PDFHandler.
The Hiring-Agent repository by InterviewStreet provides a configurable resume parsing pipeline that relies on Large Language Models to convert PDF documents into structured JSON. By tuning the generation parameters in the central configuration file, you can directly influence the determinism and accuracy of the extracted data without modifying the core extraction logic.
Understanding the Resume Extraction Pipeline
The resume extraction system operates through a three-stage architecture defined in pdf.py and prompt.py.
The Three-Stage Architecture
-
PDF Ingestion – The
PDFHandler.extract_text_from_pdfmethod uses PyMuPDF to read input files and converts each page to markdown format usingto_markdown【pdf.py†L52-L58】. -
Section-Wise LLM Processing – For every resume section (basics, work, education, skills), the handler constructs a system prompt, merges it with user content, and forwards the request via
self.provider.chat(...)【pdf.py†L66-L100】. -
Parameter Injection – Before making the LLM call, the handler retrieves generation parameters from
MODEL_PARAMETERSand injects them into the request options【pdf.py†L75-L98】.
Where LLM Parameters Are Configured
All model behavior is controlled through two central dictionaries in prompt.py that serve as the single source of truth for the entire pipeline.
DEFAULT_MODEL and MODEL_PARAMETERS
The DEFAULT_MODEL variable (environment-overridable) selects the active model name, such as gemma3:4b【prompt.py†L16-L22】. The MODEL_PARAMETERS dictionary maps each model to its generation knobs:
# prompt.py
MODEL_PARAMETERS = {
"gemma3:4b": {"temperature": 0.1, "top_p": 0.9},
# Additional model configurations...
}
During execution, PDFHandler fetches the appropriate entry using model_params = MODEL_PARAMETERS.get(DEFAULT_MODEL, {"temperature": 0.1, "top_p": 0.9}) and passes these values to the provider【pdf.py†L75-L98】【prompt.py†L27-L44】.
Key Parameters That Affect Extraction Accuracy
Three specific knobs directly influence the quality and structure of the generated JSON output.
Temperature
Temperature controls randomness in token generation. Lower values (0.0–0.1) make output more deterministic, reducing hallucinations and ensuring strict adherence to JSON schemas for sections like work history and education.
Top-p (Nucleus Sampling)
Top-p limits token sampling to the most probable nucleus. A tight nucleus (0.4) prevents the model from drifting into free-form text, while a looser value (0.9) allows richer phrasing for unstructured sections like summaries.
Model Selection
The specific model identifier in DEFAULT_MODEL determines architectural capabilities. Larger models often handle complex formatting better, though they may introduce verbosity. The repository supports Ollama-compatible models like gemma3:4b and qwen3:4b through the provider factory in llm_utils.py.
Practical Tuning Strategies
To maximize extraction fidelity, apply deterministic configurations to structured sections while allowing slight variation for narrative content.
Recommended Configuration Values
| Parameter | Effect on Extraction | High-Accuracy Setting |
|---|---|---|
temperature |
Controls randomness; lower values reduce hallucinations | 0.0 – 0.1 |
top_p |
Limits token sampling nucleus; tighter values enforce structure | 0.4 – 0.9 (tight for structured data) |
Strategy: Set temperature to 0.0 for sections requiring strict JSON schema compliance (basics, work, education). Use top_p of 0.4 to prevent the model from generating explanatory text outside the requested JSON structure.
Implementation Examples
Example 1: Basic Configuration in prompt.py
Modify the MODEL_PARAMETERS dictionary to enforce deterministic output for the default model:
# prompt.py
MODEL_PARAMETERS = {
"gemma3:4b": {"temperature": 0.0, "top_p": 0.4}, # High accuracy, tight nucleus
"qwen3:4b": {"temperature": 0.0, "top_p": 0.5},
}
Example 2: Runtime Override
Override model selection and parameters programmatically without editing source files:
import os
from pdf import PDFHandler
from prompt import MODEL_PARAMETERS
# Override via environment or direct assignment
os.environ["DEFAULT_MODEL"] = "qwen3:4b"
os.environ["LLM_PROVIDER"] = "ollama"
# Runtime parameter adjustment
MODEL_PARAMETERS["qwen3:4b"] = {"temperature": 0.0, "top_p": 0.5}
handler = PDFHandler()
resume = handler.extract_json_from_pdf("candidate_resume.pdf")
Example 3: Per-Section Tuning
For advanced use cases requiring different behaviors per section, temporarily modify parameters before specific extraction calls:
from pdf import PDFHandler
from prompt import MODEL_PARAMETERS, DEFAULT_MODEL
handler = PDFHandler()
# Store original parameters
orig_params = MODEL_PARAMETERS[DEFAULT_MODEL].copy()
# Increase temperature for skills section to capture varied phrasing
MODEL_PARAMETERS[DEFAULT_MODEL] = {"temperature": 0.2, "top_p": 0.9}
skills = handler.extract_skills_section(resume_text)
# Restore strict parameters for subsequent sections
MODEL_PARAMETERS[DEFAULT_MODEL] = orig_params
Summary
- Centralized configuration: All LLM generation parameters live in
prompt.pywithinMODEL_PARAMETERS, providing a single point of truth for the entire extraction pipeline. - Deterministic extraction: Setting
temperatureto0.0andtop_pto0.4maximizes JSON schema compliance and minimizes hallucinations in structured resume sections. - Runtime flexibility: Parameters can be overridden via environment variables or direct dictionary manipulation without modifying the core
PDFHandlerlogic inpdf.py. - Provider abstraction: The parameter injection occurs uniformly across all LLM providers (Ollama, Gemini) through the
LLMProviderprotocol defined inmodels.py.
Frequently Asked Questions
What is the default temperature setting in Hiring-Agent?
The default fallback temperature is 0.1 with a top_p of 0.9, defined in the fallback dictionary within PDFHandler when a model is not explicitly configured in MODEL_PARAMETERS【pdf.py†L75-L98】.
Can I use different parameters for different resume sections?
Yes. While the base implementation uses global parameters, you can implement per-section tuning by temporarily overriding MODEL_PARAMETERS[DEFAULT_MODEL] before calling specific extraction methods like extract_skills_section, then restoring the original values afterward.
Which file contains the LLM provider implementation?
The provider factory and model-specific implementations reside in llm_utils.py and models.py. These files handle the concrete chat interface for providers like OllamaProvider and GeminiProvider, ensuring uniform parameter consumption across different backend services.
How does the system handle unsupported models?
If DEFAULT_MODEL references a model not present in MODEL_PARAMETERS, the system falls back to {"temperature": 0.1, "top_p": 0.9} during the parameter injection phase in pdf.py【pdf.py†L75-L98】, ensuring the pipeline continues to function with conservative default values.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →