How the Hiring-Agent Adjusts Interview Question Difficulty Levels
The hiring-agent determines interview question difficulty by combining explicit skill-level metadata extracted from parsed résumés with four quantitative evaluation scores generated by a large language model, then feeding both signals into dynamic prompt templates that instruct the LLM to generate appropriately graded questions.
The interviewstreet/hiring-agent repository implements an adaptive difficulty system that automatically calibrates technical interview questions to match candidate capabilities. By analyzing both structured skill declarations and LLM-generated evaluation metrics, the system ensures that junior developers receive fundamental concept checks while senior engineers face complex, multi-step algorithmic challenges.
Extracting Skill-Level Metadata from Résumés
The difficulty adjustment process begins in transform.py, where the convert_json_resume_to_text routine parses structured résumé data and preserves optional proficiency indicators. When processing the skills section, the code checks for a level field on each skill entry—typically values such as "Beginner", "Intermediate", or "Expert"—and includes this metadata in the rendered text output.
According to the source code at lines 814-815, the transformation logic explicitly appends the skill level when present:
# Extract skill levels from the parsed résumé (transform.py)
for skill in resume_data.skills:
print(f"Skill: {skill.name}")
if skill.level:
print(f" Level: {skill.level}") # <-- used later to set difficulty
This preserved metadata serves as the first input signal for the downstream difficulty calculation, providing explicit candidate-declared proficiency data that the system uses to anchor question complexity.
Generating Quantitative Scores with ResumeEvaluator
The second signal comes from evaluator.py, where the ResumeEvaluator class analyzes the rendered résumé text using a large language model. The evaluator prompts the LLM to score the candidate across four distinct pillars: Open-Source, Self-Projects, Production, and Technical-Skills.
At lines 73-88, the code extracts these scores from the LLM response and populates an EvaluationData object with the following fields:
open_source_scoreself_projects_scoreproduction_scoretechnical_skills_score
Each score typically ranges from 0 to 10, providing normalized quantitative metrics that represent the candidate's demonstrated expertise across different dimensions of software engineering.
Mapping Evaluation Data to Difficulty Levels
The TemplateManager class in prompts/template_manager.py combines both signals—the explicit skill levels from the résumé parser and the four pillar scores from the evaluator—to determine the appropriate question difficulty. The system employs a runtime inference logic that maps combined input signals to specific difficulty tiers:
- Low skill level combined with low evaluation scores triggers easy questions testing basic concepts
- Medium skill level combined with mid-range scores triggers medium questions probing deeper understanding
- High skill level combined with high scores triggers hard questions presenting complex scenarios or multi-step algorithms
This mapping is not hard-coded into business logic but rather communicated to the LLM through the system message, allowing the model to dynamically adjust its question generation based on the quantitative and qualitative inputs provided.
Implementation in Prompt Templates
The difficulty selection logic manifests in the prompt templates managed by TemplateManager. The following pseudocode demonstrates how the system calculates an average score across the four pillars and selects a difficulty tier before constructing the final prompt:
# Decide difficulty tier (pseudo-logic used in the prompt template)
average = sum(s["score"] for s in scores.values()) / len(scores)
if average >= 8:
difficulty = "hard"
elif average >= 5:
difficulty = "medium"
else:
difficulty = "easy"
# Pass the chosen tier to the LLM when asking for interview questions
prompt = f"""\
Based on the résumé above, generate a set of interview questions.
Use the candidate's skill levels and the overall score ({average:.1f}) to decide difficulty.
Return questions grouped under Easy, Medium, and Hard sections.
"""
The prompt.py module supplies the model-specific configuration—including DEFAULT_MODEL and other constants—that the ResumeEvaluator uses when submitting these prompts to the LLM backend.
Summary
- Dual-signal architecture: The system combines explicit
skill.levelmetadata fromtransform.pywith quantitative four-pillar scores fromevaluator.pyto determine question difficulty. - Four-pillar evaluation: The
ResumeEvaluatorgenerates scores for Open-Source contributions, Self-Projects, Production experience, and Technical-Skills proficiency. - Dynamic prompt construction:
TemplateManagerfeeds both signals into LLM prompts that request difficulty-graded questions rather than applying rigid business logic. - Runtime adaptation: Difficulty tiers (easy, medium, hard) are inferred at runtime based on the candidate's specific résumé content, ensuring personalized interview experiences.
Frequently Asked Questions
How does the hiring-agent handle résumés that lack explicit skill-level metadata?
When the optional level field is absent from skill entries in transform.py, the system relies entirely on the four quantitative pillar scores generated by ResumeEvaluator. The LLM prompt templates are designed to weight the evaluation scores more heavily when explicit skill levels are unavailable, ensuring that the difficulty inference remains robust even with incomplete structured data.
What are the four evaluation pillars used to score candidates?
According to the EvaluationData structure in evaluator.py, the four pillars are Open-Source (contributions to public repositories), Self-Projects (personal or side projects), Production (professional enterprise experience), and Technical-Skills (breadth and depth of technology stack). Each pillar receives an individual score between 0 and 10, which collectively inform the difficulty calculation.
Where is the prompt template logic that requests difficulty-graded questions?
The prompt templates are managed by the TemplateManager class located in prompts/template_manager.py. This module loads template files that contain instructions for the LLM to generate questions grouped by difficulty, incorporating the skill levels extracted in transform.py and the evaluation scores computed in evaluator.py (lines 73-88).
Can the difficulty adjustment system be configured to use different scoring thresholds?
While the example logic shows thresholds of 5 and 8 for medium and hard difficulties respectively, these values are implemented within the prompt template instructions rather than hard-coded constants in the Python source. Modifying the difficulty curve requires adjusting the template files loaded by TemplateManager or updating the scoring logic in the evaluation pipeline, with the specific model configuration referenced in prompt.py determining the LLM's interpretation of these thresholds.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →