How to Interpret Hiring Agent Evaluation Output Scores and Categories

Hiring Agent evaluation scores combine four weighted category scores (capped at their respective maximums), bonus points up to 20, and deductions, with the final total clamped at the maximum possible score plus 20 bonus headroom.

The interviewstreet/hiring-agent repository provides a résumé-analysis tool that generates structured evaluation data through LLM-powered assessment. Understanding how to interpret these hiring agent evaluation output scores requires examining the EvaluationData model in main/models.py and the formatting logic in main/score.py.

Understanding the Overall Score Calculation

The final score displayed in the console output follows a specific arithmetic pipeline defined in print_evaluation_results within score.py.

For each of the four categories, the raw LLM score is capped at the category-specific maximum defined in the category_maxes dictionary (lines 78-86):

category_score = min(category_data["score"], category_data["max"])
total_score += category_score
max_score += category_data["max"]

Bonus points (evaluation.bonus_points.total) are added, while deductions (evaluation.deductions.total) are subtracted (lines 57-64). The final result is hard-limited to max_score + 20 to enforce the 20-point bonus headroom (lines 66-69).

The console displays this as:


🎯 OVERALL SCORE: 87.5/100

Breaking Down the Four Category Scores

The evaluation distributes 100 points across four weighted dimensions. Each category is defined in score.py with specific maximum values and evidence requirements:

  • Open Source (35 points): Contributions to open-source projects including pull requests, repositories, and community impact (lines 88-92)
  • Self Projects (30 points): Personal side-projects, demos, or hobby work demonstrating initiative (lines 96-102)
  • Production (25 points): Professional experience in real-world production environments (lines 106-112)
  • Technical Skills (10 points): Depth of core technical expertise covering languages, frameworks, and tools (lines 115-122)

Each category prints with evidence:


🌐 Open Source:          30/35
   Evidence: 3 merged PRs to kubernetes/kubernetes with 150+ stars

The evidence field contains the LLM-generated justification, enabling recruiters to verify the assessment basis.

Bonus Points and Deductions

Bonus Points (Up to 20 Points)

The BonusPoints model in models.py (lines 31-33) enforces ge=0 and le=20 constraints. These points reward exceptional merit such as awards or high-impact open-source leadership:


⭐ BONUS POINTS: 12
   Awarded for leading a popular open-source library with 5k+ GitHub stars.

Deductions

Negative adjustments are recorded in the Deductions model (lines 36-42) with a non-negative total and textual reasons field. These reflect gaps or concerns identified in the résumé:


⚠️  DEDUCTIONS: -5
   Missing recent work experience (gap since 2022).

Reading Qualitative Feedback

Beyond numeric scores, the EvaluationData model (lines 44-50) includes two critical list fields:

  • Key Strengths: 1-5 items highlighting the candidate's strongest attributes
  • Areas for Improvement: 1-5 items identifying gaps or weaknesses

These lists provide context for the numeric scores and guide interview focus areas.

Code Example: Parsing Evaluation Output

To programmatically access evaluation scores, instantiate the models defined in main/models.py and use print_evaluation_results from main/score.py:

from main.models import EvaluationData, Scores, CategoryScore, BonusPoints, Deductions
from main.score import print_evaluation_results

# Construct evaluation data from LLM response

scores = Scores(
    open_source=CategoryScore(score=30, max=35, evidence="3 merged PRs"),
    self_projects=CategoryScore(score=25, max=30, evidence="Personal CLI tool"),
    production=CategoryScore(score=20, max=25, evidence="2 years at Acme Corp"),
    technical_skills=CategoryScore(score=9, max=10, evidence="Python, Docker")
)

bonus = BonusPoints(total=10, breakdown="Speaker at PyCon")
deductions = Deductions(total=2, reasons="Resume missing recent project")

eval_data = EvaluationData(
    scores=scores,
    bonus_points=bonus,
    deductions=deductions,
    key_strengths=["Strong problem-solving", "Clear communication"],
    areas_for_improvement=["More recent production experience"]
)

print_evaluation_results(eval_data, candidate_name="Alice")

This outputs the formatted report with all categories, bonuses, deductions, and qualitative feedback.

Summary

  • Four weighted categories comprise the base score: Open Source (35), Self Projects (30), Production (25), and Technical Skills (10)
  • Score capping occurs at the category level using min(score, max) before summation
  • Bonus ceiling hard-limits the total to max_score + 20 regardless of bonus point calculations
  • Evidence fields provide LLM-generated justifications for each category score
  • Qualitative fields include Key Strengths and Areas for Improvement lists (1-5 items each)

Frequently Asked Questions

What is the maximum possible score in Hiring Agent evaluations?

The theoretical maximum is 120 points: 100 from the four base categories plus 20 bonus points. However, the system clamps the final output to max_score + 20 (typically 120), ensuring bonus points cannot inflate scores beyond this ceiling as implemented in score.py lines 66-69.

How are category scores capped in the evaluation logic?

Each category score is individually capped using Python's min() function before summation. In score.py lines 78-86, the code ensures no single category exceeds its predefined maximum (35, 30, 25, or 10), preventing outlier LLM responses from skewing the total.

What constitutes a deduction in the hiring agent scoring system?

Deductions represent negative adjustments based on résumé gaps or concerns, stored in the Deductions model (models.py lines 36-42). Common reasons include employment gaps, missing required skills, or unclear project descriptions. The total deduction value is subtracted from the summed category scores before the final ceiling is applied.

Where is the evaluation output formatted for display?

The print_evaluation_results function in main/score.py handles all formatting. It receives an EvaluationData Pydantic model instance and outputs the structured console report including emoji-prefixed categories, bonus points, deductions, and qualitative feedback sections.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →