Hiring Agent Evaluation Categories: Understanding the Four Resume Scoring Dimensions
The Hiring Agent evaluates résumés across four distinct categories: Open Source, Self Projects, Production Experience, and Technical Skills, each with specific maximum scores and evidence-based justifications.
The interviewstreet/hiring-agent repository implements an automated résumé evaluation system that uses an LLM to score candidates across standardized dimensions. Understanding these evaluation categories is essential for developers integrating the tool or candidates optimizing their résumés for assessment.
The Four Evaluation Categories Used by the Hiring Agent
According to the source code in models.py, the Hiring Agent analyzes résumés through four specific lenses, each represented by a CategoryScore object containing a raw score, maximum possible points, and textual evidence.
Open Source Contributions
Open Source scoring measures contributions to public open-source projects. This category carries the highest weight, with a maximum cap of 35 points as implemented in score.py (lines 87-94). The evaluator looks for demonstrated collaboration, code quality, and community engagement in publicly accessible repositories.
Self Projects
Self Projects captures personal side-projects that are not necessarily open-source but demonstrate technical initiative. Defined in models.py (lines 225-228) and rendered in score.py (lines 95-102), this category allows candidates to showcase independent work, portfolio pieces, and experimental codebases. The maximum score for this category is 30 points.
Production Experience
Production Experience evaluates work performed in professional, production-grade environments such as corporate jobs or contracted services. This category, capped at 25 points according to score.py (lines 105-112), assesses the candidate's ability to ship code in real-world scenarios with considerations for reliability, scalability, and maintainability.
Technical Skills
Technical Skills assesses demonstrated competencies in programming languages, frameworks, tools, and methodologies. With a maximum of 10 points (the lowest weight among the four), this category captures the breadth and depth of technologies mentioned in the résumé and their contextual application. The rendering logic appears in score.py (lines 113-121).
How Category Scoring Works in the Source Code
Each evaluation category is instantiated as a CategoryScore object containing three fields: score (float), max (integer), and evidence (string). These objects are assembled into the Scores model in models.py, which serves as the data structure for the ResumeEvaluator class in evaluator.py.
The ResumeEvaluator processes the résumé text, prompts the LLM to return structured JSON, and parses the results into an EvaluationData payload containing the four category scores. Before display, the CLI formatter in score.py applies hard caps to each category to ensure scores do not exceed their predefined maximums (35, 30, 25, and 10 respectively).
The transform.py module subsequently converts these scored categories into CSV rows for export and analysis, preserving the individual category scores alongside the aggregate total.
Working with Evaluation Categories in Code
Accessing Category Scores Programmatically
You can inspect individual category scores after evaluating a résumé by accessing the scores attribute on the EvaluationData object:
from evaluator import ResumeEvaluator
from models import EvaluationData
evaluator = ResumeEvaluator()
evaluation: EvaluationData = evaluator.evaluate_resume(resume_text)
# Access individual category scores
open_source = evaluation.scores.open_source
self_projects = evaluation.scores.self_projects
production = evaluation.scores.production
technical_skills = evaluation.scores.technical_skills
print(f"Open Source: {open_source.score}/{open_source.max}")
print(f"Self Projects: {self_projects.score}/{self_projects.max}")
print(f"Production: {production.score}/{production.max}")
print(f"Technical Skills: {technical_skills.score}/{technical_skills.max}")
Rendering Categories with Score Caps
When displaying results in a CLI interface, apply the category-specific maximums as implemented in score.py:
# Simplified excerpt from score.py demonstrating category caps
category_caps = {
"open_source": 35,
"self_projects": 30,
"production": 25,
"technical_skills": 10,
}
for category_name, cat in evaluation.scores.model_dump().items():
max_points = category_caps[category_name]
display_score = min(cat['score'], max_points)
print(f"{category_name.replace('_', ' ').title()}: {display_score}/{cat['max']}")
Summary
- The Hiring Agent uses four evaluation categories: Open Source (35 pts max), Self Projects (30 pts), Production Experience (25 pts), and Technical Skills (10 pts).
- Categories are defined in
models.py(lines 225-228) as part of theScoresdata model. - Each category uses a
CategoryScoreobject containingscore,max, andevidencefields. - The
ResumeEvaluatorinevaluator.pygenerates these scores, whilescore.pyapplies caps and renders CLI output. - The
transform.pymodule handles CSV serialization of category scores for downstream analysis.
Frequently Asked Questions
What are the maximum scores for each evaluation category?
The Hiring Agent enforces specific caps for each category: 35 points for Open Source, 30 points for Self Projects, 25 points for Production Experience, and 10 points for Technical Skills. These maximums are hardcoded in score.py and applied before calculating the aggregate score.
How does the Hiring Agent calculate the overall resume score?
The overall score is the sum of the four capped category scores. The ResumeEvaluator parses LLM output into CategoryScore objects, score.py ensures each score does not exceed its category maximum, and the system aggregates the values into a final percentage or point total displayed in the CLI.
Where are the evaluation categories defined in the codebase?
The four evaluation categories are defined as fields in the Scores model within models.py (lines 225-228). The scoring logic, evidence parsing, and maximum value enforcement are implemented across evaluator.py (generation), score.py (formatting and capping), and transform.py (CSV export).
Can I customize the evaluation categories or their weights?
While the current implementation in interviewstreet/hiring-agent uses fixed categories with hardcoded maximums (35/30/25/10), the modular architecture in models.py allows for modification of the Scores class and CategoryScore structure. To adjust weights, you would need to modify the caps in score.py and potentially update the LLM prompting logic in evaluator.py to reflect new evaluation criteria.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →