Main Scoring Categories in Hiring Agent: Technical Implementation Guide
The Hiring Agent résumé evaluation system uses four main scoring categories—Open Source (35 points), Self Projects (30 points), Production Experience (25 points), and Technical Skills (10 points)—defined in the category_maxes dictionary within score.py and enforced through the CategoryScore Pydantic model in models.py.
The interviewstreet/hiring-agent repository provides an automated framework for quantifying engineering experience through structured category-based assessment. Understanding the main scoring categories in Hiring Agent is essential for interpreting evaluation outputs and customizing the LLM-powered scoring pipeline according to the actual source implementation.
The Four Core Scoring Categories
Hiring Agent evaluates résumés against four distinct competency areas, each with a specific maximum point allocation hardcoded in the evaluation logic.
Open Source Contributions (35 Points)
The Open Source category, referenced by the key open_source in the codebase, assigns up to 35 points—the highest weight in the evaluation framework. This category assesses contributions to public repositories, maintenance of open-source libraries, and community engagement activities visible on platforms like GitHub. The score reflects both the quantity and demonstrable impact of a candidate's collaborative development work.
Self Projects (30 Points)
Self Projects (self_projects in score.py) cap at 30 points and evaluate personal projects that candidates build and maintain independently. This category recognizes demonstrated initiative, architectural decision-making, and end-to-end project ownership outside of professional employment contexts, prioritizing shipping complete solutions over tutorial exercises.
Production Experience (25 Points)
The Production Experience category uses the key production and allows a maximum of 25 points. This measures professional work history including full-time employment, contract roles, and other commercial development experience. The evaluation focuses on the scale, reliability requirements, and business-critical nature of production systems the candidate has architected or maintained.
Technical Skills (10 Points)
Technical Skills (technical_skills) carries a 10-point maximum and evaluates demonstrated proficiency with specific tools, programming languages, frameworks, and cloud platforms. Unlike the experience-based categories, this assesses the breadth and depth of technical toolkits rather than project outcomes or professional tenure.
Implementation in the Codebase
The scoring architecture spans three primary files that define categories, model the data, and execute evaluations.
Category Definitions in score.py
In score.py, the category_maxes dictionary establishes the authoritative scoring framework:
category_maxes = {
"open_source": 35,
"self_projects": 30,
"production": 25,
"technical_skills": 10,
}
This mapping serves as the single source of truth for maximum point allocations across the application. When the evaluator processes a résumé, it references these caps to normalize raw scores generated by the LLM.
Data Modeling in models.py
The models.py file defines the CategoryScore Pydantic class, which structures every category evaluation with three fields:
score: The raw numeric value assigned by the LLM evaluatormax: The category maximum (mirroringcategory_maxesvalues)evidence: A textual explanation generated by the LLM justifying the score
This model enforces type safety and validation across the evaluation pipeline, storing both quantitative metrics and qualitative reasoning for auditability.
Evaluation Orchestration in evaluator.py
The evaluator.py module coordinates the LLM calls that produce EvaluationData containing populated CategoryScore instances. This file handles the prompting logic that instructs the language model to assess résumés against the four defined categories and generate appropriate evidence strings supporting each numerical assignment.
Working with Category Scores
The following Python pattern from score.py demonstrates how to access and display categorized scoring data:
# Assume `evaluation` is an EvaluationData instance returned by the LLM evaluator
category_maxes = {
"open_source": 35,
"self_projects": 30,
"production": 25,
"technical_skills": 10,
}
for cat_name in category_maxes:
cat_score: CategoryScore = getattr(evaluation.scores, cat_name)
capped = min(cat_score.score, category_maxes[cat_name])
print(f"{cat_name.replace('_', ' ').title():<25} {capped}/{cat_score.max}")
print(f" Evidence: {cat_score.evidence}\n")
Running this code produces formatted output such as:
Open Source 28/35
Evidence: Contributed to three well‑known OSS libraries.
Self Projects 22/30
Evidence: Built a personal web‑scraper used by 200+ users.
Production Experience 20/25
Evidence: 3 years as a backend engineer at Acme Corp.
Technical Skills 8/10
Evidence: Proficient in Python, Docker, and AWS.
The getattr approach dynamically retrieves each CategoryScore from the evaluation.scores object, while the min() function enforces the category maximums defined in category_maxes to ensure normalized scoring.
Summary
- Four categories dominate Hiring Agent scoring: Open Source (35 pts), Self Projects (30 pts), Production Experience (25 pts), and Technical Skills (10 pts).
- Source of truth: The
category_maxesdictionary inscore.pydefines maximum point allocations using keysopen_source,self_projects,production, andtechnical_skills. - Data structure: Each category uses the
CategoryScorePydantic model inmodels.pyto store the numeric score, maximum value, and LLM-generated evidence string. - Integration:
evaluator.pyorchestrates the LLM evaluation that populates these scores, whilescore.pyhandles normalization and display logic. - Output format: Scores display as
capped_score/maximumwith accompanying evidence text explaining the evaluation rationale.
Frequently Asked Questions
What are the exact point allocations for each scoring category?
Open Source receives a maximum of 35 points, Self Projects caps at 30 points, Production Experience allows 25 points, and Technical Skills limits at 10 points. These values are hardcoded in the category_maxes dictionary within score.py according to the interviewstreet/hiring-agent source code.
Where does Hiring Agent store the scoring logic and category definitions?
The primary category definitions and maximum point mappings reside in score.py as the category_maxes dictionary. The underlying data model for individual scores is implemented in models.py through the CategoryScore class, while evaluator.py contains the orchestration logic for LLM-based assessments.
How does the system handle evidence for category scores?
Each category score includes an evidence string field within the CategoryScore Pydantic model defined in models.py. This field stores the LLM's textual justification for the assigned score, providing transparency and auditability for the evaluation results.
What prevents category scores from exceeding their maximums?
The scoring implementation in score.py uses a min() operation to cap raw scores against the category_maxes values. This ensures that even if the LLM evaluator assigns a higher raw value, the final displayed and computed score cannot exceed the predefined category limits.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →