How to Add New Evaluation Categories to the Hiring-Agent Resume Scorer
To add new evaluation categories beyond the default four in interviewstreet/hiring-agent, extend the Scores Pydantic model in models.py, update the JSON schema in resume_evaluation_criteria.jinja, and optionally adjust the CLI output in score.py and CSV export in transform.py.
The interviewstreet/hiring-agent repository evaluates candidate resumes using four hard-coded scoring dimensions by default. When you need to assess additional competencies like leadership or system design, you must extend the schema and prompts to add new evaluation categories while maintaining the pipeline's data integrity.
Understanding the Default Evaluation Structure
The hiring-agent scores resumes across four default dimensions defined in the Scores Pydantic model: open_source, self_projects, production, and technical_skills. These categories are hard-coded in models.py and referenced throughout the evaluation pipeline, including the LLM prompt template, CSV transformer, and CLI pretty-printer. Because the iteration logic uses generic model_dump().items() calls, the system can dynamically handle additional fields without modifying the core scoring algorithm.
Step-by-Step Guide to Adding Custom Evaluation Categories
To define new evaluation categories, update three critical areas of the codebase: the data schema, the LLM prompt, and the output formatters.
Step 1: Extend the Scores Model in models.py
The Scores class in models.py (lines 24-30) defines the schema that Pydantic validates. Add your new category as a CategoryScore field to include it in the EvaluationData.scores structure:
# models.py
class Scores(BaseModel):
open_source: CategoryScore
self_projects: CategoryScore
production: CategoryScore
technical_skills: CategoryScore
# New category – adjust the max value as you see fit
leadership: CategoryScore
Source: [models.py](https://github.com/interviewstreet/hiring-agent/blob/main/models.py#L24-L30)
Step 2: Update the LLM Prompt Template
The resume_evaluation_criteria.jinja template instructs the LLM to return scores in a specific JSON structure. Add your category description and extend the JSON skeleton to ensure the model outputs the new field:
## SCORING CRITERIA
### Leadership (0-15 points)
**HIGH SCORES (12-15 points):**
- Led a team of ≥3 engineers on a production-grade project.
- Demonstrated impact on product direction or organization growth.
- Organized open-source or community initiatives.
**MEDIUM SCORES (6-11 points):**
- Managed small teams or coordinated multi-person efforts.
- Showed mentorship or project-lead responsibilities.
**LOW SCORES (1-5 points):**
- Minor coordination tasks, no clear leadership impact.
...
# At the bottom, extend the JSON schema:
{
"scores": {
"open_source": {"score": 0, "max": 35, "evidence": "string"},
"self_projects": {"score": 0, "max": 30, "evidence": "string"},
"production": {"score": 0, "max": 25, "evidence": "string"},
"technical_skills": {"score": 0, "max": 10, "evidence": "string"},
"leadership": {"score": 0, "max": 15, "evidence": "string"}
},
...
}
Source: prompts/templates/resume_evaluation_criteria.jinja
Step 3: Verify CLI Output Handling in score.py
The print_evaluation_results function in score.py iterates over evaluation.scores.model_dump().items(), meaning new categories appear automatically in the console output. No code changes are required unless you want custom formatting for specific fields:
# score.py
def print_evaluation_results(evaluation: EvaluationData, candidate_name: str = "Candidate"):
total_score = 0
# Existing categories are printed in the loop below
for cat, data in evaluation.scores.model_dump().items():
print(f"🔹 {cat.replace('_', ' ').title()}: {data['score']}/{data['max']}")
total_score += data['score']
# The rest of the function stays unchanged
Step 4: Ensure CSV Export Compatibility in transform.py
The transform_evaluation_response function in transform.py uses generic iteration to build CSV rows. New categories are automatically exported as {category}_score, {category}_max, and {category}_evidence columns without hard-coded field lists:
# transform.py
def transform_evaluation_response(...):
csv_row = {}
if evaluation and hasattr(evaluation, "scores"):
for cat, cat_data in evaluation.scores.model_dump().items():
csv_row[f"{cat}_score"] = cat_data["score"]
csv_row[f"{cat}_max"] = cat_data["max"]
csv_row[f"{cat}_evidence"] = cat_data["evidence"]
return csv_row
Why the Generic Architecture Matters
The hiring-agent avoids hard-coded category lists in its processing logic. By using model_dump().items() to iterate over score fields, the code automatically accommodates schema extensions. This design means you only need to update the Scores definition and the prompt template to add new evaluation categories—no conditional logic or mapping tables require modification.
Summary
- Extend the
Scoresmodel inmodels.pywith a newCategoryScorefield to define the schema - Update
resume_evaluation_criteria.jinjato instruct the LLM on scoring criteria and JSON output format - Verify that
score.pyandtransform.pyhandle the new fields via their genericmodel_dump().items()iteration - The default four categories (
open_source,self_projects,production,technical_skills) serve as templates for new custom evaluation categories
Frequently Asked Questions
Where are the default four evaluation categories defined?
They are defined as fields in the Scores Pydantic model located in models.py at lines 24-30. Each field uses the CategoryScore type to enforce score ranges and evidence requirements.
Do I need to modify the core evaluation logic to add a new category?
No. The evaluation logic iterates generically over evaluation.scores.model_dump().items(), so it automatically processes any field present in the Scores model. You only need to update the schema, prompt template, and optionally the display formatting.
What is the maximum score value for a new evaluation category?
The maximum value is defined in the CategoryScore field definition within the Scores model and specified in the JSON skeleton within resume_evaluation_criteria.jinja. You can set any integer value appropriate for your criteria, such as 15 points for leadership or 20 points for system design.
Will adding a new category break existing CSV exports?
No. The transform.py module uses generic iteration to create CSV columns dynamically. When you add a field like leadership, the export automatically includes leadership_score, leadership_max, and leadership_evidence columns without requiring changes to the transformation logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →