How to Add New Evaluation Categories to the Hiring-Agent Resume Scorer

To add new evaluation categories beyond the default four in interviewstreet/hiring-agent, extend the Scores Pydantic model in models.py, update the JSON schema in resume_evaluation_criteria.jinja, and optionally adjust the CLI output in score.py and CSV export in transform.py.

The interviewstreet/hiring-agent repository evaluates candidate resumes using four hard-coded scoring dimensions by default. When you need to assess additional competencies like leadership or system design, you must extend the schema and prompts to add new evaluation categories while maintaining the pipeline's data integrity.

Understanding the Default Evaluation Structure

The hiring-agent scores resumes across four default dimensions defined in the Scores Pydantic model: open_source, self_projects, production, and technical_skills. These categories are hard-coded in models.py and referenced throughout the evaluation pipeline, including the LLM prompt template, CSV transformer, and CLI pretty-printer. Because the iteration logic uses generic model_dump().items() calls, the system can dynamically handle additional fields without modifying the core scoring algorithm.

Step-by-Step Guide to Adding Custom Evaluation Categories

To define new evaluation categories, update three critical areas of the codebase: the data schema, the LLM prompt, and the output formatters.

Step 1: Extend the Scores Model in models.py

The Scores class in models.py (lines 24-30) defines the schema that Pydantic validates. Add your new category as a CategoryScore field to include it in the EvaluationData.scores structure:


# models.py

class Scores(BaseModel):
    open_source: CategoryScore
    self_projects: CategoryScore
    production: CategoryScore
    technical_skills: CategoryScore
    # New category – adjust the max value as you see fit

    leadership: CategoryScore

Source: [models.py](https://github.com/interviewstreet/hiring-agent/blob/main/models.py#L24-L30)

Step 2: Update the LLM Prompt Template

The resume_evaluation_criteria.jinja template instructs the LLM to return scores in a specific JSON structure. Add your category description and extend the JSON skeleton to ensure the model outputs the new field:


## SCORING CRITERIA

### Leadership (0-15 points)

**HIGH SCORES (12-15 points):**
- Led a team of ≥3 engineers on a production-grade project.
- Demonstrated impact on product direction or organization growth.
- Organized open-source or community initiatives.

**MEDIUM SCORES (6-11 points):**
- Managed small teams or coordinated multi-person efforts.
- Showed mentorship or project-lead responsibilities.

**LOW SCORES (1-5 points):**
- Minor coordination tasks, no clear leadership impact.

...

# At the bottom, extend the JSON schema:

{
    "scores": {
        "open_source": {"score": 0, "max": 35, "evidence": "string"},
        "self_projects": {"score": 0, "max": 30, "evidence": "string"},
        "production": {"score": 0, "max": 25, "evidence": "string"},
        "technical_skills": {"score": 0, "max": 10, "evidence": "string"},
        "leadership": {"score": 0, "max": 15, "evidence": "string"}
    },
    ...
}

Source: prompts/templates/resume_evaluation_criteria.jinja

Step 3: Verify CLI Output Handling in score.py

The print_evaluation_results function in score.py iterates over evaluation.scores.model_dump().items(), meaning new categories appear automatically in the console output. No code changes are required unless you want custom formatting for specific fields:


# score.py

def print_evaluation_results(evaluation: EvaluationData, candidate_name: str = "Candidate"):
    total_score = 0
    # Existing categories are printed in the loop below

    for cat, data in evaluation.scores.model_dump().items():
        print(f"🔹 {cat.replace('_', ' ').title()}: {data['score']}/{data['max']}")
        total_score += data['score']
    # The rest of the function stays unchanged

Step 4: Ensure CSV Export Compatibility in transform.py

The transform_evaluation_response function in transform.py uses generic iteration to build CSV rows. New categories are automatically exported as {category}_score, {category}_max, and {category}_evidence columns without hard-coded field lists:


# transform.py

def transform_evaluation_response(...):
    csv_row = {}
    if evaluation and hasattr(evaluation, "scores"):
        for cat, cat_data in evaluation.scores.model_dump().items():
            csv_row[f"{cat}_score"] = cat_data["score"]
            csv_row[f"{cat}_max"] = cat_data["max"]
            csv_row[f"{cat}_evidence"] = cat_data["evidence"]
    return csv_row

Why the Generic Architecture Matters

The hiring-agent avoids hard-coded category lists in its processing logic. By using model_dump().items() to iterate over score fields, the code automatically accommodates schema extensions. This design means you only need to update the Scores definition and the prompt template to add new evaluation categories—no conditional logic or mapping tables require modification.

Summary

  • Extend the Scores model in models.py with a new CategoryScore field to define the schema
  • Update resume_evaluation_criteria.jinja to instruct the LLM on scoring criteria and JSON output format
  • Verify that score.py and transform.py handle the new fields via their generic model_dump().items() iteration
  • The default four categories (open_source, self_projects, production, technical_skills) serve as templates for new custom evaluation categories

Frequently Asked Questions

Where are the default four evaluation categories defined?

They are defined as fields in the Scores Pydantic model located in models.py at lines 24-30. Each field uses the CategoryScore type to enforce score ranges and evidence requirements.

Do I need to modify the core evaluation logic to add a new category?

No. The evaluation logic iterates generically over evaluation.scores.model_dump().items(), so it automatically processes any field present in the Scores model. You only need to update the schema, prompt template, and optionally the display formatting.

What is the maximum score value for a new evaluation category?

The maximum value is defined in the CategoryScore field definition within the Scores model and specified in the JSON skeleton within resume_evaluation_criteria.jinja. You can set any integer value appropriate for your criteria, such as 15 points for leadership or 20 points for system design.

Will adding a new category break existing CSV exports?

No. The transform.py module uses generic iteration to create CSV columns dynamically. When you add a field like leadership, the export automatically includes leadership_score, leadership_max, and leadership_evidence columns without requiring changes to the transformation logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →