How to Add Custom Evaluation Criteria to the Hiring Agent

To add custom evaluation criteria to the Hiring Agent, extend the EvaluationData Pydantic model in models.py, update the resume_evaluation_criteria.jinja template to define the new scoring rules, and ensure your template is registered in TemplateManager._load_templates if creating a new file.

The Hiring Agent from the interviewstreet/hiring-agent repository evaluates resumes using a structured LLM pipeline governed by Jinja prompt templates. When you need to assess candidates on dimensions beyond the default categories—such as teamwork, leadership, or specialized domain knowledge—you must modify the underlying data schema and prompt instructions. This guide provides the exact implementation steps based on the source code architecture.

Understanding the Evaluation Architecture

The Hiring Agent processes resumes through a coordinated pipeline of four core components:

  • ResumeEvaluator (evaluator.py): Orchestrates the evaluation by instantiating TemplateManager, rendering the criteria template, and parsing the LLM response into structured data.
  • TemplateManager (prompts/template_manager.py): Discovers and loads Jinja templates from the prompts/templates directory.
  • resume_evaluation_criteria.jinja (prompts/templates/resume_evaluation_criteria.jinja): Contains the natural-language scoring rubric and the exact JSON schema the LLM must return.
  • EvaluationData (models.py): Pydantic model that validates the JSON output from the LLM, ensuring type safety and required field presence.

When ResumeEvaluator.evaluate_resume processes a resume (see lines 41-44 of evaluator.py), it calls TemplateManager to render the resume_evaluation_criteria template and sends the resulting prompt to the LLM. The response must conform to the EvaluationData schema defined in models.py.

Step-by-Step Implementation

Step 1: Extend the EvaluationData Model

Add your new evaluation fields to the EvaluationData class in models.py. This guarantees that the JSON returned by the LLM can be parsed without validation errors.

The following example adds a "teamwork" category with a 0-20 point scale:


# models.py – add a new field to EvaluationData

from typing import Dict, Any, Optional
from pydantic import BaseModel, Field

class EvaluationData(BaseModel):
    scores: Dict[str, Any]  # existing structure

    # … existing fields …

    
    # New field: teamwork (0‑20 points)

    teamwork: Optional[Dict[str, Any]] = Field(
        default_factory=lambda: {"score": 0, "max": 20, "evidence": ""}
    )

Step 2: Update the Jinja Prompt Template

Edit prompts/templates/resume_evaluation_criteria.jinja to describe the new scoring rules and include the corresponding JSON keys in the output schema. The template enforces a strict JSON structure that must match your Pydantic model.

Add the new criterion section and update the JSON block:


# prompts/templates/resume_evaluation_criteria.jinja – add after Technical Skills

### Teamwork (0-20 points)

- Evaluate collaboration experience from work, projects, or open‑source contributions.
- Award higher scores for multi‑member projects, code reviews, and documented teamwork processes.

## JSON SECTION (add new key)

{
    "scores": {
        "open_source": {"score": 0, "max": 35, "evidence": ""},
        "self_projects": {"score": 0, "max": 30, "evidence": ""},
        "production": {"score": 0, "max": 25, "evidence": ""},
        "technical_skills": {"score": 0, "max": 10, "evidence": ""},
        "teamwork": {"score": 0, "max": 20, "evidence": ""}   # ← new entry

    },
    # … rest of the JSON unchanged …

}

Step 3: Register New Templates (Optional)

If you create a brand-new template file rather than editing the existing one, register its name in TemplateManager._load_templates (see lines 36-48 of template_manager.py). This makes the template discoverable via render_template.


# template_manager.py – add entry to the mapping in _load_templates

template_files = {
    # … existing entries …

    "resume_evaluation_criteria": "resume_evaluation_criteria.jinja",
    "resume_evaluation_extended": "resume_evaluation_extended.jinja",  # ← new name

}

You can then render the new template in ResumeEvaluator._load_evaluation_prompt:

criteria = self.template_manager.render_template(
    "resume_evaluation_extended", text_content=resume_text
)

Validation and Global Constraints

Any new score category must respect the global limits defined in evaluator.py (lines 9-12):

MAX_FINAL_SCORE = 100
MAX_BONUS_POINTS = 10

# … other constants …

The LLM must return only the JSON structure specified in the template (see the "CRITICAL REQUIREMENTS" block at the end of the Jinja file). After modifying the model or template, run the end-to-end pipeline to verify the integration:

$ python score.py example_resume.pdf

Inspect the printed JSON or the generated CSV (when DEVELOPMENT_MODE=True) to confirm that the new field appears in the output and that no Pydantic validation errors are raised.

Summary

  • Extend EvaluationData in models.py to define new fields for custom criteria, ensuring the LLM output can be parsed and validated.
  • Update resume_evaluation_criteria.jinja to include natural-language scoring instructions and the corresponding JSON schema entries.
  • Register new templates in TemplateManager._load_templates (lines 36-48) if creating separate template files for different evaluation contexts.
  • Respect global limits such as MAX_FINAL_SCORE and MAX_BONUS_POINTS defined in evaluator.py when designing scoring scales.
  • Validate changes by running python score.py <pdf> to verify the new criteria appear correctly in the output without validation errors.

Frequently Asked Questions

Where is the evaluation criteria defined in the Hiring Agent?

The evaluation criteria are defined in prompts/templates/resume_evaluation_criteria.jinja, which contains both the natural-language scoring rubric and the exact JSON schema the LLM must return. The ResumeEvaluator.evaluate_resume method (lines 41-44 of evaluator.py) loads this template via TemplateManager and sends it to the LLM for processing.

What happens if the LLM returns JSON that doesn't match the EvaluationData model?

The Hiring Agent uses Pydantic validation through the EvaluationData class in models.py. If the LLM returns JSON missing required fields or containing incorrect types, the validation will raise a Pydantic error, causing the evaluation to fail. This strict validation ensures data integrity and prevents downstream processing errors.

Can I create multiple evaluation templates for different roles?

Yes. Create separate Jinja files in prompts/templates for each role or evaluation context, then register each filename in TemplateManager._load_templates (lines 36-48 of template_manager.py). You can then instantiate ResumeEvaluator with different template names or modify the render_template call in ResumeEvaluator._load_evaluation_prompt to select the appropriate template based on the job requirements.

How do I enforce scoring limits for custom criteria?

Global scoring constraints are defined as constants in evaluator.py (lines 9-12), including MAX_FINAL_SCORE and MAX_BONUS_POINTS. When adding custom criteria in the Jinja template, ensure your maximum point values align with these global limits. The LLM prompt should explicitly state the point ranges (e.g., "0-20 points") to guide the model toward valid outputs that won't violate the global scoring constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →