How to Customize the Evaluation Criteria and Scoring Weights in Hiring Agent

You can fully customize the evaluation criteria and scoring weights in Hiring Agent by editing the Jinja rubric templates and updating the corresponding Python constants in the source code.

The Hiring Agent repository from interviewstreet provides a flexible, open-source framework for evaluating technical candidates. Because the scoring logic relies on plain-text prompt templates and modular Python constants rather than hard-coded rules, you can easily customize the evaluation criteria and scoring weights to match your organization's specific hiring standards.

Understanding the Scoring Architecture

Hiring Agent uses a two-layer approach to define scoring behavior: prompt-driven rubrics that instruct the LLM how to evaluate candidates, and Python constants that enforce mathematical constraints on the final output.

Prompt-Driven Rubric Definition

The evaluation rubric lives as a plain-text Jinja template in prompts/templates/resume_evaluation_criteria.jinja. This file contains the human-readable criteria that the LLM uses to assign scores, including explicit maximum point values for each category. The system-level instructions that wrap this rubric are stored in prompts/templates/resume_evaluation_system_message.jinja.

Because the LLM receives the rubric as rendered text, changing the numbers in the template immediately alters how the model assigns points. No hidden logic governs the weights elsewhere in the codebase.

Python Constants and Validation

The evaluator.py file defines global constraints that cap the final scores. The CategoryScore model in models.py declares a max field that validates the values returned by the LLM, ensuring they align with the template-defined limits.

Step-by-Step Customization Guide

Follow these steps to modify the evaluation criteria and scoring weights to fit your specific requirements.

1. Adjust the Rubric Template

Edit prompts/templates/resume_evaluation_criteria.jinja to change category weights, add evaluation rules, or remove sections. The template defines four default categories with the following maximum values:

  • Open Source: 35 points
  • Self Projects: 30 points
  • Production: 25 points
  • Technical Skills: 10 points

To re-weight a category, simply edit the heading line. For example, to reduce the Open Source weight from 35 to 20 points, change the template heading:


### Open Source (0-20 points)

You can also modify the bullet-point rules beneath each heading to change what evidence qualifies for high, medium, or low scores.

2. Synchronize the Python Model

When you change category definitions or add new ones, ensure the CategoryScore model in models.py can accommodate the updated structure. The max field in CategoryScore stores the maximum points per category, which downstream code in score.py uses for display logic.

If you add a new category, extend the Scores model in models.py with a new field:

class Scores(BaseModel):
    open_source: CategoryScore
    self_projects: CategoryScore
    production: CategoryScore
    technical_skills: CategoryScore
    leadership: CategoryScore  # New category added

3. Update Global Score Caps

Control the overall score boundaries and bonus-point limits by editing the constants at the top of evaluator.py:

MAX_BONUS_POINTS = 20       # Maximum bonus points allowed

MIN_FINAL_SCORE = 0         # Floor for final aggregate score

MAX_FINAL_SCORE = 100       # Ceiling for final aggregate score

Changing MAX_BONUS_POINTS allows you to increase or decrease the impact of bonus points on the final evaluation.

4. Add New Scoring Dimensions

To introduce entirely new evaluation criteria—such as "Leadership" or "Communication"—you must update three components:

  1. Template: Add a new section in resume_evaluation_criteria.jinja with the point range and evaluation rules.

  2. Model: Extend the Scores class in models.py as shown above.

  3. Presentation: Update the print_evaluation_results function in score.py to display the new category:

if hasattr(evaluation.scores, "leadership") and evaluation.scores.leadership:
    lead_score = evaluation.scores.leadership
    capped = min(lead_score.score, 15)  # Enforce the max from template

    print(f"🌟 Leadership:          {capped}/{lead_score.max}")
    print(f"   Evidence: {lead_score.evidence}\n")

The TemplateManager.render_template method in prompts/template_manager.py loads the updated template automatically on the next evaluation run, so no additional deployment steps are required.

Practical Code Examples

Example 1: Reducing Open Source Weight

Change the maximum points for Open Source contributions from 35 to 20:


### Open Source (0-20 points)

Update the sub-rules in the template to reflect the compressed scale, and ensure any validation logic in score.py referencing the old maximum is adjusted accordingly.

Example 2: Adding a Leadership Category

Add this section to resume_evaluation_criteria.jinja:


### Leadership (0-15 points)

**HIGH SCORES (10-15 points):**
- Led a team of 5+ engineers on a production-grade project
- Demonstrated measurable impact on project success

**MEDIUM SCORES (5-9 points):**
- Managed small teams or mentored interns
- Coordinated cross-functional efforts

**LOW SCORES (1-4 points):**
- Occasional leadership in class projects or hackathons

Then update models.py to include the new field in the Scores class, and add the corresponding display logic in score.py.

Example 3: Increasing the Bonus Point Cap

To allow up to 30 bonus points instead of the default 20, edit evaluator.py:

MAX_BONUS_POINTS = 30

This change immediately affects all future evaluations without requiring template modifications.

Summary

  • Edit resume_evaluation_criteria.jinja to change category weights and evaluation rules.
  • Update models.py when adding new categories to ensure the Scores model captures all dimensions.
  • Modify constants in evaluator.py to adjust global caps like MAX_BONUS_POINTS and MAX_FINAL_SCORE.
  • Extend score.py to display new categories in the evaluation output.
  • No hidden logic exists beyond these templates and constants; the LLM receives the rubric as plain text and generates scores accordingly.

Frequently Asked Questions

Can I completely remove a scoring category from the evaluation?

Yes. Remove the section from prompts/templates/resume_evaluation_criteria.jinja and delete the corresponding field from the Scores model in models.py. If you leave the field in the model but remove it from the template, the LLM will not populate it, but your code should handle optional fields using hasattr checks as shown in score.py.

Do I need to restart the application after editing the templates?

No. The TemplateManager.render_template method reads the Jinja files from disk on each evaluation run, so changes take effect immediately. If you are running a cached development mode, clear the cache to ensure the new template renders.

Where are the default point values (35, 30, 25, 10) defined?

These values exist only in prompts/templates/resume_evaluation_criteria.jinja as part of the heading text (e.g., ### Open Source (0-35 points)). The CategoryScore model in models.py stores the max value returned by the LLM, but the authoritative source is the template text itself.

Can I customize the evaluation criteria without modifying the source code?

No. You must edit the source files—specifically the Jinja templates in prompts/templates/ and potentially the constants in evaluator.py—to customize the criteria. There is no external configuration file or admin panel for adjusting weights at runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →