How to Extend Evaluation Criteria in Hiring-Agent: Adding Custom Scoring Rules

You can extend the evaluation criteria by editing the resume_evaluation_criteria.jinja template to describe new rules to the LLM, and optionally updating the EvaluationData Pydantic model in models.py if you require new top-level score fields.

The interviewstreet/hiring-agent repository uses a prompt-driven scoring architecture where the ResumeEvaluator class processes candidate resumes through a structured Jinja template. Extending evaluation criteria requires modifying the LLM instructions in the template file and, for new scoring dimensions, adjusting the validation schema to match the expected JSON output.

Understanding the Evaluation Architecture

The hiring-agent pipeline relies on three core components working in sequence. First, the Jinja template at prompts/templates/resume_evaluation_criteria.jinja contains the human-readable scoring instructions that guide the LLM. Second, the Pydantic models in models.py (specifically EvaluationData and Scores) validate the structured JSON response returned by the LLM. Third, the evaluator constants in evaluator.py define scoring limits such as MAX_BONUS_POINTS and MAX_FINAL_SCORE that constrain the final calculation.

The ResumeEvaluator.evaluate_resume() method orchestrates this flow by rendering the template, sending it to the LLM, and deserializing the response into the EvaluationData model. This design separates the scoring logic (defined in prompts) from the data validation (defined in Python), making the system highly adaptable.

Three Levels of Extension

You can customize scoring rules at three distinct levels depending on your requirements.

Prompt Level: New Rules in the Jinja Template

The most straightforward approach involves editing the Jinja template to append new rule blocks. The template already contains sections for Open Source, Self Projects, Production, and Technical Skills with specific bonus and deduction criteria. By adding new category descriptions—such as "Community Engagement" or "Technical Writing"—you provide the LLM with instructions to generate additional scores without modifying any Python code. The LLM will include these new evaluations in its JSON response as long as the instructions are clear.

Model Level: New Pydantic Fields

If your new rule introduces a new top-level score category rather than modifying existing logic, you must extend the Scores model in models.py. Add a new CategoryScore field to capture the structured output. The ResumeEvaluator automatically deserializes the LLM response into this updated schema, populating the new field without requiring changes to the evaluation logic itself.

Constants Level: Scoring Limits and Post-Processing

When new rules change the maximum achievable points, adjust the constants defined in evaluator.py. Values such as MAX_BONUS_POINTS, MIN_FINAL_SCORE, and MAX_FINAL_SCORE constrain the final calculation. Update these thresholds to maintain coherent scoring boundaries. Downstream components like score.py and CSV export utilities will automatically respect these new limits.

Practical Example: Adding a Community Engagement Category

Follow this concrete example to add a "Community Engagement" category worth up to 15 points.

Step 1: Edit the Jinja Template

Open prompts/templates/resume_evaluation_criteria.jinja and append a new section following the existing format.


## Community Engagement (0-15 points)

- Evaluate participation in developer communities (e.g., Stack Overflow, Reddit, Discord).
- **HIGH SCORES (12-15 points):** Regular contributions (≥ 50 answers/up-votes), moderation roles, or leadership in open-source communities.
- **MEDIUM SCORES (6-11 points):** Occasional helpful answers, speaking at meet-ups, or active membership in relevant forums.
- **LOW SCORES (0-5 points):** Minimal or no community activity.

### BONUS POINTS

- +2 points for mentoring junior developers (verified via LinkedIn recommendations or personal statements).

This block instructs the LLM to generate a community_engagement entry in the JSON output.

Step 2: Update the Pydantic Schema (Optional)

If you need community_engagement as a distinct top-level field in the results, modify models.py to extend the Scores class.

from pydantic import BaseModel

class CategoryScore(BaseModel):
    score: int
    max: int
    evidence: str

class Scores(BaseModel):
    open_source: CategoryScore
    self_projects: CategoryScore
    production: CategoryScore
    technical_skills: CategoryScore
    community_engagement: CategoryScore  # New field added

The ResumeEvaluator automatically populates this field when parsing the LLM response.

Step 3: Adjust Scoring Limits (Optional)

If the new category increases the theoretical maximum score, update the constants in evaluator.py.


# evaluator.py

MAX_BONUS_POINTS = 25  # Increased from 20 to accommodate community bonuses

MAX_FINAL_SCORE = 135  # Adjust if the new category pushes the total maximum higher

Step 4: Test the Pipeline

Run the evaluation pipeline to verify the integration works end-to-end.

python score.py examples/resume.pdf

Verify that the output JSON contains the new field with valid data:

{
  "scores": {
    "open_source": {"score": 30, "max": 35, "evidence": "..."},
    "self_projects": {"score": 22, "max": 30, "evidence": "..."},
    "production": {"score": 18, "max": 25, "evidence": "..."},
    "technical_skills": {"score": 9, "max": 10, "evidence": "..."},
    "community_engagement": {"score": 13, "max": 15, "evidence": "Active Stack Overflow contributor with 500+ reputation"}
  },
  "bonus_points": {"total": 6, "breakdown": "..."},
  "deductions": {"total": 2, "reasons": "..."}
}

If the field is missing or null, verify that the Jinja template indentation is correct and that the JSON schema in models.py matches the field name expected by the prompt.

Summary

  • Modify the Jinja template at prompts/templates/resume_evaluation_criteria.jinja to add new scoring rules for the LLM to evaluate.
  • Extend the Pydantic model in models.py only when introducing new top-level score categories that require structured storage.
  • Update constants in evaluator.py such as MAX_BONUS_POINTS and MAX_FINAL_SCORE if new rules alter the scoring boundaries.
  • Test with score.py to ensure the LLM returns valid JSON matching your updated schema before deploying changes.

Frequently Asked Questions

Do I need to modify Python code to add new scoring rules?

Not necessarily. If you are adjusting existing categories or adding bonus/deduction rules that fit within the current Scores structure, you only need to edit the Jinja template. The LLM will follow your new instructions and return the appropriate JSON. You only need to modify models.py when adding entirely new top-level score fields that do not exist in the current schema.

How does the ResumeEvaluator handle new fields in the LLM response?

The ResumeEvaluator.evaluate_resume() method uses the EvaluationData Pydantic model to parse and validate the LLM's JSON output. If you add a field to the Scores model in models.py, Pydantic automatically extracts and validates the corresponding data from the response. The evaluator requires no additional logic changes because it treats the model as the source of truth for the expected structure.

What happens if the LLM returns a score outside the defined range?

The Pydantic CategoryScore model validates that scores are integers, but the range constraints (e.g., 0-15 points) are enforced primarily through the prompt instructions in the Jinja template. If the LLM returns a value exceeding your specified maximum, the current implementation in evaluator.py may still process it unless you add explicit validation logic. For strict enforcement, add Pydantic validators to the CategoryScore class or clamp values during post-processing in the evaluator.

Can I add deduction rules instead of bonus points using the same method?

Yes. The Jinja template already contains a deductions section. You can append new deduction rules—such as penalties for employment gaps or missing documentation—directly to the template. Ensure you update the deductions field structure in the response handling if you introduce new deduction categories beyond the simple total, or describe them clearly in the prompt so the LLM includes them in the existing deductions breakdown.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →