How Evaluation Metrics Are Defined in the Hiring-Agent: A Technical Deep Dive

The interviewstreet/hiring-agent defines evaluation metrics through a three-layer architecture: Pydantic schemas in models.py that structure the data, global constants in evaluator.py that enforce hard limits, and a Jinja2 rubric template that instructs the LLM how to assign points across four mandatory categories.

The interviewstreet/hiring-agent repository implements a deterministic, auditable scoring system for résumé evaluation. Its evaluation metrics are not merely prompt suggestions but are codified into strict data contracts, bounded mathematical constraints, and a transparent rubric that governs LLM behavior.

Metric Schema and Data Structures (models.py)

At the foundation, the evaluation metrics are strictly typed using Pydantic models located in [models.py](https://github.com/interviewstreet/hiring-agent/blob/main/models.py#L18-L50). These schemas ensure that every LLM response conforms to a predictable structure before any scoring logic is applied.

The atomic unit is the CategoryScore class, which captures the result for a single evaluation dimension:

class CategoryScore(BaseModel):
    score: float = Field(ge=0, description="Score achieved in this category")
    max: int = Field(gt=0, description="Maximum possible score")
    evidence: str = Field(min_length=1, description="Evidence supporting the score")

Four instances of this model are composed into the Scores container, representing the mandatory evaluation dimensions:

class Scores(BaseModel):
    open_source: CategoryScore      # Range: 0-35

    self_projects: CategoryScore    # Range: 0-30

    production: CategoryScore       # Range: 0-25

    technical_skills: CategoryScore # Range: 0-10

Additional metrics for adjustments are defined separately to isolate bonuses from deductions:

class BonusPoints(BaseModel):
    total: float = Field(ge=0, le=20, description="Total bonus points")
    breakdown: str = Field(description="Breakdown of bonus points")

class Deductions(BaseModel):
    total: float = Field(ge=0, description="Total deduction points")
    reasons: str = Field(description="Reasons for deductions")

The top-level EvaluationData model aggregates these components and adds qualitative fields, creating the complete contract that the LLM must satisfy:

class EvaluationData(BaseModel):
    scores: Scores
    bonus_points: BonusPoints
    deductions: Deductions
    key_strengths: List[str] = Field(min_items=1, max_items=5)
    areas_for_improvement: List[str] = Field(min_items=1, max_items=5)

Global Scoring Constraints (evaluator.py)

Hard limits on the evaluation metrics are enforced by the ResumeEvaluator class in [evaluator.py](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py#L24-L88). These constants prevent score manipulation and keep results within a calibrated range:

MAX_BONUS_POINTS = 20
MIN_FINAL_SCORE = -20
MAX_FINAL_SCORE = 120

After parsing the LLM response into the Pydantic models, the evaluator applies these caps programmatically. If the model attempts to assign 25 bonus points, the system clamps the value to MAX_BONUS_POINTS (20). Similarly, the final aggregated score is bounded between -20 and 120, ensuring that extreme outlier judgments cannot skew the hiring pipeline.

Prompt-Driven Rubric Definition (resume_evaluation_criteria.jinja)

While the code defines the structure, the semantic meaning of the evaluation metrics is controlled by the Jinja2 template located at prompts/templates/resume_evaluation_criteria.jinja. This file serves as the authoritative rubric that instructs the LLM how to map résumé content to numerical scores.

The template mandates four specific categories with explicit maximums:

  • Open-source contributions: 0–35 points
  • Self-projects: 0–30 points
  • Production experience: 0–25 points
  • Technical skills: 0–10 points

The rubric also embeds fairness constraints (lines 5–12) that instruct the model to ignore candidate names, gender, education institutions, and location. Bonus and deduction rules are specified with concrete examples, such as deducting 2–5 points for "simple tutorial projects" and capping total bonuses at 20 points (lines 16–64).

The Evaluation Engine Workflow

The ResumeEvaluator orchestrates the transformation of a PDF into validated metrics through a six-stage pipeline:

  1. Text Extraction: PDFHandler converts the résumé into plain text.
  2. Prompt Composition: TemplateManager renders the Jinja2 rubric with the résumé content injected.
  3. LLM Invocation: The evaluator initializes either an Ollama or Gemini provider via _initialize_llm_provider(), sending the system message and user prompt.
  4. JSON Extraction: llm_utils.extract_json_from_response sanitizes the LLM output to isolate the JSON payload.
  5. Schema Validation: EvaluationData(**evaluation_dict) parses and validates the JSON against the Pydantic schema; violations raise immediate errors.
  6. Post-Processing: score.py computes the final tally using the formula sum(category scores) + bonus - deductions, applying the global caps from evaluator.py.

Practical Implementation Examples

Running Evaluation from the Command Line

The CLI entry point in score.py executes the full pipeline:

python score.py /path/to/resume.pdf

This command triggers PDF extraction, optional GitHub enrichment, LLM evaluation, and formatted console output via print_evaluation_results.

Programmatic Evaluation in Python

For integration into larger workflows, instantiate the evaluator directly:

from evaluator import ResumeEvaluator
from models import EvaluationData

# Initialize with default model from environment

evaluator = ResumeEvaluator()

# Evaluate raw résumé text

resume_text = """John Doe
Software Engineer
GitHub: https://github.com/johndoe
..."""

evaluation: EvaluationData = evaluator.evaluate_resume(resume_text)

# Access structured metrics

print(f"Open-source: {evaluation.scores.open_source.score}/{evaluation.scores.open_source.max}")
print(f"Evidence: {evaluation.scores.open_source.evidence}")

# Calculate final score manually

total = sum(c.score for c in evaluation.scores.model_dump().values())
total += evaluation.bonus_points.total
total -= evaluation.deductions.total
print(f"Final score (capped at 120): {min(total, 120)}")

Inspecting Raw LLM Responses

To debug or audit the evaluation metrics before validation:

from llm_utils import extract_json_from_response

raw_response = evaluator.provider.chat(
    model="gemma3:4b",
    messages=[...],
    options={"temperature": 0.2},
    format=EvaluationData.model_json_schema(),
)

json_str = extract_json_from_response(raw_response["message"]["content"])
print(json_str)  # Raw JSON string

evaluation = EvaluationData.parse_raw(json_str)  # Validated model

Summary

  • Structured Schema: The CategoryScore, Scores, and EvaluationData models in models.py enforce type safety and required evidence fields for every metric.
  • Bounded Ranges: Global constants MAX_BONUS_POINTS (20), MIN_FINAL_SCORE (-20), and MAX_FINAL_SCORE (120) prevent score inflation or manipulation.
  • Rubric-Driven: The Jinja2 template resume_evaluation_criteria.jinja defines four mandatory categories with specific point ranges (open-source: 35, self-projects: 30, production: 25, technical-skills: 10) and fairness constraints.
  • Validation Pipeline: The ResumeEvaluator class orchestrates extraction, LLM invocation, JSON parsing, and schema validation to ensure only conforming metrics enter the hiring workflow.

Frequently Asked Questions

What are the four mandatory evaluation categories in the hiring-agent?

The system evaluates every résumé against four fixed dimensions: open-source contributions (max 35 points), self-projects (max 30 points), production experience (max 25 points), and technical skills (max 10 points). These ranges are hard-coded in the resume_evaluation_criteria.jinja prompt template and enforced by the Scores Pydantic model.

How does the system prevent the LLM from assigning arbitrary bonus points?

The BonusPoints Pydantic model enforces an upper bound of 20 points via Field(ge=0, le=20), and the ResumeEvaluator applies the MAX_BONUS_POINTS = 20 constant during post-processing. Even if the LLM attempts to exceed this limit in its response, the schema validation and subsequent clamping logic ensure the final value never exceeds 20.

Can the evaluation metrics handle negative final scores?

Yes. The architecture supports negative outcomes through the Deductions model and the MIN_FINAL_SCORE = -20 constant. If deductions exceed the sum of category scores and bonuses, the final calculation is clamped at -20, allowing the system to flag significantly unqualified candidates while maintaining a bounded scoring range.

Which LLM providers are supported for generating evaluation metrics?

The ResumeEvaluator supports both local and cloud providers through abstraction classes defined in models.py. It can utilize Ollama for local model hosting (e.g., Gemma, Llama) or Gemini for Google's API, selectable via environment configuration or runtime parameters in _initialize_llm_provider().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →