# How Evaluation Metrics Are Defined in the Hiring-Agent: A Technical Deep Dive

> Understand how evaluation metrics are defined in the hiring-agent. Explore its three-layer architecture: Pydantic schemas, global constants, and Jinja2 rubric templates for LLM scoring.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-07-04

---

**The interviewstreet/hiring-agent defines evaluation metrics through a three-layer architecture: Pydantic schemas in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) that structure the data, global constants in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) that enforce hard limits, and a Jinja2 rubric template that instructs the LLM how to assign points across four mandatory categories.**

The interviewstreet/hiring-agent repository implements a deterministic, auditable scoring system for résumé evaluation. Its evaluation metrics are not merely prompt suggestions but are codified into strict data contracts, bounded mathematical constraints, and a transparent rubric that governs LLM behavior.

## Metric Schema and Data Structures (models.py)

At the foundation, the evaluation metrics are strictly typed using Pydantic models located in [[`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)](https://github.com/interviewstreet/hiring-agent/blob/main/models.py#L18-L50). These schemas ensure that every LLM response conforms to a predictable structure before any scoring logic is applied.

The atomic unit is the `CategoryScore` class, which captures the result for a single evaluation dimension:

```python
class CategoryScore(BaseModel):
    score: float = Field(ge=0, description="Score achieved in this category")
    max: int = Field(gt=0, description="Maximum possible score")
    evidence: str = Field(min_length=1, description="Evidence supporting the score")

```

Four instances of this model are composed into the `Scores` container, representing the mandatory evaluation dimensions:

```python
class Scores(BaseModel):
    open_source: CategoryScore      # Range: 0-35

    self_projects: CategoryScore    # Range: 0-30

    production: CategoryScore       # Range: 0-25

    technical_skills: CategoryScore # Range: 0-10

```

Additional metrics for adjustments are defined separately to isolate bonuses from deductions:

```python
class BonusPoints(BaseModel):
    total: float = Field(ge=0, le=20, description="Total bonus points")
    breakdown: str = Field(description="Breakdown of bonus points")

class Deductions(BaseModel):
    total: float = Field(ge=0, description="Total deduction points")
    reasons: str = Field(description="Reasons for deductions")

```

The top-level `EvaluationData` model aggregates these components and adds qualitative fields, creating the complete contract that the LLM must satisfy:

```python
class EvaluationData(BaseModel):
    scores: Scores
    bonus_points: BonusPoints
    deductions: Deductions
    key_strengths: List[str] = Field(min_items=1, max_items=5)
    areas_for_improvement: List[str] = Field(min_items=1, max_items=5)

```

## Global Scoring Constraints (evaluator.py)

Hard limits on the evaluation metrics are enforced by the `ResumeEvaluator` class in [[`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py#L24-L88). These constants prevent score manipulation and keep results within a calibrated range:

```python
MAX_BONUS_POINTS = 20
MIN_FINAL_SCORE = -20
MAX_FINAL_SCORE = 120

```

After parsing the LLM response into the Pydantic models, the evaluator applies these caps programmatically. If the model attempts to assign 25 bonus points, the system clamps the value to `MAX_BONUS_POINTS` (20). Similarly, the final aggregated score is bounded between -20 and 120, ensuring that extreme outlier judgments cannot skew the hiring pipeline.

## Prompt-Driven Rubric Definition (resume_evaluation_criteria.jinja)

While the code defines the structure, the semantic meaning of the evaluation metrics is controlled by the Jinja2 template located at [`prompts/templates/resume_evaluation_criteria.jinja`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/templates/resume_evaluation_criteria.jinja#L1-L66). This file serves as the authoritative rubric that instructs the LLM how to map résumé content to numerical scores.

The template mandates four specific categories with explicit maximums:

- **Open-source contributions**: 0–35 points
- **Self-projects**: 0–30 points  
- **Production experience**: 0–25 points
- **Technical skills**: 0–10 points

The rubric also embeds fairness constraints (lines 5–12) that instruct the model to ignore candidate names, gender, education institutions, and location. Bonus and deduction rules are specified with concrete examples, such as deducting 2–5 points for "simple tutorial projects" and capping total bonuses at 20 points (lines 16–64).

## The Evaluation Engine Workflow

The `ResumeEvaluator` orchestrates the transformation of a PDF into validated metrics through a six-stage pipeline:

1. **Text Extraction**: `PDFHandler` converts the résumé into plain text.
2. **Prompt Composition**: `TemplateManager` renders the Jinja2 rubric with the résumé content injected.
3. **LLM Invocation**: The evaluator initializes either an **Ollama** or **Gemini** provider via `_initialize_llm_provider()`, sending the system message and user prompt.
4. **JSON Extraction**: `llm_utils.extract_json_from_response` sanitizes the LLM output to isolate the JSON payload.
5. **Schema Validation**: `EvaluationData(**evaluation_dict)` parses and validates the JSON against the Pydantic schema; violations raise immediate errors.
6. **Post-Processing**: [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) computes the final tally using the formula `sum(category scores) + bonus - deductions`, applying the global caps from [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py).

## Practical Implementation Examples

### Running Evaluation from the Command Line

The CLI entry point in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) executes the full pipeline:

```bash
python score.py /path/to/resume.pdf

```

This command triggers PDF extraction, optional GitHub enrichment, LLM evaluation, and formatted console output via `print_evaluation_results`.

### Programmatic Evaluation in Python

For integration into larger workflows, instantiate the evaluator directly:

```python
from evaluator import ResumeEvaluator
from models import EvaluationData

# Initialize with default model from environment

evaluator = ResumeEvaluator()

# Evaluate raw résumé text

resume_text = """John Doe
Software Engineer
GitHub: https://github.com/johndoe
..."""

evaluation: EvaluationData = evaluator.evaluate_resume(resume_text)

# Access structured metrics

print(f"Open-source: {evaluation.scores.open_source.score}/{evaluation.scores.open_source.max}")
print(f"Evidence: {evaluation.scores.open_source.evidence}")

# Calculate final score manually

total = sum(c.score for c in evaluation.scores.model_dump().values())
total += evaluation.bonus_points.total
total -= evaluation.deductions.total
print(f"Final score (capped at 120): {min(total, 120)}")

```

### Inspecting Raw LLM Responses

To debug or audit the evaluation metrics before validation:

```python
from llm_utils import extract_json_from_response

raw_response = evaluator.provider.chat(
    model="gemma3:4b",
    messages=[...],
    options={"temperature": 0.2},
    format=EvaluationData.model_json_schema(),
)

json_str = extract_json_from_response(raw_response["message"]["content"])
print(json_str)  # Raw JSON string

evaluation = EvaluationData.parse_raw(json_str)  # Validated model

```

## Summary

- **Structured Schema**: The `CategoryScore`, `Scores`, and `EvaluationData` models in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) enforce type safety and required evidence fields for every metric.
- **Bounded Ranges**: Global constants `MAX_BONUS_POINTS` (20), `MIN_FINAL_SCORE` (-20), and `MAX_FINAL_SCORE` (120) prevent score inflation or manipulation.
- **Rubric-Driven**: The Jinja2 template `resume_evaluation_criteria.jinja` defines four mandatory categories with specific point ranges (open-source: 35, self-projects: 30, production: 25, technical-skills: 10) and fairness constraints.
- **Validation Pipeline**: The `ResumeEvaluator` class orchestrates extraction, LLM invocation, JSON parsing, and schema validation to ensure only conforming metrics enter the hiring workflow.

## Frequently Asked Questions

### What are the four mandatory evaluation categories in the hiring-agent?

The system evaluates every résumé against four fixed dimensions: **open-source contributions** (max 35 points), **self-projects** (max 30 points), **production experience** (max 25 points), and **technical skills** (max 10 points). These ranges are hard-coded in the `resume_evaluation_criteria.jinja` prompt template and enforced by the `Scores` Pydantic model.

### How does the system prevent the LLM from assigning arbitrary bonus points?

The `BonusPoints` Pydantic model enforces an upper bound of 20 points via `Field(ge=0, le=20)`, and the `ResumeEvaluator` applies the `MAX_BONUS_POINTS = 20` constant during post-processing. Even if the LLM attempts to exceed this limit in its response, the schema validation and subsequent clamping logic ensure the final value never exceeds 20.

### Can the evaluation metrics handle negative final scores?

Yes. The architecture supports negative outcomes through the `Deductions` model and the `MIN_FINAL_SCORE = -20` constant. If deductions exceed the sum of category scores and bonuses, the final calculation is clamped at -20, allowing the system to flag significantly unqualified candidates while maintaining a bounded scoring range.

### Which LLM providers are supported for generating evaluation metrics?

The `ResumeEvaluator` supports both local and cloud providers through abstraction classes defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py). It can utilize **Ollama** for local model hosting (e.g., Gemma, Llama) or **Gemini** for Google's API, selectable via environment configuration or runtime parameters in `_initialize_llm_provider()`.