How Evaluation Metrics Are Defined in the Hiring-Agent: A Technical Deep Dive
The interviewstreet/hiring-agent defines evaluation metrics through a three-layer architecture: Pydantic schemas in models.py that structure the data, global constants in evaluator.py that enforce hard limits, and a Jinja2 rubric template that instructs the LLM how to assign points across four mandatory categories.
The interviewstreet/hiring-agent repository implements a deterministic, auditable scoring system for résumé evaluation. Its evaluation metrics are not merely prompt suggestions but are codified into strict data contracts, bounded mathematical constraints, and a transparent rubric that governs LLM behavior.
Metric Schema and Data Structures (models.py)
At the foundation, the evaluation metrics are strictly typed using Pydantic models located in [models.py](https://github.com/interviewstreet/hiring-agent/blob/main/models.py#L18-L50). These schemas ensure that every LLM response conforms to a predictable structure before any scoring logic is applied.
The atomic unit is the CategoryScore class, which captures the result for a single evaluation dimension:
class CategoryScore(BaseModel):
score: float = Field(ge=0, description="Score achieved in this category")
max: int = Field(gt=0, description="Maximum possible score")
evidence: str = Field(min_length=1, description="Evidence supporting the score")
Four instances of this model are composed into the Scores container, representing the mandatory evaluation dimensions:
class Scores(BaseModel):
open_source: CategoryScore # Range: 0-35
self_projects: CategoryScore # Range: 0-30
production: CategoryScore # Range: 0-25
technical_skills: CategoryScore # Range: 0-10
Additional metrics for adjustments are defined separately to isolate bonuses from deductions:
class BonusPoints(BaseModel):
total: float = Field(ge=0, le=20, description="Total bonus points")
breakdown: str = Field(description="Breakdown of bonus points")
class Deductions(BaseModel):
total: float = Field(ge=0, description="Total deduction points")
reasons: str = Field(description="Reasons for deductions")
The top-level EvaluationData model aggregates these components and adds qualitative fields, creating the complete contract that the LLM must satisfy:
class EvaluationData(BaseModel):
scores: Scores
bonus_points: BonusPoints
deductions: Deductions
key_strengths: List[str] = Field(min_items=1, max_items=5)
areas_for_improvement: List[str] = Field(min_items=1, max_items=5)
Global Scoring Constraints (evaluator.py)
Hard limits on the evaluation metrics are enforced by the ResumeEvaluator class in [evaluator.py](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py#L24-L88). These constants prevent score manipulation and keep results within a calibrated range:
MAX_BONUS_POINTS = 20
MIN_FINAL_SCORE = -20
MAX_FINAL_SCORE = 120
After parsing the LLM response into the Pydantic models, the evaluator applies these caps programmatically. If the model attempts to assign 25 bonus points, the system clamps the value to MAX_BONUS_POINTS (20). Similarly, the final aggregated score is bounded between -20 and 120, ensuring that extreme outlier judgments cannot skew the hiring pipeline.
Prompt-Driven Rubric Definition (resume_evaluation_criteria.jinja)
While the code defines the structure, the semantic meaning of the evaluation metrics is controlled by the Jinja2 template located at prompts/templates/resume_evaluation_criteria.jinja. This file serves as the authoritative rubric that instructs the LLM how to map résumé content to numerical scores.
The template mandates four specific categories with explicit maximums:
- Open-source contributions: 0–35 points
- Self-projects: 0–30 points
- Production experience: 0–25 points
- Technical skills: 0–10 points
The rubric also embeds fairness constraints (lines 5–12) that instruct the model to ignore candidate names, gender, education institutions, and location. Bonus and deduction rules are specified with concrete examples, such as deducting 2–5 points for "simple tutorial projects" and capping total bonuses at 20 points (lines 16–64).
The Evaluation Engine Workflow
The ResumeEvaluator orchestrates the transformation of a PDF into validated metrics through a six-stage pipeline:
- Text Extraction:
PDFHandlerconverts the résumé into plain text. - Prompt Composition:
TemplateManagerrenders the Jinja2 rubric with the résumé content injected. - LLM Invocation: The evaluator initializes either an Ollama or Gemini provider via
_initialize_llm_provider(), sending the system message and user prompt. - JSON Extraction:
llm_utils.extract_json_from_responsesanitizes the LLM output to isolate the JSON payload. - Schema Validation:
EvaluationData(**evaluation_dict)parses and validates the JSON against the Pydantic schema; violations raise immediate errors. - Post-Processing:
score.pycomputes the final tally using the formulasum(category scores) + bonus - deductions, applying the global caps fromevaluator.py.
Practical Implementation Examples
Running Evaluation from the Command Line
The CLI entry point in score.py executes the full pipeline:
python score.py /path/to/resume.pdf
This command triggers PDF extraction, optional GitHub enrichment, LLM evaluation, and formatted console output via print_evaluation_results.
Programmatic Evaluation in Python
For integration into larger workflows, instantiate the evaluator directly:
from evaluator import ResumeEvaluator
from models import EvaluationData
# Initialize with default model from environment
evaluator = ResumeEvaluator()
# Evaluate raw résumé text
resume_text = """John Doe
Software Engineer
GitHub: https://github.com/johndoe
..."""
evaluation: EvaluationData = evaluator.evaluate_resume(resume_text)
# Access structured metrics
print(f"Open-source: {evaluation.scores.open_source.score}/{evaluation.scores.open_source.max}")
print(f"Evidence: {evaluation.scores.open_source.evidence}")
# Calculate final score manually
total = sum(c.score for c in evaluation.scores.model_dump().values())
total += evaluation.bonus_points.total
total -= evaluation.deductions.total
print(f"Final score (capped at 120): {min(total, 120)}")
Inspecting Raw LLM Responses
To debug or audit the evaluation metrics before validation:
from llm_utils import extract_json_from_response
raw_response = evaluator.provider.chat(
model="gemma3:4b",
messages=[...],
options={"temperature": 0.2},
format=EvaluationData.model_json_schema(),
)
json_str = extract_json_from_response(raw_response["message"]["content"])
print(json_str) # Raw JSON string
evaluation = EvaluationData.parse_raw(json_str) # Validated model
Summary
- Structured Schema: The
CategoryScore,Scores, andEvaluationDatamodels inmodels.pyenforce type safety and required evidence fields for every metric. - Bounded Ranges: Global constants
MAX_BONUS_POINTS(20),MIN_FINAL_SCORE(-20), andMAX_FINAL_SCORE(120) prevent score inflation or manipulation. - Rubric-Driven: The Jinja2 template
resume_evaluation_criteria.jinjadefines four mandatory categories with specific point ranges (open-source: 35, self-projects: 30, production: 25, technical-skills: 10) and fairness constraints. - Validation Pipeline: The
ResumeEvaluatorclass orchestrates extraction, LLM invocation, JSON parsing, and schema validation to ensure only conforming metrics enter the hiring workflow.
Frequently Asked Questions
What are the four mandatory evaluation categories in the hiring-agent?
The system evaluates every résumé against four fixed dimensions: open-source contributions (max 35 points), self-projects (max 30 points), production experience (max 25 points), and technical skills (max 10 points). These ranges are hard-coded in the resume_evaluation_criteria.jinja prompt template and enforced by the Scores Pydantic model.
How does the system prevent the LLM from assigning arbitrary bonus points?
The BonusPoints Pydantic model enforces an upper bound of 20 points via Field(ge=0, le=20), and the ResumeEvaluator applies the MAX_BONUS_POINTS = 20 constant during post-processing. Even if the LLM attempts to exceed this limit in its response, the schema validation and subsequent clamping logic ensure the final value never exceeds 20.
Can the evaluation metrics handle negative final scores?
Yes. The architecture supports negative outcomes through the Deductions model and the MIN_FINAL_SCORE = -20 constant. If deductions exceed the sum of category scores and bonuses, the final calculation is clamped at -20, allowing the system to flag significantly unqualified candidates while maintaining a bounded scoring range.
Which LLM providers are supported for generating evaluation metrics?
The ResumeEvaluator supports both local and cloud providers through abstraction classes defined in models.py. It can utilize Ollama for local model hosting (e.g., Gemma, Llama) or Gemini for Google's API, selectable via environment configuration or runtime parameters in _initialize_llm_provider().
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →