# How the Hiring Agent Ensures Fair Resume Evaluation: A Technical Deep Dive

> Discover how the Hiring Agent ensures fair resume evaluation. Learn about LLM prompt policies, PII stripping, and technical merit validation for unbiased hiring.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-06-28

---

**The Hiring Agent enforces fair resume evaluation by embedding strict fairness policies directly into LLM prompts, stripping all personally identifiable information before scoring, and validating outputs against a rigid Pydantic schema that only recognizes technical merit categories.**

The interviewstreet/hiring-agent repository implements a policy-first architecture that eliminates demographic bias from technical hiring decisions. By combining Jinja2 templating with structured data validation, the system ensures that candidate assessments rely solely on objective evidence such as open-source contributions and production experience rather than protected attributes like name, gender, or university.

## Policy-Driven Prompt Templates for Bias Prevention

The foundation of fair evaluation resides in the `resume_evaluation_criteria.jinja` template located at `prompts/templates/resume_evaluation_criteria.jinja`. This file explicitly instructs the LLM to ignore **personally identifiable information** and demographic signals including candidate names, gender, location, school affiliation, and grades.

The template directs the model to score exclusively on technical dimensions: skills proficiency, project complexity, open-source contributions, production experience, and problem-solving ability. By baking these constraints into the prompt itself, the fairness policy becomes immutable for each evaluation cycle, preventing the LLM from accessing protected attributes during the scoring phase.

## Structured Output Validation with Pydantic Models

### The `EvaluationData` Schema in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)

The system mandates strict output conformity through the `EvaluationData` model defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py). This Pydantic schema enforces a machine-parseable JSON structure that eliminates free-form text responses capable of reintroducing bias.

The schema requires four compulsory score categories:

- `open_source` – Contributions to public repositories and community involvement
- `self_projects` – Complexity and technical depth of personal projects  
- `production` – Evidence of production-grade system experience
- `technical_skills` – Proficiency in relevant technologies and frameworks

Each category adheres to predefined numerical constraints, ensuring that scores reflect actual technical evidence rather than subjective interpretation.

### Automated Score Caps and Deduction Rules

The Jinja template embeds **hard limits** and deduction clauses that the LLM applies before returning the JSON payload. For example, the template specifies that the `open_source` score must never exceed 10 points when only personal repositories are present. These guardrails execute within the LLM context, creating a self-policing mechanism that standardizes evaluations across different candidates and model instances.

## Deterministic LLM Orchestration and Response Handling

### The `ResumeEvaluator` Class in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)

The `ResumeEvaluator` class orchestrates the entire fair evaluation pipeline in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py). It loads the fairness policy template, injects the anonymized resume text, and transmits requests through a unified `LLMProvider` interface supporting both Ollama and Gemini backends.

The provider initializes once via `_initialize_llm_provider` and maintains consistent configuration across all evaluations. This deterministic instantiation prevents configuration drift that could inadvertently alter scoring behavior between different candidate assessments.

### JSON Extraction via `extract_json_from_response`

After receiving the LLM response, the system passes the raw output through `extract_json_from_response` in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py). This utility strips surrounding prose and conversational text, isolating only the validated JSON structure. By enforcing this extraction layer, the pipeline guarantees that downstream logic processes strictly structured data, eliminating any narrative content that might contain demographic inferences.

## Complete Implementation Example

The following implementation demonstrates how to execute a bias-free evaluation using the Hiring Agent's Python API:

```python
from evaluator import ResumeEvaluator

# Example raw resume text (PII will be ignored by the policy template)

resume_text = """
John Doe
Software Engineer
MIT Computer Science, GPA 3.9
San Francisco, CA

Experience:
- Built a distributed caching system handling 10k req/s
- Contributed to kubernetes/kubernetes (merged 12 PRs)
- Developed React frontend for fintech platform serving 1M users
"""

# Initialize the evaluator with configured LLM provider

evaluator = ResumeEvaluator()

# Execute evaluation - returns validated EvaluationData instance

evaluation = evaluator.evaluate_resume(resume_text)

# Access technical merit scores only

print("Open-source score:", evaluation.scores.open_source.score)
print("Self-project score:", evaluation.scores.self_projects.score)
print("Production experience:", evaluation.scores.production.score)
print("Technical-skills score:", evaluation.scores.technical_skills.score)

# Review bonus points and deductions applied by policy rules

print("Bonus points:", evaluation.bonus_points.total)
print("Deductions:", evaluation.deductions.total)

# Human-readable assessment summary

print("Key strengths:", evaluation.key_strengths)
print("Areas for improvement:", evaluation.areas_for_improvement)

```

This workflow transforms unstructured resume text into an objective assessment while ensuring that the `ResumeEvaluator` never logs or stores protected demographic fields for scoring purposes.

## Summary

- **Policy-first architecture**: Fairness rules live in `resume_evaluation_criteria.jinja`, explicitly prohibiting demographic signals from influencing scores.
- **Structured validation**: The `EvaluationData` schema in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) enforces JSON output containing only four technical merit categories.
- **Deterministic orchestration**: The `ResumeEvaluator` class in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) provides consistent LLM provider handling and response processing.
- **Secure extraction**: `extract_json_from_response` in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) isolates structured data from conversational LLM output.
- **Protected attribute isolation**: The system evaluates candidates solely on open-source contributions, production experience, project complexity, and technical skills.

## Frequently Asked Questions

### What specific demographic signals does the Hiring Agent remove from evaluation?

The `resume_evaluation_criteria.jinja` template explicitly instructs the LLM to disregard names, gender indicators, geographic locations, university affiliations, and GPA scores. The evaluation focuses exclusively on technical artifacts such as code repositories, system architecture descriptions, and problem-solving methodologies.

### How does the system prevent the LLM from hallucinating scores outside the defined rubric?

The template embeds hard score caps and automatic deduction rules that the LLM must apply before generating the JSON response. Combined with the strict `EvaluationData` Pydantic model in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), these constraints validate that numerical scores fall within acceptable ranges and correspond to actual evidence found in the resume.

### Can the fairness policy be customized without modifying the Python source code?

Yes. Since the core fairness rules reside in the Jinja template at `prompts/templates/resume_evaluation_criteria.jinja`, organizations can adjust scoring criteria, modify technical merit weightings, or update deduction clauses by editing this template file. The `ResumeEvaluator` dynamically loads this template at runtime, allowing policy changes without redeploying the Python application.

### Which LLM providers does the Hiring Agent support for fair resume evaluation?

The `LLMProvider` interface abstraction in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) supports both **Ollama** for local model hosting and **Gemini** for cloud-based inference. The `_initialize_llm_provider` method instantiates the selected provider once and reuses it across all evaluations, ensuring consistent behavior regardless of the underlying model.