# How Hiring Agent Ensures Fairness in Candidate Evaluation: Technical Architecture Explained

> Discover how Hiring Agent technical architecture ensures fairness in candidate evaluation. Learn about prompt constraints, JSON validation, and runtime checks to eliminate bias.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: architecture
- Published: 2026-07-02

---

**Hiring Agent enforces fairness through prompt-driven constraints, structured JSON output validation, and strict runtime checks that explicitly exclude protected attributes from scoring decisions.**

The interviewstreet/hiring-agent repository implements a multi-layered fairness architecture designed to evaluate software engineering candidates based solely on technical merit. By combining declarative prompt templates with rigid output schemas, the system ensures that protected characteristics never influence scoring while maintaining full auditability of every evaluation decision.

## Fairness Through Declarative Prompt Templates

### Explicit Exclusion of Protected Attributes

In `prompts/templates/resume_evaluation_system_message.jinja`, the system message explicitly lists every factor that **must not** influence a score. The template prohibits consideration of name, gender, school, GPA, location, and other demographic identifiers, restricting the LLM to only four admissible scoring dimensions: open-source contributions, self-projects, production experience, and technical skills.

### Normalized Scoring Criteria

The `prompts/templates/resume_evaluation_criteria.jinja` file codifies rigid definitions that prevent subjective interpretation. For example, the template specifies that "personal GitHub repos do NOT count as open-source contributions" and that "simple tutorial projects receive zero points." This declarative approach ensures that **no procedural code** can override fairness constraints, making the evaluation logic transparent and auditable.

## Structured Output Enforcement via Schema Validation

The `ResumeEvaluator` class in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) orchestrates the evaluation pipeline by loading these templates and constructing a chat payload containing both the system message (fairness constraints) and user message (resume content). According to the source code in lines 76-78, the LLM call includes a `format` argument derived from `EvaluationData.model_json_schema()`, which forces the model to respond only with the predefined JSON schema defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py).

This eliminates free-form narrative responses that could unintentionally expose bias or hallucinate irrelevant personal details. The Pydantic model acts as a strict contract that the LLM must follow.

## Runtime Validation and Audit Logging

After receiving the LLM response, the code extracts the JSON, logs the raw text for audit purposes, and parses it into the `EvaluationData` Pydantic model (lines 80-87 in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)). Any deviation from the schema raises a validation exception, ensuring the evaluator never returns malformed or potentially biased results. This runtime validation layer guarantees that every evaluation adheres to the fairness constraints defined in the prompt templates.

## Implementation: Running a Fairness-Compliant Evaluation

To execute an evaluation that respects these fairness guardrails:

```python
from evaluator import ResumeEvaluator

# Sample resume text (plain string)

resume_text = """
Senior Software Engineer
Experience: 5 years Python, 3 years distributed systems
GitHub: github.com/example contributor to kubernetes-sigs
"""

evaluator = ResumeEvaluator()  # uses DEFAULT_MODEL and built-in parameters

evaluation = evaluator.evaluate_resume(resume_text)

print(evaluation.json(indent=2))  # JSON adheres to the strict schema

```

To inspect the fairness constraints being injected into the prompt:

```python
e = ResumeEvaluator()
prompt = e._load_evaluation_prompt(resume_text)  # internal method

print(prompt)  # contains the "CRITICAL FAIRNESS REQUIREMENTS" block

```

The `ResumeEvaluator` relies on [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) to initialize the provider, ensuring that regardless of which LLM backend is selected, the same fairness constraints apply.

## Summary

- **Explicit prompt constraints** in Jinja templates exclude protected attributes (name, gender, school, GPA) from consideration
- **Declarative scoring rules** in `resume_evaluation_criteria.jinja` standardize what counts as legitimate technical experience
- **JSON schema enforcement** via `EvaluationData.model_json_schema()` prevents free-form biased outputs
- **Runtime Pydantic validation** in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) ensures only structured, schema-compliant responses are returned
- **Audit logging** captures raw LLM outputs for transparency and compliance review

## Frequently Asked Questions

### How does Hiring Agent prevent bias based on candidate names or schools?

The `resume_evaluation_system_message.jinja` template explicitly enumerates prohibited evaluation factors including name, gender, school, GPA, and location. These constraints are injected into every LLM call by the `ResumeEvaluator` class, ensuring the model treats these attributes as invisible during scoring.

### What happens if the LLM returns biased or malformed output?

The system implements strict runtime validation in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) (lines 80-87). After receiving the raw LLM response, the code attempts to parse it into the `EvaluationData` Pydantic model. Any schema violation or malformed JSON triggers a validation exception, preventing biased or corrupted evaluations from reaching the final output.

### Can the fairness constraints be customized or overridden?

While the scoring criteria are declarative and defined in Jinja templates, modifying them requires changing the template files themselves (`resume_evaluation_system_message.jinja` and `resume_evaluation_criteria.jinja`). The architecture intentionally prevents procedural overrides in the Python code to maintain auditability and prevent ad-hoc bias introduction.

### Which technical factors does Hiring Agent actually evaluate?

According to the criteria templates, the system evaluates four specific dimensions: open-source contributions (excluding personal repos), self-directed projects (excluding tutorials), production software experience, and demonstrated technical skills. This narrow scope ensures assessments focus exclusively on verifiable engineering capability.