# How hiring-agent Ensures Fairness in Resume Evaluations: A Technical Deep Dive

> Explore how hiring-agent ensures fairness in resume evaluations using prompt constraints, structured JSON, and runtime validation to eliminate bias from protected attributes.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-07-04

---

**Hiring-agent enforces fairness through prompt-driven constraints, structured JSON output, and strict runtime validation that together eliminate bias from protected attributes.**

The interviewstreet/hiring-agent repository provides an automated resume evaluation system designed to assess candidates based solely on technical merit. By embedding fairness constraints directly into Jinja2 prompt templates and enforcing strict output schemas through Pydantic validation, the system guarantees that protected attributes never influence scoring decisions.

## Prompt-Driven Fairness Constraints

The foundation of hiring-agent fairness lies in declarative prompt templates that explicitly prohibit biased evaluation criteria.

### System Message Templates

In `prompts/templates/resume_evaluation_system_message.jinja` and `prompts/templates/resume_evaluation_criteria.jinja`, the system encodes **Critical Fairness Requirements** that list every factor prohibited from influencing scores. These include:

- Name, gender, or personal identifiers
- School name, GPA, or academic institution
- Geographic location or address
- Demographic indicators

The templates define the only admissible scoring dimensions: **open-source contributions**, **self-projects**, **production experience**, and **technical skills**. By declaratively encoding these constraints in the system prompt, the architecture ensures that no procedural code can override fairness rules during evaluation.

### Declarative Scoring Rules

The criteria template further eliminates ambiguity by specifying precise scoring boundaries. For example, the template explicitly states that personal GitHub repositories do **not** count as open-source contributions, and simple tutorial projects receive zero points. This declarative approach makes the evaluation logic transparent and auditable, preventing the LLM from inferring merit based on resume formatting or candidate background.

## Structured Output Enforcement

The `ResumeEvaluator` class in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) (lines 24-66) orchestrates the evaluation pipeline by loading fairness templates and constructing the LLM request payload. To prevent free-form narrative responses that could inadvertently expose bias, the system enforces strict output schemas.

When calling the LLM provider, the code passes a `format` argument derived from `EvaluationData.model_json_schema()`:

```python

# evaluator.py lines 76-78

response = self.provider.complete(
    messages,
    format=EvaluationData.model_json_schema()  # Enforces strict JSON schema

)

```

This constraint forces the model to return only the predefined JSON structure defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), eliminating any unstructured text that might contain subjective commentary about protected attributes.

## Runtime Validation and Audit Logging

After receiving the LLM response, the system implements multiple validation layers to ensure hiring-agent fairness persists through the final output.

In [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) (lines 80-87), the code extracts the raw JSON, logs the complete text for audit purposes, and parses it into the `EvaluationData` Pydantic model:

```python

# evaluator.py lines 80-87

raw_text = response.content.strip()
logger.info(f"Raw evaluation response: {raw_text}")  # Audit trail

try:
    data = json.loads(raw_text)
    evaluation = EvaluationData(**data)  # Strict schema validation

except ValidationError as e:
    raise RuntimeError(f"Invalid evaluation format: {e}")

```

Any deviation from the expected schema raises a `ValidationError`, ensuring the evaluator never returns malformed or potentially biased results. The logging mechanism creates a transparent audit trail for compliance review.

## Implementation Example

The following example demonstrates how to run a fairness-compliant evaluation using the `ResumeEvaluator` class:

```python
from evaluator import ResumeEvaluator

# Sample resume text

resume_text = """
Jane Smith
Python Developer
=== GITHUB DATA ===
Contributed to django/django (open source)
Built personal portfolio site (tutorial project)
"""

evaluator = ResumeEvaluator()  # Uses default fairness constraints

evaluation = evaluator.evaluate_resume(resume_text)

print(evaluation.json(indent=2))

# Output contains only technical scores, no demographic data

```

To inspect the fairness constraints that will be applied before running an evaluation:

```python
e = ResumeEvaluator()
prompt = e._load_evaluation_prompt(resume_text)
print(prompt)  # Displays the "CRITICAL FAIRNESS REQUIREMENTS" block

```

When extending the system with custom LLM providers, you must preserve the fairness constraints by maintaining the same prompt structure:

```python
from llm_utils import initialize_llm_provider

class CustomEvaluator(ResumeEvaluator):
    def _initialize_llm_provider(self):
        # Custom provider implementation

        self.provider = initialize_llm_provider(self.model_name, custom=True)
        # Note: The fairness constraints in self.system_message remain unchanged

```

## Summary

- **Prompt templates** in `prompts/templates/` explicitly prohibit protected attributes (name, gender, school, GPA) from influencing scores.
- **Structured output** via `EvaluationData.model_json_schema()` forces JSON-only responses, eliminating biased narrative generation.
- **Runtime validation** in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) (lines 80-87) parses responses through Pydantic models, rejecting any malformed or non-compliant outputs.
- **Audit logging** captures raw LLM responses for transparency and compliance verification.
- **Declarative scoring rules** encode technical merit criteria directly in templates, ensuring consistent evaluation standards across all candidates.

## Frequently Asked Questions

### What specific protected attributes does hiring-agent exclude from evaluations?

According to the source code in `prompts/templates/resume_evaluation_system_message.jinja`, the system explicitly excludes name, gender, school name, GPA, and geographic location from scoring considerations. The templates instruct the LLM to evaluate candidates based solely on open-source contributions, production experience, self-projects, and technical skills.

### How does the system prevent the LLM from generating biased narrative responses?

The `ResumeEvaluator` class prevents biased narratives by passing `EvaluationData.model_json_schema()` as a format constraint to the LLM provider (lines 76-78 in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)). This forces the model to output strictly structured JSON that matches the Pydantic schema defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), eliminating free-form text where subjective bias might appear.

### Can the fairness constraints be customized or overridden?

While you can extend the `ResumeEvaluator` class to customize LLM providers or add evaluation criteria, the fairness constraints are embedded in the Jinja2 templates loaded during initialization. To modify fairness rules, you would need to edit `prompts/templates/resume_evaluation_system_message.jinja` directly, ensuring any changes remain transparent and auditable through version control.

### How does hiring-agent maintain an audit trail for fairness compliance?

The system logs all raw LLM responses before Pydantic validation (line 81 in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)). This captures the complete text returned by the model, allowing compliance teams to verify that the prompt constraints were properly enforced and that no protected attributes influenced the final structured output.