# How Hiring Agent Generates Evaluation Evidence: A Deep Dive into the LLM Pipeline

> Discover how Hiring Agent generates evaluation evidence. Learn how LLM pipelines analyze résumés and produce scored JSON with textual justifications for hiring decisions.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-06-29

---

**Hiring Agent generates evaluation evidence by prompting a large language model to analyze résumé text and return structured JSON containing scored categories with textual justifications.**

The open-source Hiring Agent repository from InterviewStreet automates résumé screening by extracting evaluation evidence directly from LLM outputs. This evidence appears as detailed textual justifications accompanying each numeric score in the final assessment. Understanding how this evidence is generated requires tracing the flow from Jinja template prompts through JSON schema enforcement to final presentation.

## The Evaluation Pipeline Architecture

Hiring Agent’s evidence generation relies on a structured pipeline that transforms raw résumé text into scored categories with explanations.

### Prompt Construction with Jinja Templates

The evaluation process begins in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) where the `ResumeEvaluator` class constructs a multi-part prompt. The system loads two critical templates from the `prompts/templates/` directory:

- `resume_evaluation_criteria.jinja` – Injects the raw résumé text into the evaluation criteria
- `resume_evaluation_system_message.jinja` – Instructs the LLM to format responses as structured JSON

According to the source code in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) (lines 40-46), these templates combine to create a prompt that explicitly requests the model to provide evidence for each scoring category.

### LLM Provider Integration

Once constructed, the evaluator passes these messages to the configured provider. The code at lines 62-78 in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) invokes `self.provider.chat()` with three critical parameters:

1. The system and user messages
2. A `format` argument enforcing the `EvaluationData` JSON schema
3. The selected model identifier

This call forces the LLM to output parseable JSON rather than free-form text, ensuring the evaluation evidence can be programmatically extracted.

## Extracting Evidence from the LLM Response

The raw LLM response contains the evaluation evidence embedded within a JSON structure. The system extracts and validates this data before presenting it to users.

### JSON Schema Enforcement

After receiving the LLM response, [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) (lines 80-86) processes the raw text through `extract_json_from_response()` to clean markdown formatting or other extraneous characters. The cleaned JSON is then deserialized into an `EvaluationData` instance defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py).

This schema validation ensures that required fields—including the `evidence` strings for each category—are present and correctly typed before the evaluation proceeds.

### Evidence Field Extraction

The evaluation evidence itself resides within the `scores` object of the `EvaluationData` structure. As implemented in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 87-92), each category (such as *open_source*, *self_projects*, *production*, or *technical_skills*) contains:

- A numeric `score`
- A `max` value indicating the category weight
- An `evidence` string containing the LLM’s written justification

This evidence represents the model’s reasoning for assigning specific scores based on the résumé content and evaluation criteria.

## Code Implementation Examples

The following examples demonstrate how to interact with the evaluation evidence generation pipeline programmatically.

### Running a Command-Line Evaluation

To generate evaluation evidence from a PDF résumé, execute:

```bash
python score.py path/to/resume.pdf

```

This script orchestrates the full pipeline: extracting PDF text, optionally enriching it with GitHub profile data, invoking `ResumeEvaluator.evaluate_resume()`, and finally displaying the evidence through `print_evaluation_results()`.

### Accessing Evidence Programmatically

For custom integrations, instantiate the evaluator directly:

```python
from evaluator import ResumeEvaluator

# Assume resume_text contains extracted PDF content

resume_text = "Jane Doe\nSenior Software Engineer…"

evaluator = ResumeEvaluator()
evaluation = evaluator.evaluate_resume(resume_text)

# Retrieve evidence for specific categories

print(evaluation.scores.open_source.evidence)
print(evaluation.scores.technical_skills.evidence)

```

### Inspecting Raw LLM Responses

To debug or audit the evidence generation, inspect the raw JSON before parsing:

```python
from evaluator import ResumeEvaluator
from llm_utils import extract_json_from_response

evaluator = ResumeEvaluator()

# Manually invoke the provider

raw_response = evaluator.provider.chat(
    model="gpt-4o-mini",
    messages=[
        {"role": "system", "content": "System prompt here"},
        {"role": "user", "content": "Resume text here"}
    ],
    format=EvaluationData.model_json_schema()
)

# Extract and display the raw JSON containing evidence

json_text = extract_json_from_response(raw_response["message"]["content"])
print(json_text)

```

## Rendering Evidence to End Users

The final presentation of evaluation evidence occurs in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) through the `print_evaluation_results()` function (lines 87-104). This utility iterates through the `EvaluationData.scores` object and prints each category’s evidence alongside its numeric score.

This rendering step transforms the structured JSON data into human-readable justifications, allowing hiring managers to understand not just *what* score a candidate received, but *why* the LLM assigned that score based on specific résumé content.

## Summary

- **Hiring Agent** generates evaluation evidence by prompting an LLM to return structured JSON containing justifications for each scoring category.
- **Prompt templates** in `prompts/templates/` (specifically `resume_evaluation_criteria.jinja` and `resume_evaluation_system_message.jinja`) guide the model to produce evidence-based responses.
- **Schema enforcement** via `EvaluationData` in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) ensures the LLM output includes required evidence fields for categories like *open_source*, *self_projects*, and *technical_skills*.
- **Extraction logic** in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) (lines 80-86) parses the LLM response and deserializes it into typed objects containing the evidence strings.
- **Presentation** in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 87-104) renders the evidence to users alongside numeric scores, providing transparent hiring decisions.

## Frequently Asked Questions

### How does Hiring Agent ensure the LLM generates consistent evidence?

Hiring Agent enforces consistency by passing a strict JSON schema (derived from `EvaluationData` in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)) to the LLM via the `format` parameter. This schema requires specific fields including `evidence` strings for each category, and the `extract_json_from_response()` function in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) validates the structure before processing.

### Can I customize the evaluation criteria that generate the evidence?

Yes. The evaluation criteria reside in Jinja templates within `prompts/templates/resume_evaluation_criteria.jinja`. Modifying this template changes the instructions given to the LLM, allowing you to adjust what aspects of a résumé trigger specific evidence generation or alter the weighting of different categories.

### What file contains the logic for displaying the evaluation evidence?

The presentation logic lives in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py), specifically within the `print_evaluation_results()` function (lines 87-104). This function iterates through the scored categories and prints the evidence strings alongside numeric scores, formatting the LLM-generated justifications for end-user review.

### Is the evaluation evidence stored or just printed to stdout?

According to the current implementation in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py), the evidence is primarily handled through the `print_evaluation_results()` function, which outputs to stdout. For persistent storage, you would need to modify the script to capture the `EvaluationData` object returned by `evaluate_resume()` and serialize it to a database or file format of your choice.