# What Is evaluator.py in the interviewstreet/hiring-agent Repository?

> Discover how evaluator.py in interviewstreet/hiring-agent orchestrates résumé evaluation using LLMs to convert unstructured data into structured EvaluationData objects.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: internals
- Published: 2026-06-28

---

**[`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) is the core orchestration module that evaluates candidate résumés by invoking a Large Language Model and converting the unstructured response into a structured `EvaluationData` object.**

The `hiring-agent` repository by InterviewStreet automates technical candidate screening through LLM-powered analysis. According to the interviewstreet/hiring-agent source code, [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) transforms free-form résumé text into quantified, machine-readable evaluation data that downstream hiring logic can consume.

## Core Responsibilities of the ResumeEvaluator Class

The `ResumeEvaluator` class defined in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) serves as the primary interface for résumé scoring. It encapsulates provider management, prompt templating, and response validation into a single workflow.

### Initialization and Model Configuration

When instantiated, `ResumeEvaluator` accepts a `model_name` parameter (defaulting to `prompt.DEFAULT_MODEL`) and resolves model-specific parameters from `prompt.MODEL_PARAMETERS`. It also initializes a `TemplateManager` instance to handle Jinja-style prompt rendering.

### LLM Provider Selection via `_initialize_llm_provider`

The private method `_initialize_llm_provider` delegates to `initialize_llm_provider` from [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py). This factory function selects the appropriate provider class—either `OllamaProvider` or `GeminiProvider`—based on the model name. Both providers implement a uniform `chat` method, ensuring the remainder of the evaluation logic remains provider-agnostic.

### Prompt Preparation with `_load_evaluation_prompt`

The evaluator renders two distinct templates using the `TemplateManager`:
- A **system message** template providing high-level instructions to the model
- A **resume evaluation criteria** template populated with the raw résumé text, instructing the LLM to score specific categories

### Structured LLM Request Construction

The `evaluate_resume` method constructs a chat payload containing:
- `model`: The selected model identifier
- `messages`: An array containing the system message and user criteria prompt
- `options`: Generation parameters including temperature and top-p settings
- `format`: A JSON schema generated from `EvaluationData.model_json_schema()`, forcing the LLM to output valid JSON matching the expected structure

### Response Handling and Data Validation

After the provider returns a response, [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) calls `extract_json_from_response` (from [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)) to strip extraneous text. The resulting JSON string is parsed into an `EvaluationData` instance defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py). This object contains structured scores for open-source contributions, self-projects, production experience, and technical skills, plus bonus points, deductions, key strengths, and improvement areas.

### Scoring Boundaries and Constants

Lines 9-11 of [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) define the constants `MAX_BONUS_POINTS`, `MIN_FINAL_SCORE`, and `MAX_FINAL_SCORE`. These boundaries normalize the final hiring score for downstream consumption.

## Integration with Auxiliary Modules

The evaluator relies on several co-located modules to function:

- **[`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)**: Defines the `EvaluationData` Pydantic schema that validates LLM output
- **[`prompt.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompt.py)**: Stores `DEFAULT_MODEL`, `MODEL_PARAMETERS`, and provider mappings
- **[`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py)**: Loads and renders Jinja templates for system messages and evaluation criteria
- **[`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)**: Provides the provider factory and JSON extraction utilities

## Practical Implementation Examples

The following example demonstrates evaluating a résumé using the Gemini provider:

```python
from evaluator import ResumeEvaluator

# Example résumé text (could be read from a file or a JSONResume instance)

resume_text = """
John Doe
Software Engineer with 5 years of experience in Python and cloud services.
...
"""

# Create the evaluator – you can specify a model or rely on defaults

evaluator = ResumeEvaluator(model_name="gemini-pro")   # uses GeminiProvider under the hood

# Run the evaluation

evaluation = evaluator.evaluate_resume(resume_text)

# The returned object is a Pydantic model; you can access fields directly

print("Open‑source score:", evaluation.scores.open_source.score)
print("Bonus points:", evaluation.bonus_points.total)
print("Key strengths:", evaluation.key_strengths)

```

To switch to a local Ollama instance, change the model name to one mapped in `prompt.MODEL_PROVIDER_MAPPING`:

```python
evaluator = ResumeEvaluator(model_name="llama3")
evaluation = evaluator.evaluate_resume(resume_text)

```

## Summary

- **[`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)** orchestrates the entire résumé evaluation pipeline by managing LLM providers, rendering prompts, and validating outputs.
- The **`ResumeEvaluator`** class supports multiple providers (Gemini and Ollama) through a unified interface implemented in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py).
- **Structured output** is enforced via JSON schema validation against the `EvaluationData` model from [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py).
- **Scoring boundaries** (`MAX_BONUS_POINTS`, `MIN_FINAL_SCORE`, `MAX_FINAL_SCORE`) ensure consistent normalization of candidate scores.
- The module transforms unstructured résumé text into machine-readable evaluation data containing category-specific scores, strengths, and improvement areas.

## Frequently Asked Questions

### What LLM providers does evaluator.py support?

According to the source code in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py), [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) supports Google Gemini via `GeminiProvider` and local Ollama models via `OllamaProvider`. The specific provider is selected at runtime based on the `model_name` parameter and the mappings defined in `prompt.MODEL_PROVIDER_MAPPING`.

### How does evaluator.py ensure the LLM returns valid JSON?

The `evaluate_resume` method passes a `format` parameter containing a JSON schema generated by `EvaluationData.model_json_schema()`. This schema forces the LLM to output a JSON object that matches the structure defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), which is then validated and parsed into a Pydantic instance.

### What information does the EvaluationData object contain?

The `EvaluationData` object includes structured scores for open-source contributions, self-projects, production experience, and technical skills. It also tracks bonus points, deductions, key strengths, and specific areas for improvement, providing a comprehensive quantified assessment of the candidate.

### How does evaluator.py handle errors during LLM calls?

Any exceptions raised during the LLM invocation are logged and re-raised within the `evaluate_resume` method. This allows calling code to implement appropriate error handling while ensuring that transient API failures or invalid responses do not propagate silently through the hiring pipeline.