What Is evaluator.py in the interviewstreet/hiring-agent Repository?

evaluator.py is the core orchestration module that evaluates candidate résumés by invoking a Large Language Model and converting the unstructured response into a structured EvaluationData object.

The hiring-agent repository by InterviewStreet automates technical candidate screening through LLM-powered analysis. According to the interviewstreet/hiring-agent source code, evaluator.py transforms free-form résumé text into quantified, machine-readable evaluation data that downstream hiring logic can consume.

Core Responsibilities of the ResumeEvaluator Class

The ResumeEvaluator class defined in evaluator.py serves as the primary interface for résumé scoring. It encapsulates provider management, prompt templating, and response validation into a single workflow.

Initialization and Model Configuration

When instantiated, ResumeEvaluator accepts a model_name parameter (defaulting to prompt.DEFAULT_MODEL) and resolves model-specific parameters from prompt.MODEL_PARAMETERS. It also initializes a TemplateManager instance to handle Jinja-style prompt rendering.

LLM Provider Selection via _initialize_llm_provider

The private method _initialize_llm_provider delegates to initialize_llm_provider from llm_utils.py. This factory function selects the appropriate provider class—either OllamaProvider or GeminiProvider—based on the model name. Both providers implement a uniform chat method, ensuring the remainder of the evaluation logic remains provider-agnostic.

Prompt Preparation with _load_evaluation_prompt

The evaluator renders two distinct templates using the TemplateManager:

  • A system message template providing high-level instructions to the model
  • A resume evaluation criteria template populated with the raw résumé text, instructing the LLM to score specific categories

Structured LLM Request Construction

The evaluate_resume method constructs a chat payload containing:

  • model: The selected model identifier
  • messages: An array containing the system message and user criteria prompt
  • options: Generation parameters including temperature and top-p settings
  • format: A JSON schema generated from EvaluationData.model_json_schema(), forcing the LLM to output valid JSON matching the expected structure

Response Handling and Data Validation

After the provider returns a response, evaluator.py calls extract_json_from_response (from llm_utils.py) to strip extraneous text. The resulting JSON string is parsed into an EvaluationData instance defined in models.py. This object contains structured scores for open-source contributions, self-projects, production experience, and technical skills, plus bonus points, deductions, key strengths, and improvement areas.

Scoring Boundaries and Constants

Lines 9-11 of evaluator.py define the constants MAX_BONUS_POINTS, MIN_FINAL_SCORE, and MAX_FINAL_SCORE. These boundaries normalize the final hiring score for downstream consumption.

Integration with Auxiliary Modules

The evaluator relies on several co-located modules to function:

  • models.py: Defines the EvaluationData Pydantic schema that validates LLM output
  • prompt.py: Stores DEFAULT_MODEL, MODEL_PARAMETERS, and provider mappings
  • prompts/template_manager.py: Loads and renders Jinja templates for system messages and evaluation criteria
  • llm_utils.py: Provides the provider factory and JSON extraction utilities

Practical Implementation Examples

The following example demonstrates evaluating a résumé using the Gemini provider:

from evaluator import ResumeEvaluator

# Example résumé text (could be read from a file or a JSONResume instance)

resume_text = """
John Doe
Software Engineer with 5 years of experience in Python and cloud services.
...
"""

# Create the evaluator – you can specify a model or rely on defaults

evaluator = ResumeEvaluator(model_name="gemini-pro")   # uses GeminiProvider under the hood

# Run the evaluation

evaluation = evaluator.evaluate_resume(resume_text)

# The returned object is a Pydantic model; you can access fields directly

print("Open‑source score:", evaluation.scores.open_source.score)
print("Bonus points:", evaluation.bonus_points.total)
print("Key strengths:", evaluation.key_strengths)

To switch to a local Ollama instance, change the model name to one mapped in prompt.MODEL_PROVIDER_MAPPING:

evaluator = ResumeEvaluator(model_name="llama3")
evaluation = evaluator.evaluate_resume(resume_text)

Summary

  • evaluator.py orchestrates the entire résumé evaluation pipeline by managing LLM providers, rendering prompts, and validating outputs.
  • The ResumeEvaluator class supports multiple providers (Gemini and Ollama) through a unified interface implemented in llm_utils.py.
  • Structured output is enforced via JSON schema validation against the EvaluationData model from models.py.
  • Scoring boundaries (MAX_BONUS_POINTS, MIN_FINAL_SCORE, MAX_FINAL_SCORE) ensure consistent normalization of candidate scores.
  • The module transforms unstructured résumé text into machine-readable evaluation data containing category-specific scores, strengths, and improvement areas.

Frequently Asked Questions

What LLM providers does evaluator.py support?

According to the source code in llm_utils.py, evaluator.py supports Google Gemini via GeminiProvider and local Ollama models via OllamaProvider. The specific provider is selected at runtime based on the model_name parameter and the mappings defined in prompt.MODEL_PROVIDER_MAPPING.

How does evaluator.py ensure the LLM returns valid JSON?

The evaluate_resume method passes a format parameter containing a JSON schema generated by EvaluationData.model_json_schema(). This schema forces the LLM to output a JSON object that matches the structure defined in models.py, which is then validated and parsed into a Pydantic instance.

What information does the EvaluationData object contain?

The EvaluationData object includes structured scores for open-source contributions, self-projects, production experience, and technical skills. It also tracks bonus points, deductions, key strengths, and specific areas for improvement, providing a comprehensive quantified assessment of the candidate.

How does evaluator.py handle errors during LLM calls?

Any exceptions raised during the LLM invocation are logged and re-raised within the evaluate_resume method. This allows calling code to implement appropriate error handling while ensuring that transient API failures or invalid responses do not propagate silently through the hiring pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →