How Hiring Agent Generates Evaluation Evidence: A Deep Dive into the LLM Pipeline
Hiring Agent generates evaluation evidence by prompting a large language model to analyze résumé text and return structured JSON containing scored categories with textual justifications.
The open-source Hiring Agent repository from InterviewStreet automates résumé screening by extracting evaluation evidence directly from LLM outputs. This evidence appears as detailed textual justifications accompanying each numeric score in the final assessment. Understanding how this evidence is generated requires tracing the flow from Jinja template prompts through JSON schema enforcement to final presentation.
The Evaluation Pipeline Architecture
Hiring Agent’s evidence generation relies on a structured pipeline that transforms raw résumé text into scored categories with explanations.
Prompt Construction with Jinja Templates
The evaluation process begins in evaluator.py where the ResumeEvaluator class constructs a multi-part prompt. The system loads two critical templates from the prompts/templates/ directory:
resume_evaluation_criteria.jinja– Injects the raw résumé text into the evaluation criteriaresume_evaluation_system_message.jinja– Instructs the LLM to format responses as structured JSON
According to the source code in evaluator.py (lines 40-46), these templates combine to create a prompt that explicitly requests the model to provide evidence for each scoring category.
LLM Provider Integration
Once constructed, the evaluator passes these messages to the configured provider. The code at lines 62-78 in evaluator.py invokes self.provider.chat() with three critical parameters:
- The system and user messages
- A
formatargument enforcing theEvaluationDataJSON schema - The selected model identifier
This call forces the LLM to output parseable JSON rather than free-form text, ensuring the evaluation evidence can be programmatically extracted.
Extracting Evidence from the LLM Response
The raw LLM response contains the evaluation evidence embedded within a JSON structure. The system extracts and validates this data before presenting it to users.
JSON Schema Enforcement
After receiving the LLM response, evaluator.py (lines 80-86) processes the raw text through extract_json_from_response() to clean markdown formatting or other extraneous characters. The cleaned JSON is then deserialized into an EvaluationData instance defined in models.py.
This schema validation ensures that required fields—including the evidence strings for each category—are present and correctly typed before the evaluation proceeds.
Evidence Field Extraction
The evaluation evidence itself resides within the scores object of the EvaluationData structure. As implemented in score.py (lines 87-92), each category (such as open_source, self_projects, production, or technical_skills) contains:
- A numeric
score - A
maxvalue indicating the category weight - An
evidencestring containing the LLM’s written justification
This evidence represents the model’s reasoning for assigning specific scores based on the résumé content and evaluation criteria.
Code Implementation Examples
The following examples demonstrate how to interact with the evaluation evidence generation pipeline programmatically.
Running a Command-Line Evaluation
To generate evaluation evidence from a PDF résumé, execute:
python score.py path/to/resume.pdf
This script orchestrates the full pipeline: extracting PDF text, optionally enriching it with GitHub profile data, invoking ResumeEvaluator.evaluate_resume(), and finally displaying the evidence through print_evaluation_results().
Accessing Evidence Programmatically
For custom integrations, instantiate the evaluator directly:
from evaluator import ResumeEvaluator
# Assume resume_text contains extracted PDF content
resume_text = "Jane Doe\nSenior Software Engineer…"
evaluator = ResumeEvaluator()
evaluation = evaluator.evaluate_resume(resume_text)
# Retrieve evidence for specific categories
print(evaluation.scores.open_source.evidence)
print(evaluation.scores.technical_skills.evidence)
Inspecting Raw LLM Responses
To debug or audit the evidence generation, inspect the raw JSON before parsing:
from evaluator import ResumeEvaluator
from llm_utils import extract_json_from_response
evaluator = ResumeEvaluator()
# Manually invoke the provider
raw_response = evaluator.provider.chat(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "System prompt here"},
{"role": "user", "content": "Resume text here"}
],
format=EvaluationData.model_json_schema()
)
# Extract and display the raw JSON containing evidence
json_text = extract_json_from_response(raw_response["message"]["content"])
print(json_text)
Rendering Evidence to End Users
The final presentation of evaluation evidence occurs in score.py through the print_evaluation_results() function (lines 87-104). This utility iterates through the EvaluationData.scores object and prints each category’s evidence alongside its numeric score.
This rendering step transforms the structured JSON data into human-readable justifications, allowing hiring managers to understand not just what score a candidate received, but why the LLM assigned that score based on specific résumé content.
Summary
- Hiring Agent generates evaluation evidence by prompting an LLM to return structured JSON containing justifications for each scoring category.
- Prompt templates in
prompts/templates/(specificallyresume_evaluation_criteria.jinjaandresume_evaluation_system_message.jinja) guide the model to produce evidence-based responses. - Schema enforcement via
EvaluationDatainmodels.pyensures the LLM output includes required evidence fields for categories like open_source, self_projects, and technical_skills. - Extraction logic in
evaluator.py(lines 80-86) parses the LLM response and deserializes it into typed objects containing the evidence strings. - Presentation in
score.py(lines 87-104) renders the evidence to users alongside numeric scores, providing transparent hiring decisions.
Frequently Asked Questions
How does Hiring Agent ensure the LLM generates consistent evidence?
Hiring Agent enforces consistency by passing a strict JSON schema (derived from EvaluationData in models.py) to the LLM via the format parameter. This schema requires specific fields including evidence strings for each category, and the extract_json_from_response() function in llm_utils.py validates the structure before processing.
Can I customize the evaluation criteria that generate the evidence?
Yes. The evaluation criteria reside in Jinja templates within prompts/templates/resume_evaluation_criteria.jinja. Modifying this template changes the instructions given to the LLM, allowing you to adjust what aspects of a résumé trigger specific evidence generation or alter the weighting of different categories.
What file contains the logic for displaying the evaluation evidence?
The presentation logic lives in score.py, specifically within the print_evaluation_results() function (lines 87-104). This function iterates through the scored categories and prints the evidence strings alongside numeric scores, formatting the LLM-generated justifications for end-user review.
Is the evaluation evidence stored or just printed to stdout?
According to the current implementation in score.py, the evidence is primarily handled through the print_evaluation_results() function, which outputs to stdout. For persistent storage, you would need to modify the script to capture the EvaluationData object returned by evaluate_resume() and serialize it to a database or file format of your choice.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →