Hiring Agent Output Evaluation Report Structure: Complete Schema Guide

The Hiring Agent returns a structured JSON evaluation report that conforms to the EvaluationData Pydantic model, containing quantitative scores across four categories, bonus points, deductions, key strengths, and improvement areas.

The InterviewStreet Hiring Agent transforms unstructured resume text into a machine-readable assessment using a strict schema-driven approach. Understanding the output evaluation report structure is essential for integrating this tool into hiring pipelines and building dashboards that consume candidate scores.

Core Schema Components

The evaluation report is defined by the EvaluationData Pydantic model located in models.py. This schema enforces type safety and guarantees that the JSON emitted by the LLM matches the expected structure.

The EvaluationData Root Model

According to the InterviewStreet Hiring Agent source code, the root model EvaluationData (defined in models.py lines 44‑50) serves as the container for the entire assessment. This model aggregates scores, bonuses, deductions, and qualitative feedback into a single serializable object that can be instantiated directly from LLM output.

Quantitative Scores Structure

The scores field contains four core categories, each implemented as a CategoryScore object (defined in models.py lines 18‑22). The Scores class (lines 24‑30) groups these into:

  • open_source – Contributions to public repositories
  • self_projects – Personal development work
  • production – Enterprise or production-grade experience
  • technical_skills – Language and framework proficiency

Each CategoryScore contains three fields:

  • score (float ≥ 0) – The calculated rating
  • max (int > 0) – The maximum possible points for the category
  • evidence (string) – Supporting rationale from the resume

Bonus Points and Deductions

The report includes adjustment mechanisms defined in models.py:

  • BonusPoints (lines 31‑34): Contains total (float 0‑20) and breakdown (string description) for exceptional achievements
  • Deductions (lines 36‑42): Contains total (float ≥ 0) and reasons (string description) for penalties applied to missing or weak aspects

Qualitative Feedback Sections

Beyond numerical scoring, the report provides narrative insights:

  • key_strengths – List of up to five concise bullet points highlighting standout attributes
  • areas_for_improvement – List of up to five bullet points indicating development opportunities

How the Report is Generated

The evaluation report structure is enforced during the LLM interaction pipeline implemented in evaluator.py. When ResumeEvaluator.evaluate_resume calls the underlying language model, it requests the response conform to EvaluationData.model_json_schema() to ensure valid JSON output.

The raw chat response is extracted using utilities from llm_utils.py, sanitized, and then instantiated as EvaluationData(**evaluation_dict) (see evaluator.py lines 78‑86). This validation step guarantees that any deviation from the schema raises a ValidationError before the data propagates to downstream systems.

Working with Evaluation Reports

Generating a JSON Report

To generate an evaluation report from resume text and serialize it to JSON:

from evaluator import ResumeEvaluator

# Initialise the evaluator (default model = Gemini)

evaluator = ResumeEvaluator()

# Example raw resume text (plain string)

resume_text = """John Doe
Software Engineer with 5 years of experience...
"""

# Perform evaluation – returns a typed `EvaluationData` instance

report = evaluator.evaluate_resume(resume_text)

# Serialize to JSON for API response or file storage

json_report = report.model_dump_json(indent=2)
print(json_report)

Sample output (condensed):

{
  "scores": {
    "open_source": {"score": 8.5, "max": 10, "evidence": "Contributed to 3 high‑profile repos"},
    "self_projects": {"score": 7.0, "max": 10, "evidence": "Built a personal portfolio site"},
    "production": {"score": 9.0, "max": 10, "evidence": "Led production‑grade microservices"},
    "technical_skills": {"score": 8.0, "max": 10, "evidence": "Proficient in Python, Go, Docker"}
  },
  "bonus_points": {"total": 5.0, "breakdown": "Open‑source maintainer award"},
  "deductions": {"total": 2.0, "reasons": "Missing CI/CD pipeline description"},
  "key_strengths": [
    "Strong open‑source contributions",
    "Leadership in production environments"
  ],
  "areas_for_improvement": [
    "Add CI/CD workflow details",
    "Provide clearer project impact metrics"
  ]
}

Accessing Individual Sections

Once instantiated, the EvaluationData object provides typed access to all fields:


# Retrieve the overall technical‑skills score

tech_score = report.scores.technical_skills.score
print(f"Technical Skills Score: {tech_score}/{report.scores.technical_skills.max}")

# List the key strengths

for i, strength in enumerate(report.key_strengths, 1):
    print(f"{i}. {strength}")

Validating External JSON Payloads

To validate external JSON against the Hiring Agent schema before processing:

from models import EvaluationData
import json

payload = json.loads(some_external_json_string)

# This will raise a ValidationError if the payload does not match the schema

validated_report = EvaluationData(**payload)

print("Payload is valid and ready for further processing.")

Summary

  • The Hiring Agent output evaluation report structure is defined by the EvaluationData Pydantic model in models.py, ensuring type-safe JSON generation.
  • The report contains four quantitative scoring categories (open_source, self_projects, production, technical_skills), each with a score, maximum value, and evidence string.
  • Adjustment fields bonus_points and deductions capture exceptional achievements and penalties with descriptive reasoning.
  • Qualitative arrays key_strengths and areas_for_improvement provide up to five bullet points each for narrative feedback.
  • The ResumeEvaluator.evaluate_resume method in evaluator.py orchestrates LLM calls and validates outputs against the schema at lines 78‑86.

Frequently Asked Questions

What Pydantic model defines the Hiring Agent evaluation report structure?

The EvaluationData model defined in models.py (lines 44‑50) serves as the root schema. It aggregates Scores, BonusPoints, Deductions, and string arrays for strengths and improvement areas, enforcing type safety through Pydantic validation.

How are the four scoring categories structured in the output?

Each category is a CategoryScore object (defined in models.py lines 18‑22) containing three fields: score (float), max (integer), and evidence (string). These are grouped under the scores field as open_source, self_projects, production, and technical_skills.

Can I validate external JSON against the Hiring Agent schema?

Yes. Import EvaluationData from models.py and instantiate it with the external dictionary using EvaluationData(**payload). This will raise a ValidationError if the JSON does not conform to the expected schema, including missing required fields or type mismatches.

Where is the evaluation report generated in the source code?

The report generation occurs in evaluator.py within the ResumeEvaluator.evaluate_resume method. Specifically, lines 78‑86 handle the extraction of raw JSON from the LLM response and its instantiation as an EvaluationData object, ensuring the output matches the defined structure before returning to the caller.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →