Pydantic Schemas Used in the Hiring Agent: Complete Data Model Reference
The Hiring Agent defines all data contracts in a single models.py module containing 20+ Pydantic schemas that validate JSON Resume input, structure LLM evaluation output, and enforce type safety across the AI-powered screening pipeline.
The interviewstreet/hiring-agent repository relies on strict Pydantic models to maintain data integrity between candidate resumes and LLM-generated evaluations. These schemas live entirely in models.py and cover everything from personal contact details to complex scoring rubrics. This guide breaks down every Pydantic schema used in the Hiring Agent, their relationships, and how they validate data at runtime.
Resume Input Schemas (JSON Resume Format)
The input layer implements the JSON Resume standard through nested Pydantic models. All classes inherit from pydantic.BaseModel and provide automatic validation for candidate data.
Core Personal Data
- Basics: Defined at line 46 in
models.py, this schema aggregates name, email, phone, summary, and nested objects. - Location: Line 28 defines the address model capturing city, country, and region fields.
- Profile: Line 38 represents social media entries with network name, username, and URL fields.
Professional History and Education
- Work: Line 58 defines the employment history model with company, position, start/end dates, and highlights.
- Volunteer: Line 70 mirrors the Work structure for volunteer experience entries.
- Education: Line 82 handles academic records including institution, degree, dates, and relevant courses.
Skills, Projects, and Achievements
- Skill: Line 123 captures technical competencies with name, proficiency level, and keyword lists.
- Project: Line 152 stores portfolio entries with dates, descriptions, tech stack, and related skills.
- Award: Line 95 for honors and distinctions.
- Certificate: Line 104 for professional certifications with issuer and date.
- Publication: Line 113 for published works including title, publisher, and release date.
- Language: Line 131 for spoken language proficiency ratings.
- Interest: Line 138 for personal interest categories.
- Reference: Line 145 for professional references including name and recommendation text.
Section Wrappers and Top-Level Assembly
To organize lists of the above models, lines 165-195 define section wrappers: BasicsSection, WorkSection, EducationSection, SkillsSection, ProjectsSection, and AwardsSection. These containers group related entries for cleaner JSON structure.
The JSONResume schema at line 201 serves as the root object, aggregating all optional sections into a single validated resume instance.
Evaluation Result Schemas
After the LLM processes a candidate, it returns structured scoring data captured by these output schemas.
Scoring Structure
- CategoryScore: Line 218 defines individual scoring buckets containing
score,max, andevidencefields. - Scores: Line 224 acts as a container for four core evaluation dimensions:
open_source,self_projects,production, andtechnical_skills.
Adjustments and Final Output
- BonusPoints: Line 231 tracks optional bonus points (capped at 20 total) with a breakdown explanation.
- Deductions: Line 236 records score reductions with detailed reasoning.
- EvaluationData: Line 244 is the top-level result schema combining scores, bonuses, deductions, plus narrative arrays for
key_strengthsandareas_for_improvement.
Auxiliary Data Models
GitHub Profile Enrichment
The GitHubProfile schema at line 252 provides a lightweight representation of GitHub user data. The GitHubProvider uses this model when enriching candidate profiles with external repository activity.
Validation Flow and Implementation
The Pydantic schemas enforce strict contracts through a four-stage validation pipeline:
- Input Validation: Raw CVs are parsed into
JSONResumeinstances, ensuring all personal data, work history, and skills conform to expected Python types. - LLM Integration: The
ResumeEvaluator.evaluate_resumemethod inevaluator.pysends resume text to the configured LLM provider with a format hint matching theEvaluationDataJSON schema. - Response Extraction: The
extract_json_from_responsefunction inllm_utils.pyisolates the JSON payload from the LLM's raw text output. - Output Validation: Pydantic validates the extracted JSON against
EvaluationData, raisingpydantic.ValidationErrorfor any malformed data before it reaches the final scoring logic.
Working with the Schemas
Constructing a Resume Instance
Create a valid JSONResume by nesting the component models:
from models import JSONResume, Basics, Location, Profile
resume = JSONResume(
basics=Basics(
name="Ada Lovelace",
email="ada@example.com",
location=Location(city="London", countryCode="GB"),
profiles=[Profile(network="GitHub", username="ada", url="https://github.com/ada")]
)
)
Simulating Evaluation Results
For testing or mocking LLM outputs, instantiate EvaluationData directly:
from models import EvaluationData, Scores, CategoryScore, BonusPoints, Deductions
evaluation = EvaluationData(
scores=Scores(
open_source=CategoryScore(score=9, max=10, evidence="Contributed to 5 OSS projects"),
self_projects=CategoryScore(score=8, max=10, evidence="Built a personal portfolio site"),
production=CategoryScore(score=7, max=10, evidence="2 years at Acme Corp"),
technical_skills=CategoryScore(score=8, max=10, evidence="Proficient in Python, Go, Rust")
),
bonus_points=BonusPoints(total=12, breakdown="Open‑source leadership + mentorship"),
deductions=Deductions(total=0, reasons=""),
key_strengths=["Technical depth", "Communication"],
areas_for_improvement=["Time management"]
)
Summary
- All Pydantic schemas are centralized in
models.pyand inherit frompydantic.BaseModel. - JSONResume and its 14+ constituent models handle standardized candidate input data according to the JSON Resume spec.
- EvaluationData structures the LLM output with strict scoring categories, bonuses, deductions, and narrative feedback.
- GitHubProfile supports external data enrichment for candidate verification.
- Runtime validation occurs in
evaluator.pyandllm_utils.py, ensuring only type-valid data flows through the screening pipeline.
Frequently Asked Questions
Where are the Pydantic schemas defined in the Hiring Agent?
All schemas are centralized in models.py at the repository root. This file contains every data model from Location to EvaluationData, keeping the data layer unified and maintainable.
What happens if the LLM returns malformed JSON?
The extract_json_from_response function in llm_utils.py isolates the JSON payload, and Pydantic validates it against the EvaluationData schema. If validation fails, Pydantic raises a ValidationError, preventing corrupted data from propagating through the evaluation pipeline.
How does the JSONResume schema relate to the evaluation output?
JSONResume serves as the input contract for candidate data, while EvaluationData serves as the output contract. The ResumeEvaluator class transforms the resume into text for the LLM, then validates the LLM's response against EvaluationData, creating a type-safe boundary between input and output.
Can I use these schemas for my own resume evaluation tool?
Yes, since the repository is open-source under interviewstreet/hiring-agent, you can import the models.py schemas into your own Python projects. All models inherit from standard pydantic.BaseModel, making them compatible with FastAPI, data validation, or JSON serialization in any Pydantic-supported workflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →