How Hiring Agent Uses the JSON Resume Format: Schema Definition and Data Transformation
Hiring Agent consumes loosely structured LLM outputs and normalizes them into the strict JSON Resume schema using Pydantic models defined in models.py and transformation utilities in transform.py.
The interviewstreet/hiring-agent repository implements a machine-readable approach to candidate evaluation by adopting the open-source JSON Resume format. This structured schema allows the system to parse unstructured PDF resumes into standardized data models that power automated scoring and evaluation pipelines.
Understanding the JSON Resume Schema Structure
Core Sections Defined in models.py
The JSON Resume format is implemented through Pydantic models in models.py (source). The schema organizes candidate data into these key sections:
- Basics: Core personal information including
name,email,phone,url,summary,location, andprofiles - Work: Employment history with fields like
name(company),position,url,startDate,endDate,summary, andhighlights - Volunteer: Voluntary activities containing
organization,position,url,startDate,endDate,summary, andhighlights - Education: Academic background with
institution,url,area,studyType,startDate,endDate,score, andcourses - Awards: Honors and recognitions including
title,date,awarder, andsummary - Certificates: Certifications earned with
name,date,issuer, andurl - Publications: Published works containing
name,publisher,releaseDate,url, andsummary - Skills: Skill categories with
name,level, andkeywords - Languages: Language proficiency with
languageandfluency - Interests: Personal interests with
nameandkeywords - References: Professional references with
nameandreference - Projects: Projects (open-source or personal) with
name,startDate,endDate,description,highlights,url,technologies, andskills
These models combine into the top-level JSONResume class:
class JSONResume(BaseModel):
basics: Optional[Basics] = None
work: Optional[List[Work]] = None
volunteer: Optional[List[Volunteer]] = None
education: Optional[List[Education]] = None
awards: Optional[List[Award]] = None
certificates: Optional[List[Certificate]] = None
publications: Optional[List[Publication]] = None
skills: Optional[List[Skill]] = None
languages: Optional[List[Language]] = None
interests: Optional[List[Interest]] = None
references: Optional[List[Reference]] = None
projects: Optional[List[Project]] = None
Transforming Raw LLM Data into JSON Resume Format
The Normalization Pipeline
When the LLM returns loosely-structured JSON from pdf.py processing, Hiring Agent normalizes it using functions in transform.py (source). The main entry point is transform_parsed_data(), which executes a three-step process:
- Detects which top-level sections are present in the raw payload
- Calls dedicated helpers like
transform_basics,transform_work_experience, andtransform_educationto rewrite each section into exact field names and data shapes expected by the schema - Instantiates a
JSONResumeobject for type-safe downstream access by the evaluator and CSV exporter
Field-Level Transformations
The transformation layer handles several specific normalization tasks to align raw LLM outputs with the JSON Resume format:
- Social Network Detection: Uses
extract_domain_from_urlandget_network_nameto infer platform names from profile URLs - Username Extraction: Applies
extract_username_from_urlto parse usernames from social links - Date Normalization: Employs
parse_date_rangeto standardize date formats into consistentstartDateandendDatefields - Skill Structuring: Converts free-form skill listings into structured
Skillobjects viatransform_skills_comprehensive - Project Merging: Unifies disparate project keys like "projects" and "projectsOpenSource" into a uniform list using
transform_projects_comprehensive
Practical Implementation Examples
Parsing and Normalizing Raw LLM Output
from transform import transform_parsed_data
from models import JSONResume
# `raw_payload` is the loosely-structured JSON returned by the LLM
raw_payload = {
"basics": {"name": "Alice", "email": "alice@example.com"},
"work_experience": [
{"company": "Acme", "title": "Engineer", "startDate": "Jan-Mar 2020"}
],
"skills": ["Python", "Docker", "Kubernetes"]
}
# Normalise the payload to the JSON-Resume schema
norm_dict = transform_parsed_data(raw_payload)
# Build a strongly-typed JSONResume object
resume = JSONResume(**norm_dict)
print(resume.basics.name) # → Alice
print(resume.work[0].position) # → Engineer
print(resume.skills[0].keywords) # → ['Python', 'Docker', 'Kubernetes']
Converting to Plain Text for LLM Prompts
from transform import convert_json_resume_to_text
text_blob = convert_json_resume_to_text(resume)
print(text_blob) # Human-readable sections such as "=== BASIC INFORMATION ==="
Exporting Evaluation Data to CSV
from evaluator import EvaluationData
from transform import transform_evaluation_response
# Assume `resume` (JSONResume), `github_data` (dict from github.py) and `eval_data` (EvaluationData)
csv_row = transform_evaluation_response(
file_name="alice_resume.pdf",
resume_data=resume,
github_data=github_data,
evaluation=eval_data,
)
# `csv_row` is a plain dict ready for writing to a CSV file (handled in `score.py`).
Summary
- Hiring Agent implements the open-source JSON Resume format through Pydantic models in
models.py, covering sections from Basics to Projects - The
transform_parsed_data()function intransform.pynormalizes unstructured LLM outputs into the strict schema - Specialized helpers handle URL parsing, date normalization, and skill structuration to ensure data consistency
- The
JSONResumeobject enables type-safe access for the evaluator and CSV exporter components - This transformation pipeline bridges the gap between unstructured PDF inputs and structured evaluation data
Frequently Asked Questions
What is the JSON Resume format?
The JSON Resume format is an open-source, machine-readable schema for representing curriculum vitae data. It defines standardized fields for personal information, work experience, education, skills, and projects, enabling consistent parsing across different systems and tools.
How does Hiring Agent handle inconsistent date formats in resumes?
Hiring Agent uses the parse_date_range helper function in transform.py to normalize various date string formats into standardized startDate and endDate fields. This ensures that temporal data from different resume styles conforms to the JSON Resume schema requirements.
Can Hiring Agent process non-standard section names from LLM outputs?
Yes, the transformation layer in transform.py includes comprehensive mapping functions like transform_projects_comprehensive that merge disparate keys (such as "projects" and "projectsOpenSource") into the standardized JSON Resume structure. This flexibility allows the system to handle variations in how LLMs label resume sections.
Where are the Pydantic models for the JSON Resume schema defined?
The Pydantic models are defined in models.py at the repository root. This file contains class definitions for all JSON Resume sections including Basics, Work, Education, Skill, and the top-level JSONResume container model that enforces type safety throughout the application.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →