How Hiring Agent Uses the JSON Resume Format: Schema Definition and Data Transformation

Hiring Agent consumes loosely structured LLM outputs and normalizes them into the strict JSON Resume schema using Pydantic models defined in models.py and transformation utilities in transform.py.

The interviewstreet/hiring-agent repository implements a machine-readable approach to candidate evaluation by adopting the open-source JSON Resume format. This structured schema allows the system to parse unstructured PDF resumes into standardized data models that power automated scoring and evaluation pipelines.

Understanding the JSON Resume Schema Structure

Core Sections Defined in models.py

The JSON Resume format is implemented through Pydantic models in models.py (source). The schema organizes candidate data into these key sections:

  • Basics: Core personal information including name, email, phone, url, summary, location, and profiles
  • Work: Employment history with fields like name (company), position, url, startDate, endDate, summary, and highlights
  • Volunteer: Voluntary activities containing organization, position, url, startDate, endDate, summary, and highlights
  • Education: Academic background with institution, url, area, studyType, startDate, endDate, score, and courses
  • Awards: Honors and recognitions including title, date, awarder, and summary
  • Certificates: Certifications earned with name, date, issuer, and url
  • Publications: Published works containing name, publisher, releaseDate, url, and summary
  • Skills: Skill categories with name, level, and keywords
  • Languages: Language proficiency with language and fluency
  • Interests: Personal interests with name and keywords
  • References: Professional references with name and reference
  • Projects: Projects (open-source or personal) with name, startDate, endDate, description, highlights, url, technologies, and skills

These models combine into the top-level JSONResume class:

class JSONResume(BaseModel):
    basics: Optional[Basics] = None
    work: Optional[List[Work]] = None
    volunteer: Optional[List[Volunteer]] = None
    education: Optional[List[Education]] = None
    awards: Optional[List[Award]] = None
    certificates: Optional[List[Certificate]] = None
    publications: Optional[List[Publication]] = None
    skills: Optional[List[Skill]] = None
    languages: Optional[List[Language]] = None
    interests: Optional[List[Interest]] = None
    references: Optional[List[Reference]] = None
    projects: Optional[List[Project]] = None

Transforming Raw LLM Data into JSON Resume Format

The Normalization Pipeline

When the LLM returns loosely-structured JSON from pdf.py processing, Hiring Agent normalizes it using functions in transform.py (source). The main entry point is transform_parsed_data(), which executes a three-step process:

  1. Detects which top-level sections are present in the raw payload
  2. Calls dedicated helpers like transform_basics, transform_work_experience, and transform_education to rewrite each section into exact field names and data shapes expected by the schema
  3. Instantiates a JSONResume object for type-safe downstream access by the evaluator and CSV exporter

Field-Level Transformations

The transformation layer handles several specific normalization tasks to align raw LLM outputs with the JSON Resume format:

  • Social Network Detection: Uses extract_domain_from_url and get_network_name to infer platform names from profile URLs
  • Username Extraction: Applies extract_username_from_url to parse usernames from social links
  • Date Normalization: Employs parse_date_range to standardize date formats into consistent startDate and endDate fields
  • Skill Structuring: Converts free-form skill listings into structured Skill objects via transform_skills_comprehensive
  • Project Merging: Unifies disparate project keys like "projects" and "projectsOpenSource" into a uniform list using transform_projects_comprehensive

Practical Implementation Examples

Parsing and Normalizing Raw LLM Output

from transform import transform_parsed_data
from models import JSONResume

# `raw_payload` is the loosely-structured JSON returned by the LLM

raw_payload = {
    "basics": {"name": "Alice", "email": "alice@example.com"},
    "work_experience": [
        {"company": "Acme", "title": "Engineer", "startDate": "Jan-Mar 2020"}
    ],
    "skills": ["Python", "Docker", "Kubernetes"]
}

# Normalise the payload to the JSON-Resume schema

norm_dict = transform_parsed_data(raw_payload)

# Build a strongly-typed JSONResume object

resume = JSONResume(**norm_dict)

print(resume.basics.name)          # → Alice

print(resume.work[0].position)     # → Engineer

print(resume.skills[0].keywords)   # → ['Python', 'Docker', 'Kubernetes']

Converting to Plain Text for LLM Prompts

from transform import convert_json_resume_to_text

text_blob = convert_json_resume_to_text(resume)
print(text_blob)   # Human-readable sections such as "=== BASIC INFORMATION ==="

Exporting Evaluation Data to CSV

from evaluator import EvaluationData
from transform import transform_evaluation_response

# Assume `resume` (JSONResume), `github_data` (dict from github.py) and `eval_data` (EvaluationData)

csv_row = transform_evaluation_response(
    file_name="alice_resume.pdf",
    resume_data=resume,
    github_data=github_data,
    evaluation=eval_data,
)

# `csv_row` is a plain dict ready for writing to a CSV file (handled in `score.py`).

Summary

  • Hiring Agent implements the open-source JSON Resume format through Pydantic models in models.py, covering sections from Basics to Projects
  • The transform_parsed_data() function in transform.py normalizes unstructured LLM outputs into the strict schema
  • Specialized helpers handle URL parsing, date normalization, and skill structuration to ensure data consistency
  • The JSONResume object enables type-safe access for the evaluator and CSV exporter components
  • This transformation pipeline bridges the gap between unstructured PDF inputs and structured evaluation data

Frequently Asked Questions

What is the JSON Resume format?

The JSON Resume format is an open-source, machine-readable schema for representing curriculum vitae data. It defines standardized fields for personal information, work experience, education, skills, and projects, enabling consistent parsing across different systems and tools.

How does Hiring Agent handle inconsistent date formats in resumes?

Hiring Agent uses the parse_date_range helper function in transform.py to normalize various date string formats into standardized startDate and endDate fields. This ensures that temporal data from different resume styles conforms to the JSON Resume schema requirements.

Can Hiring Agent process non-standard section names from LLM outputs?

Yes, the transformation layer in transform.py includes comprehensive mapping functions like transform_projects_comprehensive that merge disparate keys (such as "projects" and "projectsOpenSource") into the standardized JSON Resume structure. This flexibility allows the system to handle variations in how LLMs label resume sections.

Where are the Pydantic models for the JSON Resume schema defined?

The Pydantic models are defined in models.py at the repository root. This file contains class definitions for all JSON Resume sections including Basics, Work, Education, Skill, and the top-level JSONResume container model that enforces type safety throughout the application.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →