# How Hiring Agent Uses the JSON Resume Format: Schema Definition and Data Transformation

> Learn how Hiring Agent uses the JSON resume format to normalize LLM outputs. Discover schema definition and data transformation with Pydantic models.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-06

---

**Hiring Agent consumes loosely structured LLM outputs and normalizes them into the strict JSON Resume schema using Pydantic models defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) and transformation utilities in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py).**

The interviewstreet/hiring-agent repository implements a machine-readable approach to candidate evaluation by adopting the open-source JSON Resume format. This structured schema allows the system to parse unstructured PDF resumes into standardized data models that power automated scoring and evaluation pipelines.

## Understanding the JSON Resume Schema Structure

### Core Sections Defined in models.py

The JSON Resume format is implemented through Pydantic models in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) ([source](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)). The schema organizes candidate data into these key sections:

- **Basics**: Core personal information including `name`, `email`, `phone`, `url`, `summary`, `location`, and `profiles`
- **Work**: Employment history with fields like `name` (company), `position`, `url`, `startDate`, `endDate`, `summary`, and `highlights`
- **Volunteer**: Voluntary activities containing `organization`, `position`, `url`, `startDate`, `endDate`, `summary`, and `highlights`
- **Education**: Academic background with `institution`, `url`, `area`, `studyType`, `startDate`, `endDate`, `score`, and `courses`
- **Awards**: Honors and recognitions including `title`, `date`, `awarder`, and `summary`
- **Certificates**: Certifications earned with `name`, `date`, `issuer`, and `url`
- **Publications**: Published works containing `name`, `publisher`, `releaseDate`, `url`, and `summary`
- **Skills**: Skill categories with `name`, `level`, and `keywords`
- **Languages**: Language proficiency with `language` and `fluency`
- **Interests**: Personal interests with `name` and `keywords`
- **References**: Professional references with `name` and `reference`
- **Projects**: Projects (open-source or personal) with `name`, `startDate`, `endDate`, `description`, `highlights`, `url`, `technologies`, and `skills`

These models combine into the top-level `JSONResume` class:

```python
class JSONResume(BaseModel):
    basics: Optional[Basics] = None
    work: Optional[List[Work]] = None
    volunteer: Optional[List[Volunteer]] = None
    education: Optional[List[Education]] = None
    awards: Optional[List[Award]] = None
    certificates: Optional[List[Certificate]] = None
    publications: Optional[List[Publication]] = None
    skills: Optional[List[Skill]] = None
    languages: Optional[List[Language]] = None
    interests: Optional[List[Interest]] = None
    references: Optional[List[Reference]] = None
    projects: Optional[List[Project]] = None

```

## Transforming Raw LLM Data into JSON Resume Format

### The Normalization Pipeline

When the LLM returns loosely-structured JSON from [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) processing, Hiring Agent normalizes it using functions in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) ([source](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py)). The main entry point is `transform_parsed_data()`, which executes a three-step process:

1. Detects which top-level sections are present in the raw payload
2. Calls dedicated helpers like `transform_basics`, `transform_work_experience`, and `transform_education` to rewrite each section into exact field names and data shapes expected by the schema
3. Instantiates a `JSONResume` object for type-safe downstream access by the evaluator and CSV exporter

### Field-Level Transformations

The transformation layer handles several specific normalization tasks to align raw LLM outputs with the JSON Resume format:

- **Social Network Detection**: Uses `extract_domain_from_url` and `get_network_name` to infer platform names from profile URLs
- **Username Extraction**: Applies `extract_username_from_url` to parse usernames from social links
- **Date Normalization**: Employs `parse_date_range` to standardize date formats into consistent `startDate` and `endDate` fields
- **Skill Structuring**: Converts free-form skill listings into structured `Skill` objects via `transform_skills_comprehensive`
- **Project Merging**: Unifies disparate project keys like "projects" and "projectsOpenSource" into a uniform list using `transform_projects_comprehensive`

## Practical Implementation Examples

### Parsing and Normalizing Raw LLM Output

```python
from transform import transform_parsed_data
from models import JSONResume

# `raw_payload` is the loosely-structured JSON returned by the LLM

raw_payload = {
    "basics": {"name": "Alice", "email": "alice@example.com"},
    "work_experience": [
        {"company": "Acme", "title": "Engineer", "startDate": "Jan-Mar 2020"}
    ],
    "skills": ["Python", "Docker", "Kubernetes"]
}

# Normalise the payload to the JSON-Resume schema

norm_dict = transform_parsed_data(raw_payload)

# Build a strongly-typed JSONResume object

resume = JSONResume(**norm_dict)

print(resume.basics.name)          # → Alice

print(resume.work[0].position)     # → Engineer

print(resume.skills[0].keywords)   # → ['Python', 'Docker', 'Kubernetes']

```

### Converting to Plain Text for LLM Prompts

```python
from transform import convert_json_resume_to_text

text_blob = convert_json_resume_to_text(resume)
print(text_blob)   # Human-readable sections such as "=== BASIC INFORMATION ==="

```

### Exporting Evaluation Data to CSV

```python
from evaluator import EvaluationData
from transform import transform_evaluation_response

# Assume `resume` (JSONResume), `github_data` (dict from github.py) and `eval_data` (EvaluationData)

csv_row = transform_evaluation_response(
    file_name="alice_resume.pdf",
    resume_data=resume,
    github_data=github_data,
    evaluation=eval_data,
)

# `csv_row` is a plain dict ready for writing to a CSV file (handled in `score.py`).

```

## Summary

- Hiring Agent implements the open-source **JSON Resume format** through Pydantic models in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), covering sections from Basics to Projects
- The `transform_parsed_data()` function in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) normalizes unstructured LLM outputs into the strict schema
- Specialized helpers handle URL parsing, date normalization, and skill structuration to ensure data consistency
- The `JSONResume` object enables type-safe access for the evaluator and CSV exporter components
- This transformation pipeline bridges the gap between unstructured PDF inputs and structured evaluation data

## Frequently Asked Questions

### What is the JSON Resume format?

The JSON Resume format is an open-source, machine-readable schema for representing curriculum vitae data. It defines standardized fields for personal information, work experience, education, skills, and projects, enabling consistent parsing across different systems and tools.

### How does Hiring Agent handle inconsistent date formats in resumes?

Hiring Agent uses the `parse_date_range` helper function in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) to normalize various date string formats into standardized `startDate` and `endDate` fields. This ensures that temporal data from different resume styles conforms to the JSON Resume schema requirements.

### Can Hiring Agent process non-standard section names from LLM outputs?

Yes, the transformation layer in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) includes comprehensive mapping functions like `transform_projects_comprehensive` that merge disparate keys (such as "projects" and "projectsOpenSource") into the standardized JSON Resume structure. This flexibility allows the system to handle variations in how LLMs label resume sections.

### Where are the Pydantic models for the JSON Resume schema defined?

The Pydantic models are defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) at the repository root. This file contains class definitions for all JSON Resume sections including `Basics`, `Work`, `Education`, `Skill`, and the top-level `JSONResume` container model that enforces type safety throughout the application.