# How the JSON Resume Schema Defines the Extracted Data Model in Hiring Agent

> Discover how the JSON Resume schema defines the extracted data model in Hiring Agent. This repo uses Pydantic models for type-safe, validated data structures, ensuring consistency in workflows.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: data-model
- Published: 2026-07-05

---

**The Hiring Agent repository implements the JSON Resume specification as a hierarchy of Pydantic models in [`main/models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/models.py), providing a type-safe, validated data structure that ensures consistency across parsing, transformation, and evaluation workflows.**

The InterviewStreet hiring-agent project utilizes the JSON Resume schema as its canonical data model for representing candidate résumés. By mapping this open specification to a complete set of Pydantic classes, the codebase guarantees that every processing stage—from PDF ingestion to final scoring—operates on a fully typed and validated object graph.

## Schema Definition in models.py

The core data contract resides in [`main/models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/models.py), where the `JSONResume` class serves as the root aggregation point for all résumé sections. This class composes nested models such as `Basics`, `Work`, `Education`, and `Skill`, each mirroring the official JSON Resume specification fields.

When the system receives raw résumé data—whether extracted from a PDF, parsed from raw JSON, or generated by an LLM—it instantiates the top-level model:

```python
from models import JSONResume

resume = JSONResume(**data)

```

Pydantic validates required types, applies default values, and raises descriptive errors if the payload deviates from the schema. This strict validation guarantees downstream components receive a predictable data structure regardless of the input source.

## Parsing and Validation Patterns

Incoming résumé data flows through a consistent validation checkpoint. Functions across the repository instantiate `JSONResume` objects immediately after extraction to confirm compliance before further processing.

In [`main/pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/pdf.py), for example, the pipeline extracts JSON text from PDF documents and immediately constructs a `JSONResume` instance to verify the content adheres to the expected schema. This prevents malformed or incomplete data from propagating into scoring algorithms.

Similarly, scoring logic in [`main/score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/score.py) relies on the `is_valid_resume_data()` function to check for a valid `JSONResume` instance before computing category scores, ensuring that only properly structured résumés enter the evaluation pipeline.

## Downstream Component Integration

The validated `JSONResume` object acts as the single source of truth for multiple system components:

### Data Transformation in transform.py

The [`main/transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/transform.py) module extracts specific fields—such as `resume_data.basics`, `resume_data.work`, and `resume_data.skills`—to populate CSV export columns. Direct property access on the typed instance eliminates the need for defensive dictionary lookups and provides IDE autocomplete support.

### PDF Processing in pdf.py

After extracting raw JSON from PDF documents, [`main/pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/pdf.py) constructs a `JSONResume` object to validate the resume complies with the schema before passing it to downstream consumers.

### Resume Scoring in score.py

The scoring engine validates that a resume can be processed using `is_valid_resume_data()`, which checks for a properly instantiated `JSONResume` object. It then accesses specific sections—such as work history and skills—to compute category scores returned as an `EvaluationData` instance.

### Report Generation in evaluator.py

The [`main/evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/evaluator.py) module receives a `JSONResume` object alongside LLM-generated evaluation data to produce the final hiring report, combining structured résumé facts with qualitative assessments.

## Implementation Examples

The following patterns demonstrate how the JSON Resume schema integrates into different processing stages.

**Loading and validating raw JSON:**

```python
import json
from models import JSONResume

raw = json.loads(open("candidate_resume.json").read())
resume = JSONResume(**raw)          # Validation happens here

print(resume.basics.name)           # → candidate’s full name

```

**Accessing structured data for transformation:**

```python
from models import JSONResume
from transform import transform_evaluation_response

def demo():
    resume = JSONResume(**some_parsed_data)
    csv_row = transform_evaluation_response(
        file_name="resume.pdf",
        resume_data=resume,
        github_data={},
        evaluation=None,
    )
    print(csv_row["github_url"], csv_row["linkedin_url"])

```

**Validating before scoring:**

```python
from models import JSONResume, EvaluationData
from score import is_valid_resume_data, compute_scores

if is_valid_resume_data(resume):
    scores = compute_scores(resume)    # Returns an EvaluationData instance

    print(scores.scores.open_source.score)

```

## Key Repository Files

| File | Role |
|------|------|
| [`main/models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/models.py) | Defines the `JSONResume` Pydantic hierarchy and section models |
| [`main/transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/transform.py) | Converts `JSONResume` instances into CSV-ready column dictionaries |
| [`main/pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/pdf.py) | Extracts JSON from PDFs and validates against the `JSONResume` schema |
| [`main/score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/score.py) | Validates résumé data integrity and computes evaluation scores |
| [`main/evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/evaluator.py) | Combines `JSONResume` data with LLM outputs for final reporting |

## Summary

- The **JSON Resume schema** is implemented as a type-safe Pydantic model hierarchy in [`main/models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/models.py), providing the central data contract for the entire Hiring Agent pipeline.
- **Validation occurs at ingestion**: Functions instantiate `JSONResume(**data)` to enforce schema compliance immediately after data extraction from PDFs or JSON sources.
- **Consistent consumption**: Components in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py), [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py), and [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) access typed properties like `resume_data.basics` and `resume_data.work` without defensive coding.
- **Robustness**: Centralized schema definitions localize changes to the JSON Resume specification and prevent malformed data from reaching scoring algorithms.

## Frequently Asked Questions

### What specific sections does the JSONResume model include?

The `JSONResume` class aggregates section models including `Basics` (contact info), `Work` (employment history), `Education` (academic credentials), and `Skill` (competencies), each matching the official JSON Resume specification fields. This structure ensures every candidate profile contains the same standardized fields regardless of original document format.

### How does the system handle invalid or malformed résumé data?

When `JSONResume(**data)` is called, Pydantic validates types and required fields against the schema. If the incoming payload lacks required fields or contains type mismatches, Pydantic raises a validation error immediately. This prevents malformed data from reaching scoring functions like `compute_scores()` in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py).

### Why use Pydantic instead of standard dictionaries for résumé data?

Pydantic provides runtime type checking, IDE autocomplete, and automatic validation. By defining the JSON Resume schema as Pydantic models in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), the codebase eliminates key-access errors and ensures that properties like `resume_data.basics.name` exist and contain the expected string type, which is critical for reliable CSV generation in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py).

### Which components actually consume the JSONResume object?

The `JSONResume` object flows through multiple stages: [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) uses it for initial validation, [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) extracts fields for CSV export, [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) validates it via `is_valid_resume_data()` before scoring, and [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) combines it with LLM evaluation results to generate final hiring reports.