How the JSON Resume Schema Defines the Extracted Data Model in Hiring Agent

The Hiring Agent repository implements the JSON Resume specification as a hierarchy of Pydantic models in main/models.py, providing a type-safe, validated data structure that ensures consistency across parsing, transformation, and evaluation workflows.

The InterviewStreet hiring-agent project utilizes the JSON Resume schema as its canonical data model for representing candidate résumés. By mapping this open specification to a complete set of Pydantic classes, the codebase guarantees that every processing stage—from PDF ingestion to final scoring—operates on a fully typed and validated object graph.

Schema Definition in models.py

The core data contract resides in main/models.py, where the JSONResume class serves as the root aggregation point for all résumé sections. This class composes nested models such as Basics, Work, Education, and Skill, each mirroring the official JSON Resume specification fields.

When the system receives raw résumé data—whether extracted from a PDF, parsed from raw JSON, or generated by an LLM—it instantiates the top-level model:

from models import JSONResume

resume = JSONResume(**data)

Pydantic validates required types, applies default values, and raises descriptive errors if the payload deviates from the schema. This strict validation guarantees downstream components receive a predictable data structure regardless of the input source.

Parsing and Validation Patterns

Incoming résumé data flows through a consistent validation checkpoint. Functions across the repository instantiate JSONResume objects immediately after extraction to confirm compliance before further processing.

In main/pdf.py, for example, the pipeline extracts JSON text from PDF documents and immediately constructs a JSONResume instance to verify the content adheres to the expected schema. This prevents malformed or incomplete data from propagating into scoring algorithms.

Similarly, scoring logic in main/score.py relies on the is_valid_resume_data() function to check for a valid JSONResume instance before computing category scores, ensuring that only properly structured résumés enter the evaluation pipeline.

Downstream Component Integration

The validated JSONResume object acts as the single source of truth for multiple system components:

Data Transformation in transform.py

The main/transform.py module extracts specific fields—such as resume_data.basics, resume_data.work, and resume_data.skills—to populate CSV export columns. Direct property access on the typed instance eliminates the need for defensive dictionary lookups and provides IDE autocomplete support.

PDF Processing in pdf.py

After extracting raw JSON from PDF documents, main/pdf.py constructs a JSONResume object to validate the resume complies with the schema before passing it to downstream consumers.

Resume Scoring in score.py

The scoring engine validates that a resume can be processed using is_valid_resume_data(), which checks for a properly instantiated JSONResume object. It then accesses specific sections—such as work history and skills—to compute category scores returned as an EvaluationData instance.

Report Generation in evaluator.py

The main/evaluator.py module receives a JSONResume object alongside LLM-generated evaluation data to produce the final hiring report, combining structured résumé facts with qualitative assessments.

Implementation Examples

The following patterns demonstrate how the JSON Resume schema integrates into different processing stages.

Loading and validating raw JSON:

import json
from models import JSONResume

raw = json.loads(open("candidate_resume.json").read())
resume = JSONResume(**raw)          # Validation happens here

print(resume.basics.name)           # → candidate’s full name

Accessing structured data for transformation:

from models import JSONResume
from transform import transform_evaluation_response

def demo():
    resume = JSONResume(**some_parsed_data)
    csv_row = transform_evaluation_response(
        file_name="resume.pdf",
        resume_data=resume,
        github_data={},
        evaluation=None,
    )
    print(csv_row["github_url"], csv_row["linkedin_url"])

Validating before scoring:

from models import JSONResume, EvaluationData
from score import is_valid_resume_data, compute_scores

if is_valid_resume_data(resume):
    scores = compute_scores(resume)    # Returns an EvaluationData instance

    print(scores.scores.open_source.score)

Key Repository Files

File Role
main/models.py Defines the JSONResume Pydantic hierarchy and section models
main/transform.py Converts JSONResume instances into CSV-ready column dictionaries
main/pdf.py Extracts JSON from PDFs and validates against the JSONResume schema
main/score.py Validates résumé data integrity and computes evaluation scores
main/evaluator.py Combines JSONResume data with LLM outputs for final reporting

Summary

  • The JSON Resume schema is implemented as a type-safe Pydantic model hierarchy in main/models.py, providing the central data contract for the entire Hiring Agent pipeline.
  • Validation occurs at ingestion: Functions instantiate JSONResume(**data) to enforce schema compliance immediately after data extraction from PDFs or JSON sources.
  • Consistent consumption: Components in transform.py, score.py, and evaluator.py access typed properties like resume_data.basics and resume_data.work without defensive coding.
  • Robustness: Centralized schema definitions localize changes to the JSON Resume specification and prevent malformed data from reaching scoring algorithms.

Frequently Asked Questions

What specific sections does the JSONResume model include?

The JSONResume class aggregates section models including Basics (contact info), Work (employment history), Education (academic credentials), and Skill (competencies), each matching the official JSON Resume specification fields. This structure ensures every candidate profile contains the same standardized fields regardless of original document format.

How does the system handle invalid or malformed résumé data?

When JSONResume(**data) is called, Pydantic validates types and required fields against the schema. If the incoming payload lacks required fields or contains type mismatches, Pydantic raises a validation error immediately. This prevents malformed data from reaching scoring functions like compute_scores() in score.py.

Why use Pydantic instead of standard dictionaries for résumé data?

Pydantic provides runtime type checking, IDE autocomplete, and automatic validation. By defining the JSON Resume schema as Pydantic models in models.py, the codebase eliminates key-access errors and ensures that properties like resume_data.basics.name exist and contain the expected string type, which is critical for reliable CSV generation in transform.py.

Which components actually consume the JSONResume object?

The JSONResume object flows through multiple stages: pdf.py uses it for initial validation, transform.py extracts fields for CSV export, score.py validates it via is_valid_resume_data() before scoring, and evaluator.py combines it with LLM evaluation results to generate final hiring reports.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →