How to Convert Unstructured LLM JSON to JSON Resume Format in Python

The hiring-agent repository converts raw LLM output into structured JSON Resume format using specialized transformer functions in transform.py that normalize dates, profiles, and sections before validating against Pydantic models in models.py.

The interviewstreet/hiring-agent repository handles the conversion of unstructured LLM JSON to JSON Resume through a pipeline of specialized transformation functions. This process ensures that arbitrary JSON payloads from language models conform to the standardized JSON Resume schema used for structured candidate evaluation. The transformation logic resides primarily in transform.py with schema definitions in models.py.

Entry Point: The transform_parsed_data Function

The conversion begins in transform.py at the transform_parsed_data function, defined at line 6. This function accepts a raw dictionary from the LLM and orchestrates the entire transformation pipeline.

def transform_parsed_data(parsed_data: Dict) -> Dict:

The function first validates that the input is a dictionary, then detects which resume sections are present by checking for keys such as basics, work, education, skills, and others. Depending on the detected sections, it dispatches to specialized helper functions that reshape the data to match the JSON Resume specification. The high-level mapping logic spans lines 6 through 39 in transform.py.

Section-Specific Transformation Logic

Each resume section requires distinct normalization logic to handle the irregular structure of LLM-generated JSON.

Normalizing Basics and Social Profiles

The transform_basics function (lines 25-52 in transform.py) processes personal information and profile URLs. It extracts domain information using extract_domain_from_url and determines the social network name via get_network_name. The function also derives usernames from URLs using extract_username_from_url (lines 154-172), ensuring that profile entries conform to the JSON Resume basics schema with proper network identification.

Structuring Work Experience

For employment history, the transform_work_experience function (lines 75-121) flattens description lists and parses date ranges using parse_date_range. It produces standardized dictionaries containing name, position, url, startDate, endDate, summary, and highlights fields. This ensures that varied LLM date formats like "Jan-Mar 2021" or "2020-2021" resolve to consistent ISO-style strings.

Standardizing Education History

The transform_education function (lines 141-172) handles academic entries by extracting degree information, GPA or percentage values, and date ranges. It intelligently splits degree strings into studyType and area components while formatting dates through the shared parse_date_range utility.

Processing Skills and Technologies

Complex skill structures are handled by transform_skills_comprehensive (lines 248-374), which accepts multiple input shapes including skills, librariesFrameworks, toolsPlatforms, and databases. This function consolidates varied taxonomies into a unified list of skill objects following the JSON Resume schema, ensuring that technologies are properly categorized regardless of how the LLM structured them.

Mapping Projects and Achievements

The transform_projects_comprehensive function (lines 378-409) processes both projects and projectsOpenSource sections. It parses "name|skill list" strings and normalizes technologies fields to produce consistent project objects. For awards and achievements, transform_achievements (lines 175-194) normalizes titles, dates, awarder names, and summaries into the standard awards schema.

Handling Volunteer Organizations

The transform_organizations function (lines 124-138) maps organizational involvement to the JSON Resume volunteer schema, standardizing fields for organization name, position, URL, and date ranges.

Utility Functions for Data Normalization

The transformation relies on robust utility functions to handle messy LLM output.

URL Parsing and Network Detection

The extract_domain_from_url function (lines 98-105) identifies the base domain from profile URLs, while extract_username_from_url extracts clean usernames. These utilities enable automatic network detection (GitHub, LinkedIn, etc.) without requiring explicit network labels in the input JSON.

Date Range Parsing

The parse_date_range function (lines 412-484) recognizes diverse date formats including "Jan-Mar 2021", "2020-2021", "Jan 2021 onwards", and "Present". It returns normalized start and end date strings that comply with the JSON Resume date expectations.

Schema Validation with Pydantic

After transformation, the resulting dictionary is validated against the JSONResume Pydantic model defined in models.py (lines 1-167). This model provides typed access to fields such as resume.basics.name and resume.work[0].position while enforcing schema compliance. The validation step catches structural inconsistencies before the data proceeds to downstream evaluation in score.py and evaluator.py.

from transform import transform_parsed_data
from models import JSONResume

# Example of raw LLM output (unstructured)

raw_llm_json = {
    "basics": {
        "name": "Ada Lovelace",
        "email": "ada@example.com",
        "profiles": [
            {"url": "https://github.com/ada"},
            {"url": "https://linkedin.com/in/ada-lovelace"}
        ]
    },
    "work_experience": [
        {
            "name": "Analytical Engines Inc.",
            "position": "Senior Engineer",
            "startDate": "Jan-Mar 1842",
            "endDate": "Present",
            "highlights": ["Implemented early compiler"]
        }
    ],
    "skills": ["Python", "C++"],
    "librariesFrameworks": ["TensorFlow", "PyTorch"]
}

# Transform to JSON Resume dict

resume_dict = transform_parsed_data(raw_llm_json)

# Validate against the Pydantic model (optional but recommended)

resume = JSONResume(**resume_dict)

print(resume.json(indent=2))

Summary

  • The transform_parsed_data function in transform.py serves as the main entry point for converting unstructured LLM JSON to JSON Resume format.
  • Section-specific transformers handle normalization of basics, work experience, education, skills, projects, and volunteer work with specialized logic for each schema requirement.
  • Utility functions extract domains and usernames from URLs and parse flexible date ranges into standardized formats.
  • Pydantic validation via the JSONResume model in models.py ensures type safety and schema compliance before the data reaches evaluation functions.
  • Direct field copies occur for sections like certificates, publications, languages, interests, references, and meta that already conform to the expected structure.

Frequently Asked Questions

What is the main entry point for converting LLM JSON to JSON Resume?

The transform_parsed_data function in transform.py (line 6) serves as the primary entry point. It validates the input dictionary and dispatches to specialized transformation functions based on which resume sections are detected in the unstructured LLM output.

How does the transformer handle inconsistent date formats from LLM output?

The parse_date_range function (lines 412-484 in transform.py) recognizes multiple natural language date formats including "Jan-Mar 2021", "2020-2021", and "Present". It normalizes these variations into consistent date strings suitable for the JSON Resume schema.

What happens to sections that already match the JSON Resume schema?

Sections such as certificates, publications, languages, interests, references, and meta are passed through unchanged without transformation. The pipeline assumes these sections already conform to the standard schema and copies them directly into the output dictionary.

Why is Pydantic validation used after transformation?

The JSONResume model in models.py provides runtime type checking and schema validation. This ensures that the transformed data strictly adheres to the JSON Resume specification before being consumed by evaluation logic in score.py and evaluator.py, preventing downstream errors from malformed data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →