How transform.py Normalizes Loose LLM JSON to JSON Resume Format
The transform.py module in the interviewstreet/hiring-agent repository converts unstructured LLM-generated JSON into strict JSON-Resume schema compliance through a pipeline of specialized transformer functions that handle everything from date parsing to profile extraction.
The hiring-agent project bridges the gap between raw LLM outputs and structured resume data. When large language models return loosely-structured JSON with inconsistent field names and formats, transform.py reshapes these payloads into valid JSON-Resume documents that downstream evaluation tools can consume reliably.
Entry Point and Schema Detection
The normalization process begins with transform_parsed_data in transform.py (lines 6-39). This function inspects the incoming dictionary and determines whether it contains a complete resume structure or just isolated sections.
If the input contains a top-level basics section alongside other fields, the function orchestrates a full transformation by delegating to specialized transformers for each JSON-Resume category. When only a single section appears (such as only work or only skills), the code falls back to a minimal document construction path (lines 40-87), wrapping that solitary section into a valid JSON-Resume container.
Normalizing the Basics Section
The transform_basics function (lines 25-52) handles personal information and profile links. It performs critical cleanup on the profiles list, extracting usernames from URLs via extract_username_from_url and inferring the social network name when missing. This ensures that LinkedIn, GitHub, and other profile URLs become properly structured objects with network identifiers and usernames rather than raw strings.
Processing Work Experience
Work history normalization occurs in transform_work_experience, which guarantees that every entry contains the mandatory JSON-Resume fields: name, position, url, startDate, endDate, summary, and highlights.
The function handles heterogeneous input formats, converting description fields that arrive as lists into consolidated summary strings. It delegates date range parsing to parse_date_range, interpreting human-readable formats like "Jan-Mar 2021" or "2020 onwards" into ISO-compliant date strings.
Handling Education and Date Ranges
The transform_education function converts academic entries into the required schema fields: institution, url, area, studyType, startDate, endDate, score, and courses. It intelligently splits composite degree strings (like "B.Sc., Computer Science") into separate study type and area components.
Supporting this process, the parse_date_range utility (lines 112-184) acts as the date normalization engine, converting vague temporal expressions into concrete start and end dates suitable for JSON-Resume.
Skills and Projects Transformation
For technical competencies, transform_skills_comprehensive accepts multiple possible input keys including skills, librariesFrameworks, toolsPlatforms, and databases. When skills arrives as a flat list of strings, it wraps them under a default "Programming Languages" category. The helper transform_skills builds category objects containing name, level, and keywords arrays.
Project normalization through transform_projects_comprehensive handles both standard projects and projectsOpenSource fields. It splits composite project titles such as "MyApp | Python, Flask" into clean names and technology lists, ensuring output fields include name, startDate, endDate, description, highlights, url, technologies, and skills.
Volunteer and Achievements
The transform_organizations function maps generic organization data into the JSON-Resume volunteer format, while transform_achievements normalizes award objects to contain title, date, awarder, and summary fields.
Practical Implementation Examples
The following example demonstrates how transform_parsed_data handles a sparse LLM output with inconsistent field names:
from transform import transform_parsed_data
# Example of a loose LLM output (only a subset of fields)
loose_json = {
"name": "Jane Doe",
"email": "jane@example.com",
"work": [
{
"position": "Software Engineer",
"name": "Acme Corp",
"startDate": "Jan‑Mar 2020",
"endDate": "Present",
"description": ["Built APIs", "Improved performance"]
}
],
"skills": ["Python", "SQL", "Docker"],
"education": [
{"degree": "B.Sc., Computer Science", "institution": "MIT", "years": "2015‑2019"}
]
}
resume_json = transform_parsed_data(loose_json)
# `resume_json` now conforms to JSON‑Resume:
# {
# "basics": {"name": "Jane Doe", "email": "jane@example.com", ...},
# "work": [
# {"name": "Acme Corp", "position": "Software Engineer",
# "startDate": "Jan 2020", "endDate": "Present",
# "summary": "Built APIs Improved performance", "highlights": []}
# ],
# "skills": [
# {"name": "Programming Languages", "level": null,
# "keywords": ["Python", "SQL", "Docker"]}
# ],
# "education": [
# {"institution": "MIT", "studyType": "B.Sc.", "area": "Computer Science",
# "startDate": "2015-01", "endDate": "2019-12", "score": null, "courses": []}
# ]
# }
For inputs that already contain structured sections like basics and projectsOpenSource, the transformation handles profile inference and technology extraction:
# Using the same function with a richer input that already contains a `basics` block
loose_json2 = {
"basics": {
"name": "John Smith",
"email": "john@smith.io",
"profiles": [{"url": "https://github.com/johnsmith"}]
},
"projectsOpenSource": [
{"name": "my-tool | Python, Click", "url": "https://github.com/johnsmith/my-tool"}
]
}
resume_json2 = transform_parsed_data(loose_json2)
# The `profiles` entry now has `network: "GitHub"` and `username: "johnsmith"`
# The project is normalised with `technologies` = ["Python", "Click"] and `skills` extracted.
Target Schema and Data Models
The transformation pipeline targets the dataclasses defined in models.py, which represent the JSON-Resume specification. These models enforce type safety and required fields, ensuring that the output from transform.py satisfies the strict schema requirements for downstream CSV export and evaluation components.
Summary
transform_parsed_dataintransform.pydetects whether LLM output contains full or partial resume data and routes accordingly between complete transformation and minimal document construction.- Specialized transformers like
transform_basics,transform_work_experience, andtransform_educationenforce JSON-Resume schema compliance through explicit field mapping. - The
parse_date_rangeutility converts human-readable date ranges into standardized ISO formats suitable for the JSON-Resume specification. - Skills and projects undergo comprehensive normalization to handle flat lists, composite strings, and multiple input key variants.
- The pipeline outputs a dictionary that strictly conforms to the JSON-Resume hierarchy, making it safe for downstream processing.
Frequently Asked Questions
How does transform.py handle incomplete LLM outputs?
When the input contains only a single top-level section (such as only work or only skills), transform_parsed_data detects this condition and constructs a minimal JSON-Resume document containing just that section, preventing downstream validation errors while preserving the available data.
What date formats does the hiring-agent parser support?
The parse_date_range function interprets various human-readable formats including ranges like "Jan-Mar 2021", "2020-2021", and ongoing periods marked as "2022 onwards", converting them into ISO-compliant date strings suitable for the JSON-Resume schema.
How are social media profiles normalized in the basics section?
The transform_basics function cleans the profiles array by inferring the network name from URLs when missing and extracting usernames via extract_username_from_url, ensuring that profile links become structured objects with network, username, and url fields.
Can transform.py handle mixed skill input formats?
Yes, transform_skills_comprehensive accepts multiple input keys (skills, librariesFrameworks, toolsPlatforms, databases) and normalizes both flat string lists and category objects, wrapping simple lists under a default "Programming Languages" category when necessary.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →