How to Normalize LLM JSON Output to Standard Format Using `transform.py`
The transform.py module in the interviewstreet/hiring-agent repository converts loosely-structured LLM-generated JSON resumes into strict JSON Resume schema compliance through a pipeline of specialized normalizer functions that handle field mapping, date parsing, and structural validation.
The interviewstreet/hiring-agent repository solves the challenge of inconsistent LLM outputs by providing a robust normalization layer. When large language models generate resume data, they often produce varying field names, nested structures, and date formats that don't conform to standard schemas. The transform.py file contains the essential logic to normalize LLM JSON output to standard format using a series of dedicated transformation functions that route data through section-specific handlers.
The Entry Point: transform_parsed_data
The normalization process begins at line 6 of transform.py with the transform_parsed_data function. This function acts as a router that inspects the top-level keys of the raw LLM payload (e.g., basics, work_experience, education, skills, projects) and delegates each section to its appropriate specialized transformer.
When transform_parsed_data encounters keys it cannot recognize, it implements graceful fallback by returning the original payload untouched. This ensures the pipeline never crashes on unexpected LLM output or schema variations.
Section-Specific Normalization Strategies
Each resume section requires distinct handling to map free-form LLM output to the strict JSON Resume specification.
Standardizing Personal Information with transform_basics
Located at line 25 in transform.py, the transform_basics function resolves inconsistencies in profile data:
- Network name detection: Extracts domain names from URLs (e.g.,
github.com→GitHub) when thenetworkfield is missing - Username extraction: Parses usernames from profile URLs automatically
- Structure enforcement: Guarantees the
profilesfield contains a proper list of profile objects, even if the LLM output provided a single dict or string
Normalizing Work History via transform_work_experience
The transform_work_experience function at line 75 handles the complex task of standardizing employment records:
- Description merging: Joins free-form description arrays into coherent summary strings
- Date range parsing: Converts human-readable month ranges like
"Jan-Mar 2021"into structured date objects - Schema compliance: Outputs uniform dictionaries containing
name,position,url,startDate,endDate,summary, andhighlightsfields
Structuring Education Data with transform_education
At line 42, transform_education extracts degree components and parses education-year ranges. It transforms varied LLM representations of academic credentials into the canonical education objects expected by the JSON Resume schema.
Consolidating Skills Using transform_skills_comprehensive
The transform_skills_comprehensive function at line 48 handles the most variable aspect of LLM outputs:
- Category folding: Merges raw skill lists into the JSON Resume
skillsarray - Language-specific blocks: Normalizes specialized fields like
librariesFrameworks,toolsPlatforms, anddatabases - Type coercion: Handles boolean-or-string skill representations by converting them to standardized skill objects with appropriate keywords
Project Normalization with transform_projects_comprehensive
Found at line 78, this function normalizes both personal projects and open-source contributions. It extracts optional technology tags and skill keywords, ensuring that project entries conform to the JSON Resume schema regardless of whether the LLM provided nested objects or flat lists.
Intelligent Date Parsing with parse_date_range
The parse_date_range utility at line 12 provides sophisticated date normalization:
- Range detection: Converts strings like
"Jan-Mar 2021"or"2020-2021"into ISO-compatiblestartDateandendDatepairs - Open-ended employment: Handles the
"onwards"suffix to represent current positions - Format standardization: Transforms various human-readable formats into consistent date strings
Implementation Example
Here is a complete workflow demonstrating how to normalize LLM JSON output using transform.py:
import json
from transform import transform_parsed_data
from models import JSONResume # pydantic model that matches the JSON-Resume spec
# 1️⃣ Raw LLM output (could be a string or dict)
raw_llm_json = """
{
"basics": {
"name": "Ada Lovelace",
"email": "ada@example.com",
"profiles": [
{"url": "https://github.com/adalove"}
]
},
"work_experience": [
{
"name": "Analytical Engines Inc.",
"position": "Mathematician",
"startDate": "Jan-Mar 1845",
"description": ["Developed early algorithms."]
}
],
"skills": ["Python", "Algorithms"]
}
"""
payload = json.loads(raw_llm_json)
# 2️⃣ Normalise to the standard schema
standardised = transform_parsed_data(payload)
# 3️⃣ Validate against the pydantic model (optional but recommended)
validated_resume = JSONResume(**standardised)
print(json.dumps(validated_resume.dict(), indent=2))
Transformation details:
- Basics processing:
transform_basicsextracts the domaingithub.comand addsnetwork: "GitHub"andusername: "adalove" - Work experience:
transform_work_experiencejoins the description list and parses"Jan-Mar 1845"into"startDate": "Jan 1845"and"endDate": "Mar 1845" - Skills normalization:
transform_skills_comprehensivewraps the raw string list into a single skill category "Programming Languages"
Summary
transform.pyininterviewstreet/hiring-agentprovides the core logic to normalize LLM JSON output to standard format using specialized transformation functions.transform_parsed_data(line 6) routes data to section-specific handlers includingtransform_basics(line 25),transform_work_experience(line 75), andtransform_skills_comprehensive(line 48).- Date parsing is handled by
parse_date_range(line 12), which converts human-readable ranges into ISO-compatible formats. - Graceful degradation ensures unrecognized keys return the original payload rather than causing pipeline failures.
- The output conforms to the JSON Resume schema defined in
models.py, ready for downstream evaluation and scoring.
Frequently Asked Questions
What schema does transform.py normalize to?
According to the source code, transform.py normalizes LLM output to the JSON Resume schema (jsonresume.org), a community-driven open-source initiative to standardize resume formats. The models.py file contains the Pydantic JSONResume class that validates the transformed output against this specification.
How does transform.py handle inconsistent date formats?
The parse_date_range function at line 12 handles various date representations including ranges like "Jan-Mar 2021", "2020-2021", and open-ended periods marked with "onwards". It extracts start and end dates, converting them into standardized formats that comply with the JSON Resume schema's date requirements.
What happens when the LLM output contains unexpected fields?
The normalization pipeline implements graceful fallback: whenever transform_parsed_data encounters keys it cannot recognize, it returns the original payload untouched rather than raising exceptions. This ensures that even when LLMs produce schema variations or novel field names, the pipeline continues functioning and passes data through for manual review.
Where is the validation model defined?
The JSONResume Pydantic model is defined in models.py within the repository root. This model enforces type safety and schema compliance after transform.py completes normalization. The transform_parsed_data function produces a dictionary that, when unpacked into JSONResume(**standardised), validates against the strict JSON Resume specification.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →