How Data Normalization Transforms LLM JSON to JSON Resume Format in interviewstreet/hiring-agent

Data normalization in the hiring-agent repository converts loosely-structured LLM output into strict JSON-Resume schemas through section-specific transformers in transform.py that handle missing fields, date parsing, and profile inference.

The interviewstreet/hiring-agent project solves the challenge of converting unpredictable, free-form JSON generated by large language models into the standardized JSON-Resume format. This transformation pipeline, implemented in pure Python with Pydantic validation, ensures that partial or inconsistently structured LLM outputs become type-safe, standards-compliant resume objects.

The Entry Point: transform_parsed_data

The normalization process begins in transform.py with the transform_parsed_data function (lines 6-95). This orchestrator accepts the raw dictionary returned by the LLM and detects which top-level sections are present—such as basics, work, education, skills, and projects.

For each detected section, the function delegates to a dedicated transformer. Missing sections are left untouched, allowing the pipeline to handle partial resumes gracefully. The final output is a dictionary that aligns exactly with the JSONResume model defined in models.py.

Section-Specific Normalization Strategies

Each resume section implements specialized logic to handle the variations common in LLM output.

Basic Information and Profile Inference (transform_basics)

The transform_basics function (lines 25-52) normalizes the profiles list, which often contains URLs without clear platform identifiers. When a profile contains a URL but no network name, the system extracts the domain using extract_domain_from_url, maps it to a friendly name via get_network_name, and derives the username using extract_username_from_url.

This ensures every profile entry satisfies the JSON-Resume requirement for both network and username fields, even when the LLM provides only a raw link.

Work Experience Structuring (transform_work_experience)

The transform_work_experience function (lines 75-122) guarantees each work item contains the mandatory fields: name, position, url, startDate, endDate, summary, and highlights.

When the LLM provides description as a list of strings, the transformer concatenates it into a single summary string. It also handles date ranges supplied as human-readable strings (e.g., "Jan‑Mar 2021") by delegating to parse_date_range, converting ambiguous temporal data into standardized ISO-like strings.

Education Parsing (transform_education)

In transform_education (lines 42-71), the system extracts degree information and splits it into studyType and area when a comma is present (e.g., "B.Sc., Computer Science" becomes studyType: "B.Sc." and area: "Computer Science").

The function also parses GPA or percentage values into a string score field and uses parse_date_range to transform flexible year specifications (like "2020‑2021") into precise startDate and endDate values.

Skills Aggregation (transform_skills_comprehensive)

The transform_skills_comprehensive function (lines 48-74) merges several possible field names that LLMs might use—skills, librariesFrameworks, toolsPlatforms, and databases—into a unified list of skill categories.

Each category becomes an object with name, level, and keywords properties, matching the JSON-Resume Skill type. The transformer also handles the special case where skills is a flat list of programming languages, wrapping them into the proper categorical structure.

Project Normalization (transform_projects_comprehensive)

For the transform_projects_comprehensive function (lines 78-89), the system normalizes both generic projects and open-source-specific projectsOpenSource fields.

It splits project titles formatted as "My App | Python, Flask" into a clean name and an inferred skills list. The output guarantees all project objects contain name, startDate, endDate, description, highlights, url, technologies, and skills, regardless of the input structure.

Date Parsing and Standardization

Central to the normalization pipeline is parse_date_range (lines 112-179), which accepts diverse human-readable date specifications such as "Jan‑Mar 2021", "2020‑2021", or "Jan‑Mar 2021 onwards".

The function returns a tuple of (startDate, endDate) in consistent ISO-like strings (e.g., "Jan 2021", "Mar 2021", "Present"). By centralizing date logic, every resume section receives a uniform temporal representation, eliminating the formatting inconsistencies common in LLM outputs.

Validation and Final Assembly

After all individual sections are normalized, transform_parsed_data assembles the final dictionary. This structure is then validated against the JSONResume class defined in models.py (lines 201-216).

The Pydantic model enforces type safety and schema compliance, guaranteeing that downstream consumers—such as CSV exporters or text conversion utilities—receive a well-typed resume object. This validation step catches any edge cases missed during transformation, ensuring 100% compatibility with the JSON-Resume standard.

Practical Implementation Example


# Example raw LLM output (shortened for clarity)

raw_llm = {
    "basics": {
        "name": "Ada Lovelace",
        "email": "ada@example.com",
        "profiles": [
            {"url": "https://github.com/ada", "username": "ada"}
        ]
    },
    "work_experience": [
        {
            "name": "Analytical Engines Inc.",
            "position": "Research Engineer",
            "description": ["Built early computing prototypes."],
            "startDate": "Jan‑Mar 1845"
        }
    ],
    "skills": ["Python", "C++"],
    "education": [
        {"degree": "B.Sc., Computer Science", "years": "1840‑1842", "gpa": 3.9}
    ]
}

# Normalisation step

from transform import transform_parsed_data
normalized = transform_parsed_data(raw_llm)

# Build a JSON‑Resume object (validation via Pydantic)

from models import JSONResume
resume = JSONResume(**normalized)

print(resume.json(indent=2))

Output (abridged):

{
  "basics": {
    "name": "Ada Lovelace",
    "email": "ada@example.com",
    "profiles": [
      {
        "url": "https://github.com/ada",
        "network": "GitHub",
        "username": "ada"
      }
    ]
  },
  "work": [
    {
      "name": "Analytical Engines Inc.",
      "position": "Research Engineer",
      "summary": "Built early computing prototypes.",
      "startDate": "Jan 1845",
      "endDate": null,
      "highlights": []
    }
  ],
  "skills": [
    {
      "name": "Programming Languages",
      "level": null,
      "keywords": ["Python", "C++"]
    }
  ],
  "education": [
    {
      "institution": "",
      "area": "Computer Science",
      "studyType": "B.Sc.",
      "startDate": "1840-01",
      "endDate": "1842-12",
      "score": "3.9",
      "courses": []
    }
  ]
}

Summary

  • Data normalization in interviewstreet/hiring-agent transforms unstructured LLM JSON into the strict JSON-Resume schema through coordinated transformations in transform.py.
  • Section-specific transformers handle unique normalization challenges: profile inference for URLs, date range parsing for experiences, and degree string splitting for education.
  • Centralized date parsing via parse_date_range standardizes temporal data into ISO-like formats, handling human-readable ranges like "Jan‑Mar 2021".
  • Pydantic validation through the JSONResume model in models.py ensures type safety and guarantees downstream compatibility with resume processing tools.
  • The pipeline gracefully handles partial resumes and missing fields, leaving undetected sections untouched while normalizing available data.

Frequently Asked Questions

How does the system handle missing network names in profile URLs?

The transform_basics function detects profiles containing URLs but lacking network identifiers. It extracts the domain via extract_domain_from_url, maps it to a friendly platform name using get_network_name, and derives the username through extract_username_from_url. This ensures every profile entry meets the JSON-Resume requirement for both network and username fields.

What happens when the LLM provides dates in non-standard formats?

The parse_date_range utility accepts diverse human-readable date strings such as "Jan‑Mar 2021", "2020‑2021", or "Jan‑Mar 2021 onwards". It parses these into standardized tuples of (startDate, endDate) in ISO-like format, allowing the system to normalize temporal data from work experience, education, and project sections consistently.

Can the pipeline process resumes with missing sections?

Yes. The transform_parsed_data entry point detects which sections are present and processes only those found in the input. Missing sections are left untouched in the output dictionary, allowing the caller to handle partial resumes. The Pydantic JSONResume model validates whatever data is present, making the system robust against incomplete LLM outputs.

How are skills normalized when the LLM uses different field names?

The transform_skills_comprehensive function merges multiple possible field names—including skills, librariesFrameworks, toolsPlatforms, and databases—into a unified list of skill categories. It also handles flat string lists by wrapping them into the proper categorical structure with name, level, and keywords properties required by the JSON-Resume schema.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →