How to Support New Resume Sections in Hiring-Agent: Adding Projects, Awards, and Custom Fields

Add new resume sections by extending the JSONResume Pydantic model, creating a Jinja prompt template, and wiring the extraction logic in PDFExtractor—no changes to the core scoring engine required.

The interviewstreet/hiring-agent repository parses PDF résumés using LLM-driven prompts and converts them into structured JSON before scoring candidates. To support new resume sections like Projects, Awards, or custom fields, you modify the data models, prompts, and extraction pipeline while reusing the existing transformation and scoring infrastructure.

Understanding the Resume Parsing Pipeline

The Hiring-Agent follows a six-stage data flow that makes adding sections straightforward:

  1. PDF Ingestion – PDFHandler.extract_text_from_pdf extracts raw text from the uploaded file.
  2. Section Extraction – PDFExtractor loads Jinja templates (e.g., projects.jinja) and calls the LLM via _call_llm_for_section to parse specific segments.
  3. Model Assembly – Parsed sections populate a JSONResume Pydantic model instance.
  4. Text Transformation – transform.convert_json_resume_to_text renders the model as markdown-style text for evaluation.
  5. CSV Generation – The same transformation logic flattens the model into a spreadsheet row.
  6. AI Scoring – score._evaluate_resume concatenates the text representation and sends it to the Evaluator.

Because each stage operates on the JSONResume model rather than hard-coded fields, you can introduce new attributes without touching the scoring algorithm.

Step 1: Define the Data Model in models.py

First, declare the new section's structure in models.py by creating a Pydantic class and adding it to the JSONResume root model.


# models.py

from typing import List, Optional
from pydantic import BaseModel

class ProjectSection(BaseModel):
    name: str
    description: Optional[str] = None
    url: Optional[str] = None
    startDate: Optional[str] = None
    endDate: Optional[str] = None

class JSONResume(BaseModel):
    basics: Optional[BasicsSection] = None
    work: Optional[List[WorkSection]] = None
    education: Optional[List[EducationSection]] = None
    skills: Optional[List[SkillSection]] = None
    projects: Optional[List[ProjectSection]] = None   # ← new field

    awards: Optional[List[AwardSection]] = None       # ← existing example

The optional List[<SectionModel>] pattern allows the parser to handle résumés that lack the section without validation errors.

Step 2: Create a Jinja Extraction Prompt

Create a new template file in the templates folder (referenced by prompts/template_manager.py) that instructs the LLM how to structure the extracted data.

{# templates/projects.jinja #}

You are an expert résumé parser. Extract each project from the following résumé text and output a JSON array where each element contains:
- name
- description (optional)
- url (optional)
- startDate (YYYY-MM-DD, optional)
- endDate (YYYY-MM-DD, optional)

Return ONLY valid JSON. If no projects are present, return an empty array.

Register the template in TemplateManager.TEMPLATES so template_manager.render_template("projects", ...) can locate it.

Step 3: Implement the Section Extractor in pdf.py

Add a method to the PDFExtractor class that binds the template to the model. Follow the existing pattern used for work experience or education sections.


# pdf.py

from models import ProjectSection
from typing import Optional, Dict

class PDFExtractor:
    def extract_projects_section(self, resume_text: str) -> Optional[Dict]:
        """Extract the Projects section using the new template."""
        prompt = self.template_manager.render_template(
            "projects", 
            text_content=resume_text
        )
        return self._call_llm_for_section(
            "projects", 
            resume_text, 
            prompt, 
            ProjectSection
        )
    
    def _call_llm_for_section(self, section_name: str, text: str, 
                              prompt: str, model_class):
        # Existing generic LLM caller implementation

        ...

This method reuses _call_llm_for_section to handle the LLM API call, JSON parsing, and Pydantic validation.

Step 4: Assemble the Complete JSONResume

Inside PDFExtractor.extract_all_sections (or the equivalent aggregation method), collect the new section and attach it to the root model before returning.


# pdf.py – inside the extraction orchestration method

projects = self.extract_projects_section(resume_text)
resume = JSONResume(
    basics=basics_data,
    work=work_data,
    education=edu_data,
    skills=skills_data,
    projects=projects,   # ← attach new section

    awards=awards_data
)

If the extraction returns None or an empty list, the optional field in JSONResume handles it gracefully.

Step 5: Transform to Text and CSV

Update transform.py to render the new section into the text representation used by the evaluator and CSV builder. The existing code already iterates over optional attributes using hasattr checks.


# transform.py – inside convert_json_resume_to_text (around lines 652-660)

lines = []

if resume_data and hasattr(resume_data, "projects") and resume_data.projects:
    lines.append("\n### Projects")

    for i, project in enumerate(resume_data.projects, 1):
        lines.append(f"{i}. **{project.name}** – {project.description or ''}")
        if project.url:
            lines.append(f"   URL: {project.url}")
        if project.startDate or project.endDate:
            lines.append(f"   Dates: {project.startDate or ''} – {project.endDate or ''}")

return "\n".join(lines)

For CSV output, extend the row-building logic to include columns for the new fields (e.g., project_names, project_count).

Step 6: Configure Scoring Weights (Optional)

The score.py module automatically includes all rendered text in the evaluation prompt via convert_json_resume_to_text. If you need bespoke weighting—such as prioritizing project descriptions—prepend a formatted header before the standard text conversion.


# score.py – inside resume preparation logic (around lines 170-176)

resume_text = convert_json_resume_to_text(resume_data)

if resume_data.projects:
    highlight = "\n\n---\n**Projects Highlight**\n" + "\n".join(
        f"- {p.name}: {p.description or ''}" for p in resume_data.projects
    )
    resume_text = highlight + "\n\n" + resume_text

This ensures the LLM evaluator sees the projects first without modifying the underlying Evaluator class.

Summary

  • Data Definition – Add List[SectionModel] fields to JSONResume in models.py with strict Pydantic typing.
  • Prompt Engineering – Create specialized .jinja templates and register them in TemplateManager to guide LLM extraction.
  • Extraction Logic – Implement extract_<section>_section methods in pdf.py using the reusable _call_llm_for_section helper.
  • Pipeline Integration – Merge new sections into the root model during assembly and handle them in transform.py for text/CSV output.
  • Scoring Compatibility – The existing scoring pipeline ingests the text representation automatically; optional custom weighting can be added in score.py.

Frequently Asked Questions

How do I add a section that is not part of the standard JSON Resume schema?

Create a custom Pydantic model in models.py (e.g., CertificationSection) and add it as an optional field to JSONResume. The pipeline treats custom fields identically to standard ones—just define the template, extractor method, and transformation logic following the six-step pattern above.

Will adding new resume sections break existing CSV exports?

No. The CSV builder in transform.py iterates over known attributes, so new fields only appear if you explicitly extend the row-building logic. Existing exports continue to work unchanged because optional fields default to None and are skipped during text conversion.

How does the LLM know the correct format for a new section?

The Jinja prompt template (e.g., projects.jinja) provides explicit instructions and requested JSON keys. TemplateManager.render_template injects the raw résumé text into this template, and _call_llm_for_section validates the LLM output against your Pydantic model, retrying if the structure is invalid.

Can I assign higher scoring weight to specific sections like Awards?

Yes. While the default convert_json_resume_to_text flattens all sections equally, you can prepend emphasized text in score.py before calling _evaluate_resume. Insert priority sections at the top of the prompt or duplicate critical fields with special headers to influence the LLM evaluator's attention.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →