How to Add Custom Resume Sections to the Hiring-Agent Resume Parser

To add custom resume sections in the interviewstreet/hiring-agent repository, you must extend the three-layer architecture by creating a Pydantic model in models.py, adding a Jinja template in prompts/templates/, and registering the new extraction logic in pdf.py and template_manager.py.

The Hiring-Agent parser extracts structured data from PDF resumes by invoking LLM prompts for each predefined section. While the default pipeline handles Basics, Work, Education, Skills, Projects, and Awards, the modular design allows you to add organization-specific sections like Certificates, Publications, or Volunteer Experience. This requires synchronized updates to the data models, prompt templates, and extraction orchestration.

Understanding the Three-Layer Architecture

The extraction pipeline relies on three coordinated components that must all be updated when adding custom resume sections:

  1. Data Layer (main/models.py) – Pydantic models define the JSON schema for each section and the final JSONResume container
  2. Template Layer (main/prompts/) – Jinja templates instruct the LLM how to extract and format data for specific sections
  3. Orchestration Layer (main/pdf.py) – The PDFHandler class iterates through sections, renders prompts, and assembles the final object

Each default section appears in the hard-coded list within PDFHandler._extract_all_sections_separately (lines 271-272 in main/pdf.py), has a corresponding template loaded by TemplateManager, and maps to a typed attribute in the JSONResume class.

Step 1: Define the Pydantic Model

Create a section-specific model in main/models.py that mirrors the structure of existing sections like ProjectsSection or AwardsSection.


# main/models.py

from typing import List, Optional
from pydantic import BaseModel

class Certificate(BaseModel):
    name: str
    date: Optional[str] = None
    issuer: Optional[str] = None
    url: Optional[str] = None

class CertificateSection(BaseModel):
    """Certificates section containing a list of professional certifications."""
    certificates: Optional[List[Certificate]] = None

The CertificateSection class enables PDFHandler._call_llm_for_section to generate a JSON schema via return_model.model_json_schema(), which is passed to the LLM to enforce structured output.

Step 2: Extend the JSONResume Container

Add the new section as an optional attribute to the JSONResume class so the final assembled object can hold the extracted data.


# main/models.py – inside the JSONResume class

class JSONResume(BaseModel):
    basics: Optional[Basics] = None
    work: Optional[List[Work]] = None
    education: Optional[List[Education]] = None
    skills: Optional[List[Skill]] = None
    projects: Optional[List[Project]] = None
    awards: Optional[List[Award]] = None
    certificates: Optional[List[Certificate]] = None  # Add this line

This ensures that when _extract_all_sections_separately merges section dictionaries into complete_resume, the new field is properly typed and accessible.

Step 3: Create the Jinja Template

Create a new template file in main/prompts/templates/ that instructs the LLM how to extract your custom section.

{# main/prompts/templates/certificates.jinja #}

You are an expert resume parser.
Extract every certificate entry from the following raw resume text.
Return a JSON object matching the following schema:
{
  "certificates": [
    {
      "name": "<certificate name>",
      "date": "<completion date>",
      "issuer": "<issuing organization>",
      "url": "<optional link>"
    }
  ]
}

Resume text:
{{ text_content }}

TemplateManager loads these files dynamically, so the filename must match the key you will register in the next step.

Step 4: Register with TemplateManager

Update the _load_templates method in main/prompts/template_manager.py to include your new template file.


# main/prompts/template_manager.py

def _load_templates(self):
    template_files = {
        "basics": "basics.jinja",
        "work": "work.jinja",
        "education": "education.jinja",
        "skills": "skills.jinja",
        "projects": "projects.jinja",
        "awards": "awards.jinja",
        "certificates": "certificates.jinja",  # Add this entry

    }
    # ... rest of loading logic

Only sections listed in this dictionary are available for rendering via self.template_manager.render_template("certificates", ...).

Step 5: Add the Extraction Method

Implement a dedicated extraction method in main/pdf.py within the PDFHandler class, following the pattern of existing methods like extract_work_section.


# main/pdf.py – inside class PDFHandler

def extract_certificates_section(self, resume_text: str) -> Optional[Dict]:
    prompt = self.template_manager.render_template(
        "certificates", text_content=resume_text
    )
    if not prompt:
        logger.error("❌ Failed to render certificates template")
        return None
    return self._call_llm_for_section(
        "certificates", resume_text, prompt, CertificateSection
    )

This method renders the prompt, validates the LLM response against your CertificateSection model, and returns the parsed dictionary.

Step 6: Wire into the Extraction Loop

Modify PDFHandler._extract_all_sections_separately to include your new section in both the iteration list and the dispatcher dictionary.


# main/pdf.py – inside _extract_all_sections_separately

def _extract_all_sections_separately(self, resume_text: str) -> JSONResume:
    # Update the sections list (around line 271)

    sections = ["basics", "work", "education", "skills", "projects", "awards", "certificates"]
    
    # Update the dispatcher mapping in _extract_section_data or similar

    section_extractors = {
        "basics": self.extract_basics_section,
        "work": self.extract_work_section,
        "education": self.extract_education_section,
        "skills": self.extract_skills_section,
        "projects": self.extract_projects_section,
        "awards": self.extract_awards_section,
        "certificates": self.extract_certificates_section,  # Add this

    }
    
    # ... extraction loop logic

The parser will now automatically invoke the LLM for your custom section during the extraction workflow.

Complete Working Example: Adding a Volunteer Section

Here is a minimal implementation adding a Volunteer section using the same six-step pattern:


# 1. main/models.py

class Volunteer(BaseModel):
    organization: str
    position: Optional[str] = None
    startDate: Optional[str] = None
    endDate: Optional[str] = None
    summary: Optional[str] = None
    highlights: Optional[List[str]] = None

class VolunteerSection(BaseModel):
    volunteer: Optional[List[Volunteer]] = None

# Add to JSONResume:

volunteer: Optional[List[Volunteer]] = None

# 2. main/prompts/templates/volunteer.jinja

"""
Extract volunteer experiences from the resume text.
Return JSON matching: {"volunteer": [{"organization": "...", "position": "..."}]}
Text: {{ text_content }}
"""

# 3. main/prompts/template_manager.py

template_files["volunteer"] = "volunteer.jinja"

# 4. main/pdf.py

def extract_volunteer_section(self, resume_text: str) -> Optional[Dict]:
    prompt = self.template_manager.render_template("volunteer", text_content=resume_text)
    return self._call_llm_for_section("volunteer", resume_text, prompt, VolunteerSection)

# 5. Register in section_extractors and sections list

section_extractors["volunteer"] = self.extract_volunteer_section
sections.append("volunteer")

Running python main/score.py resume.pdf will now include a volunteer array in the output cache file.

Key Files and Their Roles

  • main/models.py – Defines JSONResume and section-specific Pydantic models that generate LLM JSON schemas
  • main/pdf.py – Contains PDFHandler class with _extract_all_sections_separately (lines 271-272) and section extraction methods
  • main/prompts/template_manager.py – Loads Jinja templates via _load_templates (lines 35-44)
  • main/prompts/templates/ – Directory containing .jinja files for each section's LLM prompt
  • main/score.py – Entry point that orchestrates extraction and caches results to cache/resumecache_<name>.json

Summary

  • Data models must be added to main/models.py to define the schema and type validation for custom sections
  • Jinja templates must be created in main/prompts/templates/ and registered in TemplateManager._load_templates to generate LLM prompts
  • Extraction methods in main/pdf.py follow the pattern extract_{section}_section and use _call_llm_for_section with your Pydantic model
  • Orchestration updates require adding the section name to the sections list and section_extractors dictionary in PDFHandler
  • Validation is handled automatically by the Pydantic model when _call_llm_for_section parses the LLM response

Frequently Asked Questions

Can I add multiple custom sections at once?

Yes. Simply repeat the six-step process for each new section. Each requires its own Pydantic model, Jinja template, and extraction method registration. Ensure each section name is unique in the sections list and section_extractors dictionary to avoid naming collisions.

Do I need to modify the scoring logic to use custom sections?

Not necessarily. The JSONResume model will include your new fields, and downstream components that ignore unknown fields will continue functioning. However, if you want the scoring algorithm in main/score.py to evaluate the new section (e.g., weighting certificates), you must explicitly add that logic to the evaluation functions.

What if the LLM returns malformed data for my custom section?

The _call_llm_for_section method handles validation by passing your Pydantic model (e.g., CertificateSection) to enforce the JSON schema. If the LLM returns invalid JSON or missing required fields, Pydantic will raise a validation error, and the method will return None for that section, preventing the malformed data from corrupting the final resume object.

Can I use this approach for sections with nested complex structures?

Absolutely. The Pydantic models support arbitrary nesting. Define nested models for complex entities (like Certificate containing Issuer details), and the model_json_schema() method will automatically generate the appropriate JSON schema for the LLM prompt. Ensure your Jinja template explicitly describes the nested structure to guide the LLM extraction accurately.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →