How to Add Custom Resume Sections to the Hiring-Agent Resume Parser
To add custom resume sections in the interviewstreet/hiring-agent repository, you must extend the three-layer architecture by creating a Pydantic model in models.py, adding a Jinja template in prompts/templates/, and registering the new extraction logic in pdf.py and template_manager.py.
The Hiring-Agent parser extracts structured data from PDF resumes by invoking LLM prompts for each predefined section. While the default pipeline handles Basics, Work, Education, Skills, Projects, and Awards, the modular design allows you to add organization-specific sections like Certificates, Publications, or Volunteer Experience. This requires synchronized updates to the data models, prompt templates, and extraction orchestration.
Understanding the Three-Layer Architecture
The extraction pipeline relies on three coordinated components that must all be updated when adding custom resume sections:
- Data Layer (
main/models.py) – Pydantic models define the JSON schema for each section and the finalJSONResumecontainer - Template Layer (
main/prompts/) – Jinja templates instruct the LLM how to extract and format data for specific sections - Orchestration Layer (
main/pdf.py) – ThePDFHandlerclass iterates through sections, renders prompts, and assembles the final object
Each default section appears in the hard-coded list within PDFHandler._extract_all_sections_separately (lines 271-272 in main/pdf.py), has a corresponding template loaded by TemplateManager, and maps to a typed attribute in the JSONResume class.
Step 1: Define the Pydantic Model
Create a section-specific model in main/models.py that mirrors the structure of existing sections like ProjectsSection or AwardsSection.
# main/models.py
from typing import List, Optional
from pydantic import BaseModel
class Certificate(BaseModel):
name: str
date: Optional[str] = None
issuer: Optional[str] = None
url: Optional[str] = None
class CertificateSection(BaseModel):
"""Certificates section containing a list of professional certifications."""
certificates: Optional[List[Certificate]] = None
The CertificateSection class enables PDFHandler._call_llm_for_section to generate a JSON schema via return_model.model_json_schema(), which is passed to the LLM to enforce structured output.
Step 2: Extend the JSONResume Container
Add the new section as an optional attribute to the JSONResume class so the final assembled object can hold the extracted data.
# main/models.py – inside the JSONResume class
class JSONResume(BaseModel):
basics: Optional[Basics] = None
work: Optional[List[Work]] = None
education: Optional[List[Education]] = None
skills: Optional[List[Skill]] = None
projects: Optional[List[Project]] = None
awards: Optional[List[Award]] = None
certificates: Optional[List[Certificate]] = None # Add this line
This ensures that when _extract_all_sections_separately merges section dictionaries into complete_resume, the new field is properly typed and accessible.
Step 3: Create the Jinja Template
Create a new template file in main/prompts/templates/ that instructs the LLM how to extract your custom section.
{# main/prompts/templates/certificates.jinja #}
You are an expert resume parser.
Extract every certificate entry from the following raw resume text.
Return a JSON object matching the following schema:
{
"certificates": [
{
"name": "<certificate name>",
"date": "<completion date>",
"issuer": "<issuing organization>",
"url": "<optional link>"
}
]
}
Resume text:
{{ text_content }}
TemplateManager loads these files dynamically, so the filename must match the key you will register in the next step.
Step 4: Register with TemplateManager
Update the _load_templates method in main/prompts/template_manager.py to include your new template file.
# main/prompts/template_manager.py
def _load_templates(self):
template_files = {
"basics": "basics.jinja",
"work": "work.jinja",
"education": "education.jinja",
"skills": "skills.jinja",
"projects": "projects.jinja",
"awards": "awards.jinja",
"certificates": "certificates.jinja", # Add this entry
}
# ... rest of loading logic
Only sections listed in this dictionary are available for rendering via self.template_manager.render_template("certificates", ...).
Step 5: Add the Extraction Method
Implement a dedicated extraction method in main/pdf.py within the PDFHandler class, following the pattern of existing methods like extract_work_section.
# main/pdf.py – inside class PDFHandler
def extract_certificates_section(self, resume_text: str) -> Optional[Dict]:
prompt = self.template_manager.render_template(
"certificates", text_content=resume_text
)
if not prompt:
logger.error("❌ Failed to render certificates template")
return None
return self._call_llm_for_section(
"certificates", resume_text, prompt, CertificateSection
)
This method renders the prompt, validates the LLM response against your CertificateSection model, and returns the parsed dictionary.
Step 6: Wire into the Extraction Loop
Modify PDFHandler._extract_all_sections_separately to include your new section in both the iteration list and the dispatcher dictionary.
# main/pdf.py – inside _extract_all_sections_separately
def _extract_all_sections_separately(self, resume_text: str) -> JSONResume:
# Update the sections list (around line 271)
sections = ["basics", "work", "education", "skills", "projects", "awards", "certificates"]
# Update the dispatcher mapping in _extract_section_data or similar
section_extractors = {
"basics": self.extract_basics_section,
"work": self.extract_work_section,
"education": self.extract_education_section,
"skills": self.extract_skills_section,
"projects": self.extract_projects_section,
"awards": self.extract_awards_section,
"certificates": self.extract_certificates_section, # Add this
}
# ... extraction loop logic
The parser will now automatically invoke the LLM for your custom section during the extraction workflow.
Complete Working Example: Adding a Volunteer Section
Here is a minimal implementation adding a Volunteer section using the same six-step pattern:
# 1. main/models.py
class Volunteer(BaseModel):
organization: str
position: Optional[str] = None
startDate: Optional[str] = None
endDate: Optional[str] = None
summary: Optional[str] = None
highlights: Optional[List[str]] = None
class VolunteerSection(BaseModel):
volunteer: Optional[List[Volunteer]] = None
# Add to JSONResume:
volunteer: Optional[List[Volunteer]] = None
# 2. main/prompts/templates/volunteer.jinja
"""
Extract volunteer experiences from the resume text.
Return JSON matching: {"volunteer": [{"organization": "...", "position": "..."}]}
Text: {{ text_content }}
"""
# 3. main/prompts/template_manager.py
template_files["volunteer"] = "volunteer.jinja"
# 4. main/pdf.py
def extract_volunteer_section(self, resume_text: str) -> Optional[Dict]:
prompt = self.template_manager.render_template("volunteer", text_content=resume_text)
return self._call_llm_for_section("volunteer", resume_text, prompt, VolunteerSection)
# 5. Register in section_extractors and sections list
section_extractors["volunteer"] = self.extract_volunteer_section
sections.append("volunteer")
Running python main/score.py resume.pdf will now include a volunteer array in the output cache file.
Key Files and Their Roles
main/models.py– DefinesJSONResumeand section-specific Pydantic models that generate LLM JSON schemasmain/pdf.py– ContainsPDFHandlerclass with_extract_all_sections_separately(lines 271-272) and section extraction methodsmain/prompts/template_manager.py– Loads Jinja templates via_load_templates(lines 35-44)main/prompts/templates/– Directory containing.jinjafiles for each section's LLM promptmain/score.py– Entry point that orchestrates extraction and caches results tocache/resumecache_<name>.json
Summary
- Data models must be added to
main/models.pyto define the schema and type validation for custom sections - Jinja templates must be created in
main/prompts/templates/and registered inTemplateManager._load_templatesto generate LLM prompts - Extraction methods in
main/pdf.pyfollow the patternextract_{section}_sectionand use_call_llm_for_sectionwith your Pydantic model - Orchestration updates require adding the section name to the
sectionslist andsection_extractorsdictionary inPDFHandler - Validation is handled automatically by the Pydantic model when
_call_llm_for_sectionparses the LLM response
Frequently Asked Questions
Can I add multiple custom sections at once?
Yes. Simply repeat the six-step process for each new section. Each requires its own Pydantic model, Jinja template, and extraction method registration. Ensure each section name is unique in the sections list and section_extractors dictionary to avoid naming collisions.
Do I need to modify the scoring logic to use custom sections?
Not necessarily. The JSONResume model will include your new fields, and downstream components that ignore unknown fields will continue functioning. However, if you want the scoring algorithm in main/score.py to evaluate the new section (e.g., weighting certificates), you must explicitly add that logic to the evaluation functions.
What if the LLM returns malformed data for my custom section?
The _call_llm_for_section method handles validation by passing your Pydantic model (e.g., CertificateSection) to enforce the JSON schema. If the LLM returns invalid JSON or missing required fields, Pydantic will raise a validation error, and the method will return None for that section, preventing the malformed data from corrupting the final resume object.
Can I use this approach for sections with nested complex structures?
Absolutely. The Pydantic models support arbitrary nesting. Define nested models for complex entities (like Certificate containing Issuer details), and the model_json_schema() method will automatically generate the appropriate JSON schema for the LLM prompt. Ensure your Jinja template explicitly describes the nested structure to guide the LLM extraction accurately.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →