How to Support New Resume Sections in Hiring-Agent: Adding Projects, Awards, and Custom Fields
Add new resume sections by extending the JSONResume Pydantic model, creating a Jinja prompt template, and wiring the extraction logic in PDFExtractor—no changes to the core scoring engine required.
The interviewstreet/hiring-agent repository parses PDF résumés using LLM-driven prompts and converts them into structured JSON before scoring candidates. To support new resume sections like Projects, Awards, or custom fields, you modify the data models, prompts, and extraction pipeline while reusing the existing transformation and scoring infrastructure.
Understanding the Resume Parsing Pipeline
The Hiring-Agent follows a six-stage data flow that makes adding sections straightforward:
- PDF Ingestion –
PDFHandler.extract_text_from_pdfextracts raw text from the uploaded file. - Section Extraction –
PDFExtractorloads Jinja templates (e.g.,projects.jinja) and calls the LLM via_call_llm_for_sectionto parse specific segments. - Model Assembly – Parsed sections populate a
JSONResumePydantic model instance. - Text Transformation –
transform.convert_json_resume_to_textrenders the model as markdown-style text for evaluation. - CSV Generation – The same transformation logic flattens the model into a spreadsheet row.
- AI Scoring –
score._evaluate_resumeconcatenates the text representation and sends it to theEvaluator.
Because each stage operates on the JSONResume model rather than hard-coded fields, you can introduce new attributes without touching the scoring algorithm.
Step 1: Define the Data Model in models.py
First, declare the new section's structure in models.py by creating a Pydantic class and adding it to the JSONResume root model.
# models.py
from typing import List, Optional
from pydantic import BaseModel
class ProjectSection(BaseModel):
name: str
description: Optional[str] = None
url: Optional[str] = None
startDate: Optional[str] = None
endDate: Optional[str] = None
class JSONResume(BaseModel):
basics: Optional[BasicsSection] = None
work: Optional[List[WorkSection]] = None
education: Optional[List[EducationSection]] = None
skills: Optional[List[SkillSection]] = None
projects: Optional[List[ProjectSection]] = None # ← new field
awards: Optional[List[AwardSection]] = None # ← existing example
The optional List[<SectionModel>] pattern allows the parser to handle résumés that lack the section without validation errors.
Step 2: Create a Jinja Extraction Prompt
Create a new template file in the templates folder (referenced by prompts/template_manager.py) that instructs the LLM how to structure the extracted data.
{# templates/projects.jinja #}
You are an expert résumé parser. Extract each project from the following résumé text and output a JSON array where each element contains:
- name
- description (optional)
- url (optional)
- startDate (YYYY-MM-DD, optional)
- endDate (YYYY-MM-DD, optional)
Return ONLY valid JSON. If no projects are present, return an empty array.
Register the template in TemplateManager.TEMPLATES so template_manager.render_template("projects", ...) can locate it.
Step 3: Implement the Section Extractor in pdf.py
Add a method to the PDFExtractor class that binds the template to the model. Follow the existing pattern used for work experience or education sections.
# pdf.py
from models import ProjectSection
from typing import Optional, Dict
class PDFExtractor:
def extract_projects_section(self, resume_text: str) -> Optional[Dict]:
"""Extract the Projects section using the new template."""
prompt = self.template_manager.render_template(
"projects",
text_content=resume_text
)
return self._call_llm_for_section(
"projects",
resume_text,
prompt,
ProjectSection
)
def _call_llm_for_section(self, section_name: str, text: str,
prompt: str, model_class):
# Existing generic LLM caller implementation
...
This method reuses _call_llm_for_section to handle the LLM API call, JSON parsing, and Pydantic validation.
Step 4: Assemble the Complete JSONResume
Inside PDFExtractor.extract_all_sections (or the equivalent aggregation method), collect the new section and attach it to the root model before returning.
# pdf.py – inside the extraction orchestration method
projects = self.extract_projects_section(resume_text)
resume = JSONResume(
basics=basics_data,
work=work_data,
education=edu_data,
skills=skills_data,
projects=projects, # ← attach new section
awards=awards_data
)
If the extraction returns None or an empty list, the optional field in JSONResume handles it gracefully.
Step 5: Transform to Text and CSV
Update transform.py to render the new section into the text representation used by the evaluator and CSV builder. The existing code already iterates over optional attributes using hasattr checks.
# transform.py – inside convert_json_resume_to_text (around lines 652-660)
lines = []
if resume_data and hasattr(resume_data, "projects") and resume_data.projects:
lines.append("\n### Projects")
for i, project in enumerate(resume_data.projects, 1):
lines.append(f"{i}. **{project.name}** – {project.description or ''}")
if project.url:
lines.append(f" URL: {project.url}")
if project.startDate or project.endDate:
lines.append(f" Dates: {project.startDate or ''} – {project.endDate or ''}")
return "\n".join(lines)
For CSV output, extend the row-building logic to include columns for the new fields (e.g., project_names, project_count).
Step 6: Configure Scoring Weights (Optional)
The score.py module automatically includes all rendered text in the evaluation prompt via convert_json_resume_to_text. If you need bespoke weighting—such as prioritizing project descriptions—prepend a formatted header before the standard text conversion.
# score.py – inside resume preparation logic (around lines 170-176)
resume_text = convert_json_resume_to_text(resume_data)
if resume_data.projects:
highlight = "\n\n---\n**Projects Highlight**\n" + "\n".join(
f"- {p.name}: {p.description or ''}" for p in resume_data.projects
)
resume_text = highlight + "\n\n" + resume_text
This ensures the LLM evaluator sees the projects first without modifying the underlying Evaluator class.
Summary
- Data Definition – Add
List[SectionModel]fields toJSONResumeinmodels.pywith strict Pydantic typing. - Prompt Engineering – Create specialized
.jinjatemplates and register them inTemplateManagerto guide LLM extraction. - Extraction Logic – Implement
extract_<section>_sectionmethods inpdf.pyusing the reusable_call_llm_for_sectionhelper. - Pipeline Integration – Merge new sections into the root model during assembly and handle them in
transform.pyfor text/CSV output. - Scoring Compatibility – The existing scoring pipeline ingests the text representation automatically; optional custom weighting can be added in
score.py.
Frequently Asked Questions
How do I add a section that is not part of the standard JSON Resume schema?
Create a custom Pydantic model in models.py (e.g., CertificationSection) and add it as an optional field to JSONResume. The pipeline treats custom fields identically to standard ones—just define the template, extractor method, and transformation logic following the six-step pattern above.
Will adding new resume sections break existing CSV exports?
No. The CSV builder in transform.py iterates over known attributes, so new fields only appear if you explicitly extend the row-building logic. Existing exports continue to work unchanged because optional fields default to None and are skipped during text conversion.
How does the LLM know the correct format for a new section?
The Jinja prompt template (e.g., projects.jinja) provides explicit instructions and requested JSON keys. TemplateManager.render_template injects the raw résumé text into this template, and _call_llm_for_section validates the LLM output against your Pydantic model, retrying if the structure is invalid.
Can I assign higher scoring weight to specific sections like Awards?
Yes. While the default convert_json_resume_to_text flattens all sections equally, you can prepend emphasized text in score.py before calling _evaluate_resume. Insert priority sections at the top of the prompt or duplicate critical fields with special headers to influence the LLM evaluator's attention.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →