How the Hiring Agent Handles Missing or Incomplete Resume Sections
The hiring agent processes each resume section independently through optional LLM extraction calls, validates that at least one core section exists before proceeding, and substitutes empty strings or zeroes for missing data during CSV generation to ensure robust handling of incomplete candidate resumes.
The interviewstreet/hiring-agent repository implements a fault-tolerant pipeline designed to extract structured data from PDF resumes even when candidates submit incomplete documents. Understanding how the system handles missing or incomplete resume sections is essential for developers extending the evaluation logic or debugging extraction failures. The architecture deliberately isolates section extraction, implements validation gates, and provides safe defaults to prevent downstream errors.
Optional Section Extraction via Independent LLM Calls
In main/pdf.py, the PDFHandler class uses TemplateManager from main/prompts/template_manager.py to render section-specific prompts. Each extraction method—such as extract_work_section—returns Optional[Dict], allowing the pipeline to continue when a specific section cannot be parsed (lines 36‑50).
def extract_work_section(self, resume_text: str) -> Optional[Dict]:
prompt = self.template_manager.render_template("work", text_content=resume_text)
# ...
return self._call_llm_for_section("work", resume_text, prompt, WorkSection)
The extract_json_from_text method orchestrates these individual calls, building a JSONResume object only from sections that succeed. If the LLM fails to produce valid JSON for a section, the method returns None rather than raising an exception. This ensures that the absence of a "Projects" or "Awards" section does not halt the entire extraction process.
# Example: extracting a PDF that lacks a "projects" section
pdf = PDFHandler()
resume = pdf.extract_json_from_pdf("candidate.pdf") # returns JSONResume
# `resume.projects` will be None, but other sections may be populated
Core Section Validation Before Processing
Before proceeding to caching, GitHub enrichment, or evaluation, the pipeline validates that the resume contains at least some usable data. In main/score.py, the is_valid_resume_data function checks the JSONResume object for the presence of any core section (lines 191‑203).
core_sections = [
resume_data.basics,
resume_data.work,
resume_data.education,
resume_data.skills,
resume_data.projects,
]
return any(section is not None for section in core_sections)
If none of these core sections are present, the pipeline aborts early with a clear warning log. This prevents downstream errors when processing completely empty or corrupted PDFs while still allowing partial resumes to proceed.
Safe Defaults for CSV Generation
The transform_evaluation_response function in main/transform.py handles the conversion of parsed resume data into flat CSV rows. For every optional field, it checks for existence and supplies empty strings or zeroes when sections are missing (lines 15‑95, 96‑115, 124‑136).
if linkedin_profile:
csv_row["linkedin_url"] = linkedin_profile.url
else:
csv_row["linkedin_url"] = ""
This defensive pattern applies across work experience, education, skills, projects, and GitHub-related metrics. Missing sections never cause KeyError exceptions during batch processing.
# Example: converting to CSV, safely handling missing sections
csv_row = transform_evaluation_response(
file_name="candidate.pdf",
resume_data=resume,
github_data={},
evaluation=None,
)
# Missing sections appear as empty strings / zeroes in `csv_row`
Cache Validation and Error Recovery
The system implements safeguards against stale or invalid cached data. When loading cached JSON files, main/score.py re-runs the is_valid_resume_data check (lines 30‑38). If the cache contains only empty sections—indicating a previous failed extraction—it is discarded and the PDF is re-processed automatically. This ensures that temporary extraction failures do not persist across runs.
Summary
- Independent section extraction: Each resume section is processed via separate LLM calls in
PDFHandler, returningNonefor missing sections without stopping the pipeline. - Validation gate: The
is_valid_resume_datafunction inscore.pyensures at least one core section exists before caching or evaluation. - Defensive CSV generation:
transform.pysubstitutes empty strings and zeroes for missing data, preventing downstream errors during output generation. - Cache invalidation: Invalid cached results are automatically detected and re-processed to maintain data integrity.
Frequently Asked Questions
What happens if a resume is completely empty?
If none of the core sections (basics, work, education, skills, projects) are present, the is_valid_resume_data check in main/score.py returns False and the pipeline aborts early with a warning log. This prevents processing of completely invalid resumes while allowing partial data to proceed.
Does the system raise exceptions for missing sections?
No. Individual section extraction methods in main/pdf.py return None when sections cannot be parsed. Exceptions are only raised if the entire resume is invalid (no core sections present) or during catastrophic LLM failures. Warning logs indicate which specific sections failed to parse.
How are missing LinkedIn URLs or GitHub profiles handled in the final CSV?
The transform_evaluation_response function in main/transform.py checks for attribute existence before accessing nested fields. Missing values are replaced with empty strings for URLs and zeroes for numeric metrics, ensuring the CSV output remains structurally consistent even with sparse input data.
Can the pipeline resume if a cached extraction is incomplete?
Yes. When loading cached JSON files, the system re-validates the data using is_valid_resume_data. If the cache contains only empty sections (indicating a previous failed extraction), it is discarded and the PDF is re-processed automatically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →