How Resume Validation Ensures Extracted Data Has Core Content in InterviewStreet’s Hiring-Agent
The hiring-agent validates résumé extraction by requiring at least one of five core sections—basics, work, education, skills, or projects—to be present before proceeding with caching or evaluation.
The interviewstreet/hiring-agent repository implements a robust safeguard against empty or malformed PDF extractions. By enforcing strict validation rules in score.py, the system guarantees that downstream AI evaluation only receives résumés containing meaningful structured data.
The Core Validation Logic in score.py
The validation mechanism centers on a single helper function designed to inspect the JSONResume object returned by the PDF parser.
The is_valid_resume_data Function
Located at lines 91–102 in score.py, is_valid_resume_data performs a two-tier check on extracted data:
# score.py – validation helper (lines 91-102)
def is_valid_resume_data(resume_data: JSONResume) -> bool:
"""Check if the resume data has at least some extracted core content."""
if not resume_data:
return False
core_sections = [
resume_data.basics,
resume_data.work,
resume_data.education,
resume_data.skills,
resume_data.projects,
]
return any(section is not None for section in core_sections)
The function first defends against null inputs by checking if resume_data is None. It then constructs a core_sections list containing the five essential attributes of the JSONResume model: basics, work, education, skills, and projects. Using Python’s any() function, it returns True if at least one of these sections contains extracted data (i.e., is not None).
Where Resume Validation Protects the Pipeline
According to the hiring-agent source code, this validator is invoked at two critical junctions to prevent empty data from corrupting the evaluation workflow.
Cache Integrity Checks
When loading a previously cached résumé, the system validates the restored data before trusting it. As implemented in score.py lines 30–34:
loaded_resume = JSONResume(**cached_data)
if not is_valid_resume_data(loaded_resume):
raise ValueError("Cached resume data contains no core content")
If the cached JSON lacks core sections, the function raises a ValueError, triggering a cache miss and forcing a fresh PDF extraction.
Post-Extraction Validation
After PDF processing completes, newly extracted data undergoes the same validation before being written to disk. Lines 58–66 in score.py implement this safeguard:
if is_valid_resume_data(resume_data):
Path(cache_filename).write_text(json.dumps(resume_data.dict(), indent=2))
else:
logger.warning("Newly extracted resume data is empty/invalid.")
Only valid résumés are cached; invalid extractions trigger a warning and are discarded, ensuring the cache never stores meaningless data.
Practical Implementation Examples
Developers working with the hiring-agent can leverage is_valid_resume_data directly in custom scripts to ensure extracted data has core content before processing.
Validating Parsed Resumes Manually
from main.score import is_valid_resume_data
from main.models import JSONResume
# Assume `parsed` is a JSONResume returned by PDFHandler.extract_json_from_pdf(...)
if is_valid_resume_data(parsed):
print("Resume contains core content – safe to evaluate.")
else:
print("Resume is empty; abort or request a new PDF.")
Integrating Validation in Custom Workflows
def process_resume(pdf_path: str):
from main.pdf import PDFHandler
from main.score import is_valid_resume_data
handler = PDFHandler()
resume = handler.extract_json_from_pdf(pdf_path)
if not is_valid_resume_data(resume):
raise RuntimeError("Extracted résumé lacks core sections")
# Continue with evaluation, scoring, etc.
# ...
Key Files in the Validation Architecture
The resume validation system spans several modules:
score.py– Definesis_valid_resume_dataand integrates validation with caching and evaluation logic.models.py– Contains the JSONResume Pydantic model that defines the structure of the five core sections.pdf.py– HousesPDFHandler.extract_json_from_pdf, the source of raw extracted data fed into the validator.transform.py– Utilities for converting validatedJSONResumeobjects into CSV rows for downstream analysis.
Summary
- Validation requires at least one core section: The
is_valid_resume_datafunction returnsTrueonly ifbasics,work,education,skills, orprojectsis present in the extracted data. - Null safety is mandatory: The validator rejects
Noneinputs entirely before checking section contents. - Cache protection is bidirectional: Validation occurs both when reading cached data (triggering re-extraction if invalid) and when writing new data (preventing invalid cache entries).
- Downstream evaluation is protected: By raising
ValueErroror warnings early, the system guarantees that_evaluate_resumereceives only meaningful résumé objects.
Frequently Asked Questions
What constitutes "core content" in resume validation?
Core content refers to any non-null data within the five essential sections defined in the JSONResume model: basics (contact information), work (employment history), education (academic background), skills (competencies), and projects (portfolio work). The validation passes if at least one of these fields contains extracted data.
How does the hiring-agent handle invalid cached resume data?
When loading a cache file in score.py (lines 30–34), the system constructs a JSONResume instance and runs is_valid_resume_data. If validation fails, the code raises a ValueError with the message "Cached resume data contains no core content", causing the system to ignore the cache and re-process the original PDF.
Can the validation logic be customized to require specific sections?
Yes. While the default implementation uses any() to accept at least one section, developers can modify the core_sections list logic in score.py lines 95–100. For example, changing any() to all() would require every section to be present, or you could check for specific combinations like work AND education before returning True.
Which file contains the is_valid_resume_data function?
The function is defined in score.py at lines 91–102. This file also contains the integration points where the validator protects cache operations at lines 30–34 (loading) and lines 58–66 (writing).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →