Complete Data Flow from PDF Input to Evaluation Output in Hiring Agent
Hiring Agent transforms a raw résumé PDF into a structured evaluation score through eight deterministic stages: PDF extraction, markdown conversion, section-wise LLM parsing, JSONResume assembly, optional GitHub enrichment, text transformation, LLM-based evaluation, and final scoring presentation.
The interviewstreet/hiring-agent repository implements a fully automated pipeline that converts unstructured résumé documents into quantified hiring assessments. Understanding the complete data flow from PDF input to evaluation output reveals how modular Python components and Large Language Model (LLM) integrations work together to produce standardized candidate scores.
Step 1: PDF Ingestion and Markdown Extraction
CLI Entry Point in score.py
The pipeline begins in score.py where the main(pdf_path) function serves as the command-line interface. When invoked via python score.py résumé.pdf, the function creates a cache identifier, checks for existing cached JSON when DEVELOPMENT_MODE=True, and instantiates a PDFHandler to begin extraction.
Low-Level Text Extraction with PyMuPDF
Inside pdf.py, the PDFHandler.extract_text_from_pdf method opens the PDF using PyMuPDF and delegates markdown conversion to to_markdown defined in pymupdf_rag.py. This function iterates through each page, detects structural elements including headings, tables, and images, and returns a Markdown-formatted string that preserves the document's semantic layout.
Step 2: Section-Wise Structured Parsing
Template-Driven LLM Calls in pdf.py
The extracted markdown feeds into a sophisticated parsing system within pdf.py. The handler renders Jinja templates stored under prompts/templates/—including basics.jinja, work.jinja, education.jinja, skills.jinja, projects.jinja, and awards.jinja—to create section-specific prompts. Each request combines the rendered template with a system message from system_message.jinja and sends it to the active LLM provider via self.provider.chat.
Data Transformation and JSONResume Assembly
The LLM response undergoes cleaning and JSON decoding before transform_parsed_data from transform.py post-processes the output to match the JSONResume Pydantic schema. The _extract_all_sections_separately method orchestrates this process across all six core sections, assembling a complete JSONResume instance that structures the candidate's information into typed fields.
Step 3: GitHub Profile Enrichment
After constructing the base resume object, score.py examines resume_data.basics.profiles for entries where the network field equals "Github". When found, fetch_and_display_github_info from github.py retrieves the public user profile and repository list. The system then applies the github_project_selection.jinja template to prompt the LLM to identify up to seven representative projects, enriching the candidate profile with verifiable open-source contributions.
Step 4: Text Conversion for Evaluation
Before final scoring, structured data converts back to plain text using helper functions in transform.py. The convert_json_resume_to_text, convert_github_data_to_text, and convert_blog_data_to_text functions flatten the JSONResume object and GitHub data into concatenated text blocks. These strings form the contextual foundation for the evaluation prompts.
Step 5: Final Evaluation and Scoring
LLM-Based Assessment in evaluator.py
The ResumeEvaluator.evaluate_resume method in evaluator.py loads two critical templates: resume_evaluation_criteria.jinja (containing the detailed scoring rubric) and resume_evaluation_system_message.jinja (providing system-level instructions). The method sends a structured output request to the LLM provider with format=EvaluationData.model_json_schema(), ensuring the response conforms to the expected schema.
Score Computation and Output
The returned JSON payload parses into an EvaluationData Pydantic model. Back in score.py, print_evaluation_results calculates total scores, category breakdowns, bonuses, and deductions, presenting a human-readable report. When running in DEVELOPMENT_MODE, the system appends results to resume_evaluations.csv for batch analysis and tracking.
Complete Code Examples
Run the full pipeline from the command line:
python score.py path/to/resume.pdf
Use the library programmatically to access the evaluation data:
from score import main
evaluation = main("resume.pdf")
# evaluation is an instance of models.EvaluationData
print(evaluation.scores.open_source.score)
Extract the intermediate JSON resume without invoking the evaluator:
from pdf import PDFHandler
handler = PDFHandler()
json_resume = handler.extract_json_from_pdf("resume.pdf")
print(json_resume.model_dump())
Fetch GitHub enrichment data independently:
from github import fetch_and_display_github_info
github_info = fetch_and_display_github_info("https://github.com/username")
print(github_info["profile"]["login"])
Summary
- Entry Point: The
main(pdf_path)function inscore.pyorchestrates the entire pipeline with optional caching viaDEVELOPMENT_MODE. - PDF Processing:
pymupdf_rag.pyconverts PDF bytes to Markdown, preserving document structure through PyMuPDF. - Structured Extraction:
pdf.pyuses section-specific Jinja templates to prompt LLMs, withtransform.pyensuring schema compliance viatransform_parsed_data. - External Enrichment:
github.pyconditionally augments profiles usingfetch_and_display_github_infoand project selection prompts. - Evaluation:
evaluator.pygenerates standardized scores using structured output requests toEvaluationDatamodels. - Output: Final results display through
print_evaluation_resultsinscore.py, with optional CSV export toresume_evaluations.csv.
Frequently Asked Questions
How does Hiring Agent handle different PDF formats?
Hiring Agent uses PyMuPDF in pymupdf_rag.py to normalize diverse PDF layouts into Markdown. The to_markdown function detects headings, tables, and images across pages, creating a consistent text representation that subsequent LLM prompts can process regardless of the original document's formatting.
What LLM providers does Hiring Agent support?
The codebase abstracts provider selection through self.provider.chat interfaces in pdf.py and evaluator.py. The system supports Ollama and Gemini configurations, with provider mapping and API keys managed centrally in prompt.py and config.py.
Can I run the pipeline without the GitHub enrichment step?
Yes. The GitHub enrichment in score.py only executes when the candidate's basics.profiles contains a GitHub URL. If no such profile exists or if the network field does not match "Github", the pipeline proceeds directly to evaluation using only the parsed resume data.
Where is the evaluation data cached?
When DEVELOPMENT_MODE=True, score.py caches intermediate JSON results and writes final evaluations to resume_evaluations.csv. The cache name derives from the input PDF path, allowing rapid re-running of the pipeline without repeating LLM calls during development and testing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →