Main Components of the Hiring-Agent Project: Architecture and Pipeline Stages
The hiring-agent project consists of five core pipeline stages—PDF extraction, section parsing, GitHub enrichment, evaluation, and orchestration—supported by LLM utility modules and configuration management that transform résumé PDFs into structured, scored evaluations.
The hiring-agent repository by InterviewStreet is a modular Python application designed to automate technical candidate evaluation through LLM-powered pipeline processing. Understanding the main components of the hiring-agent project reveals a clean architecture where each module handles specific responsibilities, from ingesting raw PDF documents to generating fairness-aware scoring reports.
PDF Extraction and Document Processing
The pipeline initiates in pymupdf_rag.py and pdf.py, which collaborate to convert résumé PDFs into structured Markdown-like text. The pymupdf_rag.py module leverages PyMuPDF to read each page and construct a markdown representation that preserves headings, tables, and links. This content then passes to pdf.py, which orchestrates section-wise LLM calls to parse the document.
The PDFHandler class in pdf.py serves as the primary interface for this stage. It internally invokes pymupdf_rag.to_markdown, then processes each section through the LLM using strict Jinja templates.
Structured Resume Parsing and Validation
Once extracted, raw markdown transforms into JSON-Resume style objects through the parsing layer. This stage relies on three key components:
prompts/templates/*.jinja: Provider-agnostic Jinja templates that define strict prompts for sections including Basics, Work, Education, Skills, Projects, and Awardsprompt.py: Dispatches prompts to the LLM and manages provider-specific formattingmodels.py: Defines Pydantic schemas that validate LLM-returned JSON, ensuring type safety and structural integrity
The interaction between these files ensures that unstructured résumé text converts into validated Python objects before proceeding to enrichment.
GitHub Profile Enrichment
The github.py module augments résumé data with signals from the candidate’s GitHub profile. The GitHubEnricher class detects GitHub usernames within the parsed résumé, then fetches profile and repository data via the GitHub API.
This component classifies each repository and invokes the LLM to select the seven most relevant projects based on minimum author-commit thresholds and diversity of contribution types. The enriched data attaches directly to the résumé object, providing concrete evidence of open-source contributions for subsequent evaluation.
Evaluation and Scoring Engine
Fairness-aware scoring occurs in evaluator.py, which implements the Evaluator class. This module uses Jinja templates embedding fairness constraints and scoring rubrics to compute category-specific scores.
The evaluation produces multiple outputs:
- Category scores:
open_source,self_projects,production,technical_skills - Bonuses and deductions: Adjustments based on specific criteria
- Natural-language evidence: Human-readable reasoning for each score
The evaluate() method returns a Pydantic EvaluationResult object containing structured scores and explanatory text.
Pipeline Orchestration and Output
The score.py module functions as the end-to-end pipeline driver, tying all stages together while managing caching and output generation. When executed from the command line, it triggers the full flow from PDF ingestion to final report.
Key capabilities include:
- Writing human-readable summaries to stdout
- Appending CSV rows to
resume_evaluations.csvwhenDEVELOPMENT_MODE=True - Caching intermediate JSON blobs under
cache/for inspection and debugging
This module centralizes workflow management, ensuring that extraction, parsing, enrichment, and evaluation execute in sequence with proper error handling.
LLM Utilities and Configuration
Supporting infrastructure abstracts technical complexity across the pipeline:
llm_utils.py: Hides implementation differences between Ollama and Google Gemini, providing a unified interface for LLM interactions regardless of providertransform.py: Normalizes loosely-structured LLM output into strict JSON-Resume format, handling edge cases in model responsesconfig.py: Centralizes runtime configuration through environment variables includingLLM_PROVIDER,DEFAULT_MODEL, andDEVELOPMENT_MODE, plus feature flags that control pipeline behavior
These utilities ensure that core pipeline components remain agnostic to specific LLM implementations or deployment environments.
Practical Code Examples
Running a Complete Evaluation
Execute the full pipeline from the command line:
python score.py path/to/resume.pdf
This triggers PDF extraction, section parsing, GitHub enrichment, and scoring, printing a summary to stdout and appending results to the CSV cache when development mode is enabled.
Processing PDFs Programmatically
Use the PDFHandler class directly for custom integrations:
from pdf import PDFHandler
pdf_path = "example_resume.pdf"
handler = PDFHandler(pdf_path)
# Returns a Pydantic JSONResume object populated with all sections
resume_data = handler.process()
print(resume_data.basics.name)
The process() method handles the complete extraction and parsing flow, returning a validated data model.
Enriching with GitHub Data
Augment parsed résumés with repository intelligence:
from github import GitHubEnricher
# Assume resume_data is the output of PDFHandler
enricher = GitHubEnricher(resume_data)
enriched = enricher.enrich()
print(enriched.github.projects) # List of the 7 selected projects
The enrich() method identifies GitHub profiles, fetches metadata, and applies LLM-based project selection criteria.
Direct Evaluation Access
Compute scores independently using the evaluation engine:
from evaluator import Evaluator
evaluator = Evaluator(enriched) # enriched from the previous step
result = evaluator.evaluate()
print(result.scores) # Dict of category scores
print(result.explanation) # Human-readable reasoning
The evaluate() method returns structured scoring data suitable for automated decision pipelines or human review.
Summary
The hiring-agent project implements a clean, modular architecture for automated résumé evaluation:
- PDF extraction via
pymupdf_rag.pyandpdf.pyconverts documents to structured Markdown - Section parsing through Jinja templates and Pydantic models in
prompts/templates/,prompt.py, andmodels.pyvalidates resume structure - GitHub enrichment in
github.pyadds external validation signals through repository analysis - Fairness-aware scoring in
evaluator.pyproduces explainable category scores - Orchestration in
score.pymanages end-to-end execution and output persistence - Provider abstraction in
llm_utils.pyand configuration inconfig.pyensure deployment flexibility
This separation of concerns allows individual components to be tested, swapped, or extended without disrupting the broader pipeline.
Frequently Asked Questions
What is the entry point for running the hiring-agent pipeline?
The score.py module serves as the primary entry point. Executing python score.py path/to/resume.pdf initiates the complete pipeline, orchestrating PDF extraction, parsing, GitHub enrichment, and evaluation while handling output formatting and caching.
How does the hiring-agent project handle different LLM providers?
The project abstracts provider differences through llm_utils.py, which implements a unified interface for Ollama and Google Gemini. The config.py module reads the LLM_PROVIDER and DEFAULT_MODEL environment variables to determine which backend to instantiate, allowing the rest of the pipeline to remain provider-agnostic.
What schema validates the structured resume data?
The models.py file defines Pydantic schemas that enforce JSON-Resume compliance. These models validate LLM outputs during the section parsing stage, ensuring that fields like Basics, Work, Education, and Skills conform to expected types and structures before proceeding to evaluation.
How does the GitHub enrichment component select relevant repositories?
The github.py module fetches all candidate repositories, then applies a classification model and LLM-based selection criteria. It specifically selects the seven most relevant projects by evaluating author-commit thresholds, contribution diversity, and repository significance, filtering out forks and low-activity projects to highlight meaningful open-source contributions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →