Main Components of the Hiring-Agent Project: Architecture and Pipeline Stages

The hiring-agent project consists of five core pipeline stages—PDF extraction, section parsing, GitHub enrichment, evaluation, and orchestration—supported by LLM utility modules and configuration management that transform résumé PDFs into structured, scored evaluations.

The hiring-agent repository by InterviewStreet is a modular Python application designed to automate technical candidate evaluation through LLM-powered pipeline processing. Understanding the main components of the hiring-agent project reveals a clean architecture where each module handles specific responsibilities, from ingesting raw PDF documents to generating fairness-aware scoring reports.

PDF Extraction and Document Processing

The pipeline initiates in pymupdf_rag.py and pdf.py, which collaborate to convert résumé PDFs into structured Markdown-like text. The pymupdf_rag.py module leverages PyMuPDF to read each page and construct a markdown representation that preserves headings, tables, and links. This content then passes to pdf.py, which orchestrates section-wise LLM calls to parse the document.

The PDFHandler class in pdf.py serves as the primary interface for this stage. It internally invokes pymupdf_rag.to_markdown, then processes each section through the LLM using strict Jinja templates.

Structured Resume Parsing and Validation

Once extracted, raw markdown transforms into JSON-Resume style objects through the parsing layer. This stage relies on three key components:

  • prompts/templates/*.jinja: Provider-agnostic Jinja templates that define strict prompts for sections including Basics, Work, Education, Skills, Projects, and Awards
  • prompt.py: Dispatches prompts to the LLM and manages provider-specific formatting
  • models.py: Defines Pydantic schemas that validate LLM-returned JSON, ensuring type safety and structural integrity

The interaction between these files ensures that unstructured résumé text converts into validated Python objects before proceeding to enrichment.

GitHub Profile Enrichment

The github.py module augments résumé data with signals from the candidate’s GitHub profile. The GitHubEnricher class detects GitHub usernames within the parsed résumé, then fetches profile and repository data via the GitHub API.

This component classifies each repository and invokes the LLM to select the seven most relevant projects based on minimum author-commit thresholds and diversity of contribution types. The enriched data attaches directly to the résumé object, providing concrete evidence of open-source contributions for subsequent evaluation.

Evaluation and Scoring Engine

Fairness-aware scoring occurs in evaluator.py, which implements the Evaluator class. This module uses Jinja templates embedding fairness constraints and scoring rubrics to compute category-specific scores.

The evaluation produces multiple outputs:

  • Category scores: open_source, self_projects, production, technical_skills
  • Bonuses and deductions: Adjustments based on specific criteria
  • Natural-language evidence: Human-readable reasoning for each score

The evaluate() method returns a Pydantic EvaluationResult object containing structured scores and explanatory text.

Pipeline Orchestration and Output

The score.py module functions as the end-to-end pipeline driver, tying all stages together while managing caching and output generation. When executed from the command line, it triggers the full flow from PDF ingestion to final report.

Key capabilities include:

  • Writing human-readable summaries to stdout
  • Appending CSV rows to resume_evaluations.csv when DEVELOPMENT_MODE=True
  • Caching intermediate JSON blobs under cache/ for inspection and debugging

This module centralizes workflow management, ensuring that extraction, parsing, enrichment, and evaluation execute in sequence with proper error handling.

LLM Utilities and Configuration

Supporting infrastructure abstracts technical complexity across the pipeline:

  • llm_utils.py: Hides implementation differences between Ollama and Google Gemini, providing a unified interface for LLM interactions regardless of provider
  • transform.py: Normalizes loosely-structured LLM output into strict JSON-Resume format, handling edge cases in model responses
  • config.py: Centralizes runtime configuration through environment variables including LLM_PROVIDER, DEFAULT_MODEL, and DEVELOPMENT_MODE, plus feature flags that control pipeline behavior

These utilities ensure that core pipeline components remain agnostic to specific LLM implementations or deployment environments.

Practical Code Examples

Running a Complete Evaluation

Execute the full pipeline from the command line:

python score.py path/to/resume.pdf

This triggers PDF extraction, section parsing, GitHub enrichment, and scoring, printing a summary to stdout and appending results to the CSV cache when development mode is enabled.

Processing PDFs Programmatically

Use the PDFHandler class directly for custom integrations:

from pdf import PDFHandler

pdf_path = "example_resume.pdf"
handler = PDFHandler(pdf_path)

# Returns a Pydantic JSONResume object populated with all sections

resume_data = handler.process()
print(resume_data.basics.name)

The process() method handles the complete extraction and parsing flow, returning a validated data model.

Enriching with GitHub Data

Augment parsed résumés with repository intelligence:

from github import GitHubEnricher

# Assume resume_data is the output of PDFHandler

enricher = GitHubEnricher(resume_data)
enriched = enricher.enrich()
print(enriched.github.projects)      # List of the 7 selected projects

The enrich() method identifies GitHub profiles, fetches metadata, and applies LLM-based project selection criteria.

Direct Evaluation Access

Compute scores independently using the evaluation engine:

from evaluator import Evaluator

evaluator = Evaluator(enriched)     # enriched from the previous step

result = evaluator.evaluate()
print(result.scores)                # Dict of category scores

print(result.explanation)           # Human-readable reasoning

The evaluate() method returns structured scoring data suitable for automated decision pipelines or human review.

Summary

The hiring-agent project implements a clean, modular architecture for automated résumé evaluation:

  • PDF extraction via pymupdf_rag.py and pdf.py converts documents to structured Markdown
  • Section parsing through Jinja templates and Pydantic models in prompts/templates/, prompt.py, and models.py validates resume structure
  • GitHub enrichment in github.py adds external validation signals through repository analysis
  • Fairness-aware scoring in evaluator.py produces explainable category scores
  • Orchestration in score.py manages end-to-end execution and output persistence
  • Provider abstraction in llm_utils.py and configuration in config.py ensure deployment flexibility

This separation of concerns allows individual components to be tested, swapped, or extended without disrupting the broader pipeline.

Frequently Asked Questions

What is the entry point for running the hiring-agent pipeline?

The score.py module serves as the primary entry point. Executing python score.py path/to/resume.pdf initiates the complete pipeline, orchestrating PDF extraction, parsing, GitHub enrichment, and evaluation while handling output formatting and caching.

How does the hiring-agent project handle different LLM providers?

The project abstracts provider differences through llm_utils.py, which implements a unified interface for Ollama and Google Gemini. The config.py module reads the LLM_PROVIDER and DEFAULT_MODEL environment variables to determine which backend to instantiate, allowing the rest of the pipeline to remain provider-agnostic.

What schema validates the structured resume data?

The models.py file defines Pydantic schemas that enforce JSON-Resume compliance. These models validate LLM outputs during the section parsing stage, ensuring that fields like Basics, Work, Education, and Skills conform to expected types and structures before proceeding to evaluation.

How does the GitHub enrichment component select relevant repositories?

The github.py module fetches all candidate repositories, then applies a classification model and LLM-based selection criteria. It specifically selects the seven most relevant projects by evaluating author-commit thresholds, contribution diversity, and repository significance, filtering out forks and low-activity projects to highlight meaningful open-source contributions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →