Hiring-Agent Project Structure: Modular Python Architecture for Resume Scoring
The hiring-agent repository uses a flat, self-contained Python package structure where each module handles a specific stage of the resume-to-score pipeline, eliminating nested packages to simplify LLM-driven evaluation workflows.
The InterviewStreet hiring-agent is an open-source Python application that automates technical resume evaluation through a linear pipeline of PDF extraction, GitHub enrichment, and fairness-aware LLM scoring. This guide examines the project structure of hiring-agent, breaking down how each file in the root directory contributes to the end-to-end scoring system.
Flat Module Architecture
The repository deliberately avoids nested sub-packages, placing all logic at the repository root to create a clear 1:1 mapping between pipeline stages and Python modules. According to the InterviewStreet source code, this design choice makes the codebase immediately navigable for contributors extending the resume evaluation logic.
Core Processing Modules
PDF Ingestion and Text Extraction
The pdf.py module serves as the primary entry point for document processing, utilizing PyMuPDF to convert resume pages into Markdown-like text. It delegates low-level PDF-to-text conversion to pymupdf_rag.py, which returns a formatted string representation using Markdown styling.
As implemented in pdf.py, the module drives section-by-section LLM extraction using templates stored in the prompts/ directory, populating the JSON-Resume schemas defined in models.py.
Data Models and Provider Abstraction
models.py defines Pydantic schemas for the JSON-Resume format, ensuring structured data consistency throughout the pipeline. This file contains provider-specific wrapper classes including OllamaProvider and GeminiProvider that standardize interactions with different LLM backends.
Configuration flexibility is managed by llm_utils.py, which abstracts provider-specific initialization logic and normalizes LLM responses. The global DEVELOPMENT_MODE flag in config.py toggles caching behavior and CSV export functionality for debugging purposes.
GitHub Profile Enrichment
github.py integrates external signals by fetching GitHub profile and repository data. It classifies projects by relevance, asks the LLM to select the top-7 contributions, and appends these signals to the candidate profile before scoring begins.
Scoring and Evaluation Layer
Fairness-Aware Evaluation
The evaluator.py module executes the core scoring logic using Jinja-templated prompts to ensure consistent, bias-aware candidate assessment. It consumes normalized JSON-Resume objects produced by earlier stages and returns structured score dictionaries with human-readable summaries.
Prompt Engineering Infrastructure
prompt.py centralizes prompt-generation logic including system message construction and model selection parameters. The prompts/ directory contains all Jinja2 templates (*.jinja) used for extraction tasks, GitHub project selection, and final scoring rubrics.
Data Transformation Pipeline
transform.py acts as a normalization layer, converting loosely-structured LLM JSON outputs into the strict JSON-Resume schema. This ensures type safety before data reaches the evaluation stage in evaluator.py.
Command-Line Interface and Orchestration
score.py functions as the CLI entry point and pipeline orchestrator, wiring together PDF extraction, GitHub enrichment, and fairness evaluation into a unified execution flow. When executed with a PDF path argument, it prints human-readable reports and optionally writes structured CSV output when DEVELOPMENT_MODE is enabled in config.py.
Configuration and Dependencies
The repository includes .env.example as a template for required environment variables including LLM_PROVIDER, DEFAULT_MODEL, GEMINI_API_KEY, and GITHUB_TOKEN. Dependencies are pinned in requirements.txt, listing critical packages such as pydantic, jinja2, pymupdf, and LLM client libraries.
End-to-End Execution Flow
Understanding the data flow between modules clarifies the project structure of hiring-agent:
- Document Ingestion –
pymupdf_rag.pyconverts PDF pages to Markdown-like text, consumed bypdf.pyfor section extraction - Structured Parsing –
pdf.pyinvokes LLM calls using templates fromprompts/to populate JSON-Resume schemas - External Enrichment –
github.pyadds repository signals and top contribution identification - Fairness Evaluation –
evaluator.pyapplies scoring rubrics via Jinja templates fromprompts/ - Report Generation –
score.pyaggregates results and outputs CLI reports or CSV exports based onconfig.pysettings
Running the Pipeline
Install dependencies and configure environment variables before executing the scoring workflow:
# Install third-party requirements
pip install -r requirements.txt
# Configure environment variables
cp .env.example .env
# Edit .env to set LLM_PROVIDER, GEMINI_API_KEY, etc.
# Execute scoring in development mode (enables CSV export)
python score.py path/to/resume.pdf
Programmatic access to the evaluation engine is available by importing the core classes:
from evaluator import Evaluator
from models import JSONResume
evaluator = Evaluator()
result = evaluator.evaluate(resume) # resume is a JSONResume instance
print(result.summary) # Human-readable evaluation summary
print(result.scores) # Structured scoring dictionary
Summary
- The hiring-agent repository uses a flat module structure with all Python files at the root level to simplify pipeline navigation
- PDF processing is split between
pymupdf_rag.py(low-level conversion) andpdf.py(LLM-driven section extraction) - Data models in
models.pyenforce JSON-Resume schemas and provider abstractions (OllamaProvider,GeminiProvider) - GitHub enrichment via
github.pyadds repository classification and contribution ranking to candidate profiles - Fairness-aware scoring is implemented in
evaluator.pyusing Jinja templates stored in theprompts/directory - Pipeline orchestration happens through
score.py, which supports development mode CSV exports viaconfig.py
Frequently Asked Questions
What is the project structure of hiring-agent?
The hiring-agent project follows a flat Python architecture where each root-level file represents a distinct pipeline stage. This linear layout maps PDF ingestion (pdf.py), GitHub enrichment (github.py), and LLM scoring (evaluator.py) into a sequential workflow without nested package directories, making the codebase accessible for maintenance and extension.
Which file serves as the main entry point for the hiring-agent pipeline?
score.py acts as the CLI entry point and orchestration layer. It wires together document extraction, external data fetching, and evaluation logic, accepting a PDF file path as an argument and outputting either human-readable reports to stdout or structured CSV data when DEVELOPMENT_MODE is enabled in config.py.
How does the hiring-agent handle LLM provider configuration?
Provider abstraction is implemented across llm_utils.py and models.py. The llm_utils.py module handles initialization for Ollama and Gemini backends, while models.py defines the OllamaProvider and GeminiProvider wrapper classes. Environment variables for API keys and model selection are configured via the .env file based on the .env.example template.
Where are the prompt templates stored in the hiring-agent repository?
All Jinja2 prompt templates reside in the prompts/ directory with .jinja extensions. These templates are referenced by evaluator.py for fairness-aware scoring, by pdf.py for resume section extraction, and by github.py for project selection criteria, centralizing the LLM interaction logic away from hardcoded strings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →