Hiring-Agent Project Structure: Modular Python Architecture for Resume Scoring

The hiring-agent repository uses a flat, self-contained Python package structure where each module handles a specific stage of the resume-to-score pipeline, eliminating nested packages to simplify LLM-driven evaluation workflows.

The InterviewStreet hiring-agent is an open-source Python application that automates technical resume evaluation through a linear pipeline of PDF extraction, GitHub enrichment, and fairness-aware LLM scoring. This guide examines the project structure of hiring-agent, breaking down how each file in the root directory contributes to the end-to-end scoring system.

Flat Module Architecture

The repository deliberately avoids nested sub-packages, placing all logic at the repository root to create a clear 1:1 mapping between pipeline stages and Python modules. According to the InterviewStreet source code, this design choice makes the codebase immediately navigable for contributors extending the resume evaluation logic.

Core Processing Modules

PDF Ingestion and Text Extraction

The pdf.py module serves as the primary entry point for document processing, utilizing PyMuPDF to convert resume pages into Markdown-like text. It delegates low-level PDF-to-text conversion to pymupdf_rag.py, which returns a formatted string representation using Markdown styling.

As implemented in pdf.py, the module drives section-by-section LLM extraction using templates stored in the prompts/ directory, populating the JSON-Resume schemas defined in models.py.

Data Models and Provider Abstraction

models.py defines Pydantic schemas for the JSON-Resume format, ensuring structured data consistency throughout the pipeline. This file contains provider-specific wrapper classes including OllamaProvider and GeminiProvider that standardize interactions with different LLM backends.

Configuration flexibility is managed by llm_utils.py, which abstracts provider-specific initialization logic and normalizes LLM responses. The global DEVELOPMENT_MODE flag in config.py toggles caching behavior and CSV export functionality for debugging purposes.

GitHub Profile Enrichment

github.py integrates external signals by fetching GitHub profile and repository data. It classifies projects by relevance, asks the LLM to select the top-7 contributions, and appends these signals to the candidate profile before scoring begins.

Scoring and Evaluation Layer

Fairness-Aware Evaluation

The evaluator.py module executes the core scoring logic using Jinja-templated prompts to ensure consistent, bias-aware candidate assessment. It consumes normalized JSON-Resume objects produced by earlier stages and returns structured score dictionaries with human-readable summaries.

Prompt Engineering Infrastructure

prompt.py centralizes prompt-generation logic including system message construction and model selection parameters. The prompts/ directory contains all Jinja2 templates (*.jinja) used for extraction tasks, GitHub project selection, and final scoring rubrics.

Data Transformation Pipeline

transform.py acts as a normalization layer, converting loosely-structured LLM JSON outputs into the strict JSON-Resume schema. This ensures type safety before data reaches the evaluation stage in evaluator.py.

Command-Line Interface and Orchestration

score.py functions as the CLI entry point and pipeline orchestrator, wiring together PDF extraction, GitHub enrichment, and fairness evaluation into a unified execution flow. When executed with a PDF path argument, it prints human-readable reports and optionally writes structured CSV output when DEVELOPMENT_MODE is enabled in config.py.

Configuration and Dependencies

The repository includes .env.example as a template for required environment variables including LLM_PROVIDER, DEFAULT_MODEL, GEMINI_API_KEY, and GITHUB_TOKEN. Dependencies are pinned in requirements.txt, listing critical packages such as pydantic, jinja2, pymupdf, and LLM client libraries.

End-to-End Execution Flow

Understanding the data flow between modules clarifies the project structure of hiring-agent:

  1. Document Ingestion – pymupdf_rag.py converts PDF pages to Markdown-like text, consumed by pdf.py for section extraction
  2. Structured Parsing – pdf.py invokes LLM calls using templates from prompts/ to populate JSON-Resume schemas
  3. External Enrichment – github.py adds repository signals and top contribution identification
  4. Fairness Evaluation – evaluator.py applies scoring rubrics via Jinja templates from prompts/
  5. Report Generation – score.py aggregates results and outputs CLI reports or CSV exports based on config.py settings

Running the Pipeline

Install dependencies and configure environment variables before executing the scoring workflow:


# Install third-party requirements

pip install -r requirements.txt

# Configure environment variables

cp .env.example .env

# Edit .env to set LLM_PROVIDER, GEMINI_API_KEY, etc.

# Execute scoring in development mode (enables CSV export)

python score.py path/to/resume.pdf

Programmatic access to the evaluation engine is available by importing the core classes:

from evaluator import Evaluator
from models import JSONResume

evaluator = Evaluator()
result = evaluator.evaluate(resume)  # resume is a JSONResume instance

print(result.summary)  # Human-readable evaluation summary

print(result.scores)   # Structured scoring dictionary

Summary

  • The hiring-agent repository uses a flat module structure with all Python files at the root level to simplify pipeline navigation
  • PDF processing is split between pymupdf_rag.py (low-level conversion) and pdf.py (LLM-driven section extraction)
  • Data models in models.py enforce JSON-Resume schemas and provider abstractions (OllamaProvider, GeminiProvider)
  • GitHub enrichment via github.py adds repository classification and contribution ranking to candidate profiles
  • Fairness-aware scoring is implemented in evaluator.py using Jinja templates stored in the prompts/ directory
  • Pipeline orchestration happens through score.py, which supports development mode CSV exports via config.py

Frequently Asked Questions

What is the project structure of hiring-agent?

The hiring-agent project follows a flat Python architecture where each root-level file represents a distinct pipeline stage. This linear layout maps PDF ingestion (pdf.py), GitHub enrichment (github.py), and LLM scoring (evaluator.py) into a sequential workflow without nested package directories, making the codebase accessible for maintenance and extension.

Which file serves as the main entry point for the hiring-agent pipeline?

score.py acts as the CLI entry point and orchestration layer. It wires together document extraction, external data fetching, and evaluation logic, accepting a PDF file path as an argument and outputting either human-readable reports to stdout or structured CSV data when DEVELOPMENT_MODE is enabled in config.py.

How does the hiring-agent handle LLM provider configuration?

Provider abstraction is implemented across llm_utils.py and models.py. The llm_utils.py module handles initialization for Ollama and Gemini backends, while models.py defines the OllamaProvider and GeminiProvider wrapper classes. Environment variables for API keys and model selection are configured via the .env file based on the .env.example template.

Where are the prompt templates stored in the hiring-agent repository?

All Jinja2 prompt templates reside in the prompts/ directory with .jinja extensions. These templates are referenced by evaluator.py for fairness-aware scoring, by pdf.py for resume section extraction, and by github.py for project selection criteria, centralizing the LLM interaction logic away from hardcoded strings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →