# Hiring-Agent Technology Stack: Python 3.11, LLMs, and PDF Parsing Architecture

> Explore the Hiring-Agent technology stack: Python 3.11, LLMs (Ollama/Gemini), PyMuPDF, and Pydantic. Discover its modular, provider-agnostic resume evaluation pipeline.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: architecture
- Published: 2026-07-04

---

**Hiring-Agent is a Python 3.11+ application that stitches together PyMuPDF for resume extraction, Ollama or Google Gemini for LLM inference, and Pydantic for type-safe data modeling to create a modular, provider-agnostic evaluation pipeline.**

The hiring-agent repository maintained by InterviewStreet demonstrates a complete technology stack for automated candidate screening. This open-source Python application processes PDF resumes, enriches candidate profiles via GitHub APIs, and applies fairness-aware scoring using large language models. Examining the underlying architecture reveals how the system balances local execution capabilities with cloud-based LLM flexibility through clean provider abstractions.

## Core Runtime and Language

The foundation of the hiring-agent technology stack is **Python 3.11**, pinned explicitly in the `.python-version` file. This version provides the execution environment for all downstream processing, from PDF text extraction to structured JSON generation. The codebase leverages modern Python features including type hints and dataclasses, with **Black** enforcing consistent code formatting across the entire repository.

## PDF Processing Pipeline

Resume ingestion begins with **PyMuPDF** (fitz) and the **pymupdf4llm** helper library, implemented in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py). This layer reads PDF pages and converts them to a Markdown-like text format suitable for downstream LLM processing.

The [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) module orchestrates this extraction, handling the conversion from raw binary documents to structured text chunks. This approach preserves document formatting cues—such as headers and bullet points—that help the LLM accurately identify work history, education, and skills sections.

## LLM Integration Layer

The architecture supports dual LLM providers through a unified abstraction defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py).

### Local Execution with Ollama

For privacy-sensitive deployments, the system integrates **Ollama** via the `OllamaProvider` class. This wrapper handles local model inference, allowing the entire pipeline to run offline after pulling models like `gemma3:4b`.

### Cloud Processing via Google Gemini

Alternatively, the `GeminiProvider` class integrates **Google Gemini** (including `gemini-2.5-pro`) via API calls. This provider requires a `GEMINI_API_KEY` environment variable and offers higher throughput for batch processing tasks.

### Unified Provider Interface

Both providers implement a common interface exposed through the `get_provider()` factory function. As implemented in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), this function reads the `LLM_PROVIDER` and `DEFAULT_MODEL` environment variables to instantiate the appropriate backend:

```python
from models import get_provider

provider = get_provider()               # picks Ollama or Gemini from env

response = provider.chat(messages)      # unified chat API

structured = provider.parse(response)   # returns validated Pydantic model

```

## Data Modeling and Validation

**Pydantic** serves as the data validation backbone, defined entirely within [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py). The schemas enforce **JSON-Resume** compliant structures, ensuring that LLM outputs conform to expected field types for work experience, education, and skills. This type safety extends to the GitHub enrichment data, where repository metadata and commit statistics undergo strict validation before entering the scoring pipeline.

## Prompt Engineering with Jinja2

The system uses **Jinja2** templates stored in `prompts/templates/*.jinja` (including `work.jinja`) to maintain provider-agnostic prompts. These templates define strict extraction instructions for each resume section, separating prompt logic from Python code. This modular approach allows prompt versioning and A/B testing without modifying the application source.

## GitHub Integration and HTTP Layer

Candidate enrichment relies on the **requests** library within [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) to call the GitHub REST API. This module fetches user profiles, repository lists, and commit statistics, then uses the LLM to select the most relevant projects for evaluation. The integration respects the optional `GITHUB_TOKEN` environment variable for authenticated API access, increasing rate limits and enabling private repository visibility.

## Configuration and Orchestration

### Environment Management

The **python-dotenv** library loads configuration from `.env` files, managing variables such as `LLM_PROVIDER`, `DEFAULT_MODEL`, `GEMINI_API_KEY`, and `DEVELOPMENT_MODE`. The `.env.example` file documents all required and optional settings, ensuring reproducible deployments across environments.

### Pipeline Flow

The orchestration layer chains modules in a deterministic sequence:

1. **[`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py)** serves as the CLI entry point, parsing arguments and initializing the pipeline
2. **[`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)** handles document extraction and section segmentation
3. **[`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)** enriches the profile with repository data and LLM-based project selection
4. **[`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)** applies fairness-aware scoring templates

Utility modules [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) (provider initialization and response cleaning) and [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) (normalizing LLM JSON to JSON-Resume) support this flow while maintaining modularity.

### Running the Pipeline

To execute the full workflow locally:

```bash

# Clone and configure

git clone https://github.com/interviewstreet/hiring-agent
cd hiring-agent
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

# Configure for local Ollama

ollama pull gemma3:4b
export LLM_PROVIDER=ollama
export DEFAULT_MODEL=gemma3:4b

# Execute scoring

python score.py path/to/resume.pdf

```

To switch to Gemini:

```bash
export LLM_PROVIDER=gemini
export DEFAULT_MODEL=gemini-2.5-pro
export GEMINI_API_KEY=your-key-here
python score.py path/to/resume.pdf

```

## Summary

- **Hiring-Agent** runs on **Python 3.11** with dependencies managed via [`requirements.txt`](https://github.com/interviewstreet/hiring-agent/blob/main/requirements.txt)
- **PyMuPDF** and **pymupdf4llm** handle PDF-to-text conversion in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)
- **Ollama** and **Google Gemini** providers in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) offer flexible LLM execution via a unified interface
- **Pydantic** schemas ensure type-safe validation of resume and GitHub data structures
- **Jinja2** templates in `prompts/templates/` separate prompt engineering from application logic
- **requests** powers GitHub API integration within [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)
- **python-dotenv** manages configuration through environment variables
- The orchestration chain ([`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) → [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) → [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) → [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)) processes resumes end-to-end with optional CSV export when `DEVELOPMENT_MODE=True`

## Frequently Asked Questions

### What version of Python is required for hiring-agent?

The application requires **Python 3.11 or higher**, as specified in the `.python-version` file. This version ensures compatibility with the type hinting and Pydantic validation features used throughout the codebase.

### Can hiring-agent run without an internet connection?

Yes, when configured with **Ollama** as the `LLM_PROVIDER`, the entire pipeline executes locally. The system pulls models via `ollama pull` and processes PDFs and GitHub data (if cached) without external API calls, though initial GitHub enrichment requires connectivity.

### How does the system handle different LLM providers?

The `get_provider()` function in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) implements a factory pattern that instantiates either `OllamaProvider` or `GeminiProvider` based on the `LLM_PROVIDER` environment variable. Both classes expose identical `.chat()` and `.parse()` methods, ensuring the rest of the codebase remains agnostic to the underlying model.

### Where are the prompt templates stored?

All Jinja2 prompt templates reside in the `prompts/templates/` directory, including specific templates like `work.jinja` for employment history extraction. This file structure allows version control of prompts independently from the Python application logic.