How to Set Up the Hiring-Agent Environment for Development
You can set up the hiring-agent development environment by cloning the repository, installing Python 3.11+ dependencies from requirements.txt, configuring your .env file for either local Ollama or cloud Gemini LLM providers, and running python score.py on a résumé PDF.
Hiring-agent is an open-source Python pipeline from InterviewStreet that extracts résumé data from PDFs, enriches it with GitHub signals, and generates fair, explainable evaluations using large language models. To set up the hiring-agent environment for development, you need Python 3.11+, the dependencies listed in requirements.txt, and either a local Ollama instance or Google Gemini API credentials.
Prerequisites
Before installing the hiring-agent pipeline, ensure your system meets these requirements:
- Python 3.11 or higher (required for Pydantic and modern async features used in
models.py) - Git for cloning the repository
- Ollama (optional) for local LLM inference, or a Google Gemini API key for cloud-based inference
Step 1: Clone the Repository and Install Dependencies
Start by cloning the interviewstreet/hiring-agent repository and creating an isolated Python environment:
git clone https://github.com/interviewstreet/hiring-agent
cd hiring-agent
python -m venv .venv
source .venv/bin/activate # macOS/Linux
# .venv\Scripts\activate # Windows
pip install -r requirements.txt
The requirements.txt file pins critical dependencies including PyMuPDF (for PDF processing in pymupdf_rag.py and pdf.py), ollama (for local LLM calls), and pydantic (for data validation in models.py).
Step 2: Configure Environment Variables
Copy the template environment file and customize it for your setup:
cp .env.example .env
Edit .env to set the following variables as defined in config.py:
DEVELOPMENT_MODE: Set toTrueto enable JSON caching and CSV export for rapid iterationLLM_PROVIDER: Chooseollamafor local development orgeminifor cloud API accessDEFAULT_MODEL: Specify the model name (e.g.,gemma3:4bfor Ollama orgemini-profor Gemini)GEMINI_API_KEY: Required only if using the Gemini provider
Step 3: Set Up Your LLM Provider
The models.py file provides provider-agnostic wrappers that translate requests to either ollama.chat or google.generativeai based on your configuration.
Local Development with Ollama
For fully offline development, install and start Ollama:
# Install Ollama from https://ollama.com/ first
ollama serve # Starts the Ollama daemon
ollama pull gemma3:4b
Set your .env file to:
LLM_PROVIDER=ollama
DEFAULT_MODEL=gemma3:4b
DEVELOPMENT_MODE=True
Cloud Development with Gemini
For cloud-based inference without local GPU requirements:
- Obtain a Gemini API key from Google AI Studio
- Set your
.envfile to:
LLM_PROVIDER=gemini
DEFAULT_MODEL=gemini-pro
GEMINI_API_KEY=your_key_here
DEVELOPMENT_MODE=True
Step 4: Run the End-to-End Pipeline
Execute the orchestration script score.py to process a résumé PDF:
python score.py path/to/resume.pdf
The pipeline executes the following sequence implemented across the source files:
pdf.pyandpymupdf_rag.pyextract text from PDF pages using PyMuPDF and convert them to Markdown-like text for LLM processinggithub.pydetects GitHub profile URLs in the résumé, fetches profile and repository data, and caches results ascache/githubcache_<basename>.jsonevaluator.pyapplies fairness-aware scoring rules evaluating open-source contributions, self-projects, production experience, and technical skillsscore.pyaggregates results, prints a human-readable summary, and writes a CSV row whenDEVELOPMENT_MODEis enabled
Development Mode Features
When DEVELOPMENT_MODE=True in config.py, the pipeline activates several developer-friendly features:
- JSON Caching: Intermediate extraction results are stored under
cache/to avoid re-processing PDFs during iterative development - GitHub Data Caching: Profile and repository data fetched by
github.pypersists locally to respect API rate limits - CSV Export: Evaluation results append to a CSV file for easy analysis and comparison across multiple résumés
Summary
- Clone the interviewstreet/hiring-agent repository and install Python 3.11+ dependencies via
pip install -r requirements.txt - Configure your
.envfile by copying.env.exampleand settingLLM_PROVIDER,DEFAULT_MODEL, and optionalGEMINI_API_KEY - Select either local Ollama inference (offline) or cloud Gemini API (remote) in
models.pyvia the provider configuration - Enable
DEVELOPMENT_MODE=Trueinconfig.pyto activate JSON caching and CSV exports while iterating on thescore.pypipeline - Execute
python score.py path/to/resume.pdfto run the full résumé extraction, GitHub enrichment, and fairness-aware evaluation pipeline
Frequently Asked Questions
What Python version is required for hiring-agent?
Hiring-agent requires Python 3.11 or higher to support the Pydantic schemas and type hints used in models.py and the async patterns in the LLM orchestration layer. Earlier versions may fail when validating the data models that structure résumé sections and GitHub repository metadata.
Can I run hiring-agent without an internet connection?
Yes, you can run the pipeline entirely offline by configuring Ollama as your LLM provider in .env with LLM_PROVIDER=ollama. However, the GitHub enrichment feature in github.py requires internet access to fetch profile and repository data. If offline, the pipeline will skip GitHub analysis and proceed with PDF-based evaluation only.
Where does hiring-agent store cached data during development?
When DEVELOPMENT_MODE=True in config.py, the pipeline stores intermediate JSON files in a cache/ directory at the project root. Specifically, github.py saves profile data as cache/githubcache_<basename>.json, while processed résumé sections are cached to avoid redundant LLM calls during iterative testing of score.py.
How do I switch between Ollama and Gemini providers?
Edit the .env file and change the LLM_PROVIDER value to either ollama or gemini. The models.py file dynamically imports the appropriate client library based on this setting, translating your prompts to either ollama.chat() for local models or google.generativeai for cloud APIs. Ensure you have the corresponding API key or local daemon running before executing score.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →