How to Set Up Hiring Agent Locally with Ollama

Configure LLM_PROVIDER=ollama and DEFAULT_MODEL=gemma3:4b, start the Ollama server on port 11434, and run python score.py <resume.pdf> to evaluate software engineering resumes entirely offline without external API keys.

Hiring Agent is an open-source Python 3.11+ application developed by interviewstreet that parses resume PDFs, enriches candidate profiles with GitHub data, and generates fair, explainable scores using LLM-based evaluation. By leveraging the OllamaProvider class implemented in models.py, you can execute the complete pipeline locally, keeping all inference on your machine while maintaining full compatibility with the cloud-based Gemini provider.

Prerequisites and Installation

Hiring Agent requires Python 3.11 or newer and the Ollama binary installed on your system.

  1. Clone the repository and navigate to the project directory:
git clone https://github.com/interviewstreet/hiring-agent
cd hiring-agent
  1. Create and activate a virtual environment:
python -m venv .venv
source .venv/bin/activate

# On Windows: .venv\Scripts\activate
  1. Install the Python dependencies:
pip install -r requirements.txt

Configuring Ollama as the LLM Backend

Start the Ollama Server

Install Ollama from the official website, then launch the daemon:

ollama serve

This starts the HTTP API on localhost:11434, which the OllamaProvider class expects by default.

Pull a Compatible Model

Pull a model capable of structured JSON output. The repository recommends Gemma-3 4B for its balance of speed and quality:

ollama pull gemma3:4b

Environment Configuration

Create a .env file from the example template:

cp .env.example .env

Edit the file to specify Ollama as the provider:

LLM_PROVIDER=ollama
DEFAULT_MODEL=gemma3:4b

# GITHUB_TOKEN is optional but helps avoid rate limits

# GEMINI_API_KEY is not required for local mode

Local Pipeline Architecture

When running with Ollama, Hiring Agent processes resumes through a five-stage pipeline that remains entirely within your local environment:

  • PDF Extraction – pymupdf_rag.py reads PDF pages using PyMuPDF and converts them to Markdown-like text, splitting documents into logical sections.
  • Section Parsing – pdf.py sends each section to the LLM using Jinja templates stored in prompts/templates/ to extract structured JSON Resume data.
  • GitHub Enrichment – github.py detects GitHub usernames in the parsed resume, fetches profile and repository data via the GitHub API, and uses the LLM to select the top 7 most relevant projects.
  • Fairness Scoring – evaluator.py applies the open-source scoring rubric, evaluating categories like production code, technical skills, and project ownership while applying bonus and deduction logic.
  • Orchestration – score.py coordinates the pipeline, prints human-readable reports, and writes resume_evaluations.csv when DEVELOPMENT_MODE=True.

The OllamaProvider class in models.py (lines 71-96) handles the integration by constructing HTTP requests compatible with the Ollama API, setting a 32 KB context window via num_ctx=32768, and disabling streaming mode.

Running Resume Evaluations

Execute the end-to-end pipeline by providing a path to a resume PDF:

python score.py /path/to/resume.pdf

Expected output: A concise score summary prints to the console. If DEVELOPMENT_MODE=True in your environment, the system also writes a CSV file and caches intermediate JSON structures under the cache/ directory.

To verify connectivity without a real resume, you can test the parser initialization:

import subprocess
import pathlib

dummy = pathlib.Path("test.pdf")
dummy.touch()
subprocess.run(["python", "score.py", str(dummy)], check=True)

Summary

  • Hiring Agent supports fully local operation via the OllamaProvider class in models.py, requiring only the LLM_PROVIDER and DEFAULT_MODEL environment variables.
  • The pipeline processes PDFs through pymupdf_rag.py and pdf.py, enriches them via github.py, and scores them using evaluator.py.
  • Ollama integration uses a 32 KB context window (num_ctx=32768) and disables streaming to ensure compatibility with the application's synchronous JSON parsing.
  • No cloud API keys are required for local mode, though a GITHUB_TOKEN is recommended to avoid GitHub rate limits during enrichment.

Frequently Asked Questions

Do I need a GPU to run Hiring Agent with Ollama?

No, Ollama supports CPU-only inference, though performance will be significantly slower compared to GPU acceleration. For the Gemma-3 4B model recommended in the repository, a modern CPU with sufficient RAM (8GB+) can process a single resume in 30-60 seconds, while a GPU reduces this to under 10 seconds.

Which Ollama models work best with resume evaluation?

The repository specifically recommends Gemma-3 4B (gemma3:4b) as it provides the optimal balance between inference speed and JSON output quality for structured resume parsing. Other models with strong instruction-following capabilities and JSON mode support (such as Llama 3.1 or Mistral) will also work, but may require adjustments to the context window settings in models.py.

How do I troubleshoot connection errors to the Ollama server?

Verify that ollama serve is running and accessible via curl http://localhost:11434/api/tags. If the service runs on a different port or host, update the OLLAMA_HOST environment variable before starting the application. The OllamaProvider class expects the standard Ollama HTTP API format, so ensure your pulled model name exactly matches the DEFAULT_MODEL value in your .env file.

Can I switch between Ollama and Gemini without modifying code?

Yes. The LLMProvider abstraction in models.py allows provider swapping by changing only environment variables. To switch from Ollama to Gemini, set LLM_PROVIDER=gemini and provide a GEMINI_API_KEY. To return to local mode, revert to LLM_PROVIDER=ollama and remove or comment out the Gemini API key. No changes to the pipeline logic in score.py or pdf.py are required.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →