# How to Set Up Hiring Agent Locally with Ollama

> Learn how to set up Hiring Agent locally with Ollama for offline resume evaluation. Configure LLM_PROVIDER and DEFAULT_MODEL, then run the script without API keys.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Configure `LLM_PROVIDER=ollama` and `DEFAULT_MODEL=gemma3:4b`, start the Ollama server on port 11434, and run `python score.py <resume.pdf>` to evaluate software engineering resumes entirely offline without external API keys.**

Hiring Agent is an open-source Python 3.11+ application developed by interviewstreet that parses resume PDFs, enriches candidate profiles with GitHub data, and generates fair, explainable scores using LLM-based evaluation. By leveraging the `OllamaProvider` class implemented in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), you can execute the complete pipeline locally, keeping all inference on your machine while maintaining full compatibility with the cloud-based Gemini provider.

## Prerequisites and Installation

Hiring Agent requires Python 3.11 or newer and the Ollama binary installed on your system.

1. Clone the repository and navigate to the project directory:

```bash
git clone https://github.com/interviewstreet/hiring-agent
cd hiring-agent

```

2. Create and activate a virtual environment:

```bash
python -m venv .venv
source .venv/bin/activate

# On Windows: .venv\Scripts\activate

```

3. Install the Python dependencies:

```bash
pip install -r requirements.txt

```

## Configuring Ollama as the LLM Backend

### Start the Ollama Server

Install Ollama from the official website, then launch the daemon:

```bash
ollama serve

```

This starts the HTTP API on `localhost:11434`, which the `OllamaProvider` class expects by default.

### Pull a Compatible Model

Pull a model capable of structured JSON output. The repository recommends Gemma-3 4B for its balance of speed and quality:

```bash
ollama pull gemma3:4b

```

### Environment Configuration

Create a `.env` file from the example template:

```bash
cp .env.example .env

```

Edit the file to specify Ollama as the provider:

```bash
LLM_PROVIDER=ollama
DEFAULT_MODEL=gemma3:4b

# GITHUB_TOKEN is optional but helps avoid rate limits

# GEMINI_API_KEY is not required for local mode

```

## Local Pipeline Architecture

When running with Ollama, Hiring Agent processes resumes through a five-stage pipeline that remains entirely within your local environment:

- **PDF Extraction** – [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) reads PDF pages using PyMuPDF and converts them to Markdown-like text, splitting documents into logical sections.
- **Section Parsing** – [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) sends each section to the LLM using Jinja templates stored in `prompts/templates/` to extract structured JSON Resume data.
- **GitHub Enrichment** – [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) detects GitHub usernames in the parsed resume, fetches profile and repository data via the GitHub API, and uses the LLM to select the top 7 most relevant projects.
- **Fairness Scoring** – [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) applies the open-source scoring rubric, evaluating categories like production code, technical skills, and project ownership while applying bonus and deduction logic.
- **Orchestration** – [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) coordinates the pipeline, prints human-readable reports, and writes `resume_evaluations.csv` when `DEVELOPMENT_MODE=True`.

The `OllamaProvider` class in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) (lines 71-96) handles the integration by constructing HTTP requests compatible with the Ollama API, setting a **32 KB context window** via `num_ctx=32768`, and disabling streaming mode.

## Running Resume Evaluations

Execute the end-to-end pipeline by providing a path to a resume PDF:

```bash
python score.py /path/to/resume.pdf

```

**Expected output:** A concise score summary prints to the console. If `DEVELOPMENT_MODE=True` in your environment, the system also writes a CSV file and caches intermediate JSON structures under the `cache/` directory.

To verify connectivity without a real resume, you can test the parser initialization:

```python
import subprocess
import pathlib

dummy = pathlib.Path("test.pdf")
dummy.touch()
subprocess.run(["python", "score.py", str(dummy)], check=True)

```

## Summary

- Hiring Agent supports fully local operation via the `OllamaProvider` class in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), requiring only the `LLM_PROVIDER` and `DEFAULT_MODEL` environment variables.
- The pipeline processes PDFs through [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) and [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py), enriches them via [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), and scores them using [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py).
- Ollama integration uses a 32 KB context window (`num_ctx=32768`) and disables streaming to ensure compatibility with the application's synchronous JSON parsing.
- No cloud API keys are required for local mode, though a `GITHUB_TOKEN` is recommended to avoid GitHub rate limits during enrichment.

## Frequently Asked Questions

### Do I need a GPU to run Hiring Agent with Ollama?

No, Ollama supports CPU-only inference, though performance will be significantly slower compared to GPU acceleration. For the Gemma-3 4B model recommended in the repository, a modern CPU with sufficient RAM (8GB+) can process a single resume in 30-60 seconds, while a GPU reduces this to under 10 seconds.

### Which Ollama models work best with resume evaluation?

The repository specifically recommends **Gemma-3 4B** (`gemma3:4b`) as it provides the optimal balance between inference speed and JSON output quality for structured resume parsing. Other models with strong instruction-following capabilities and JSON mode support (such as Llama 3.1 or Mistral) will also work, but may require adjustments to the context window settings in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py).

### How do I troubleshoot connection errors to the Ollama server?

Verify that `ollama serve` is running and accessible via `curl http://localhost:11434/api/tags`. If the service runs on a different port or host, update the `OLLAMA_HOST` environment variable before starting the application. The `OllamaProvider` class expects the standard Ollama HTTP API format, so ensure your pulled model name exactly matches the `DEFAULT_MODEL` value in your `.env` file.

### Can I switch between Ollama and Gemini without modifying code?

Yes. The `LLMProvider` abstraction in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) allows provider swapping by changing only environment variables. To switch from Ollama to Gemini, set `LLM_PROVIDER=gemini` and provide a `GEMINI_API_KEY`. To return to local mode, revert to `LLM_PROVIDER=ollama` and remove or comment out the Gemini API key. No changes to the pipeline logic in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) or [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) are required.