How to use Ollama with the Hiring Agent for Local LLM Inference

Set the LLM_PROVIDER environment variable to ollama and DEFAULT_MODEL to your desired local model (e.g., gemma3:4b) to route all inference through the OllamaProvider class and keep resume data entirely on your local machine.

The Hiring Agent from interviewstreet/hiring-agent is an open-source resume evaluation pipeline that abstracts LLM interactions behind a provider-agnostic interface. By configuring the OllamaProvider in models.py, you can execute the full workflow—from PDF parsing to GitHub enrichment—without transmitting sensitive candidate data to external APIs.

Understanding the LLM Provider Architecture

The codebase isolates all language-model-specific logic behind a clean protocol-based interface, allowing you to swap between local Ollama instances and cloud providers by changing environment variables rather than refactoring business logic.

The LLMProvider Protocol

In models.py, the LLMProvider protocol defines the contract that all model implementations must satisfy. The repository ships with two concrete classes: OllamaProvider for local inference and GeminiProvider for Google's API. According to the source code in models.py (lines 71-115), the OllamaProvider class wraps the Ollama Python client to handle synchronous chat completion.

OllamaProvider Implementation Details

The OllamaProvider configures a 32,000 token context window (num_ctx=32768) and explicitly strips streaming flags before invoking ollama.chat, as implemented in models.py (lines 92-101). This ensures compatibility with the pipeline's synchronous expectations while maximizing the amount of resume text that can be processed in a single request.

Configuring Ollama for Local Inference

Before running the Hiring Agent, install the Ollama daemon, pull your desired model, and configure the environment variables that control provider selection and model mapping.

Installation and Model Setup

Install Ollama and pull a compatible model such as Gemma 3 4B:


# Install Ollama

curl -fsSL https://ollama.com/install.sh | sh

# Start the daemon in the background

ollama serve &

# Pull the model

ollama pull gemma3:4b

Environment Configuration

The system determines which provider to instantiate based on the LLM_PROVIDER environment variable, while the specific model mapping resides in prompt.py (lines 46-64). Set these variables in your .env file:

cp .env.example .env

# Edit .env

LLM_PROVIDER=ollama
DEFAULT_MODEL=gemma3:4b

The default values are defined in prompt.py (lines 15-22), where DEFAULT_MODEL_NAME and DEFAULT_PROVIDER establish fallback behavior when environment variables are unset.

Running the Complete Pipeline

Once configured, the score.py entry point automatically routes all LLM calls through your local Ollama instance via the llm_utils.py factory.

Execute the evaluation with:

source .venv/bin/activate
python score.py /path/to/resume.pdf

This command triggers the following provider-agnostic pipeline:

  1. PDF extraction – pymupdf_rag.py converts the resume to Markdown.
  2. Section parsing – pdf.py feeds structured sections to the LLM using Jinja templates (prompts/templates/*.jinja), calling OllamaProvider.chat through llm_utils.py.
  3. GitHub enrichment – github.py fetches public repository data mentioned in the resume.
  4. Evaluation – evaluator.py applies fairness-aware scoring rules.
  5. Output – Results print to console and, when DEVELOPMENT_MODE=True, append to resume_evaluations.csv.

Programmatic Integration

For custom workflows, instantiate the provider directly through the factory function in llm_utils.py:

import os
from llm_utils import get_llm
from pdf import PDFHandler

os.environ["LLM_PROVIDER"] = "ollama"
os.environ["DEFAULT_MODEL"] = "gemma3:4b"

llm = get_llm()  # Returns OllamaProvider instance

handler = PDFHandler(llm=llm)
resume_json = handler.process("resume.pdf")

Summary

  • The LLMProvider protocol in models.py abstracts all model interactions, enabling seamless swapping between OllamaProvider and GeminiProvider.
  • Set LLM_PROVIDER=ollama and DEFAULT_MODEL to your pulled model name to enable local inference without code changes.
  • The OllamaProvider configures a 32k context window (num_ctx=32768) and handles synchronous chat completion via ollama.chat.
  • Runtime provider selection logic resides in prompt.py (lines 46-64), with instantiation handled by llm_utils.py.
  • All pipeline stages—from pdf.py to evaluator.py—remain provider-agnostic, processing data through the standardized chat method interface.

Frequently Asked Questions

What Ollama models work with Hiring Agent?

Any model supported by your local Ollama installation works, provided you specify the correct model tag in DEFAULT_MODEL. The codebase has been tested with gemma3:4b, but you can substitute llama3, mistral, or other models by updating the environment variable and ensuring the model is pulled via ollama pull.

How does the 32k context window affect local performance?

The OllamaProvider explicitly sets num_ctx=32768 to accommodate large resumes and detailed GitHub profiles. This increases RAM usage on your local machine compared to default Ollama settings, but prevents truncation of long-form content during the evaluation phase.

Can I switch between local Ollama and Google Gemini without modifying code?

Yes. The architecture isolates provider-specific logic entirely within models.py. Changing LLM_PROVIDER from ollama to gemini and setting the appropriate API key immediately switches the pipeline to Google's API, with all other files—including pdf.py, github.py, and evaluator.py—requiring no modifications.

Where does the model-to-provider mapping occur?

The mapping logic resides in prompt.py (lines 46-64), where the system reads LLM_PROVIDER and resolves it to ModelProvider.OLLAMA or ModelProvider.GEMINI. This enum value then determines which concrete class llm_utils.py instantiates at runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →