Performance Implications of Using Local LLMs vs Gemini in the Hiring Agent Pipeline
Local Ollama models offer sub-100ms latency and zero per-call costs but require significant GPU/CPU resources, while Google Gemini provides managed scalability and larger models at the expense of network latency, API costs, and potential rate-limiting delays.
The interviewstreet/hiring-agent repository implements a flexible abstraction layer that supports both local Ollama instances and the Google Gemini API for resume processing. Understanding the performance implications of using local LLMs vs Gemini is critical for optimizing hiring pipelines, as the LLM call represents the most time-consuming step when extracting data from each resume section. Both backends implement the LLMProvider protocol defined in models.py, enabling seamless switching via the initialize_llm_provider function in llm_utils.py.
Latency and Cold Start Characteristics
Local Ollama models communicate via a local HTTP API, typically delivering sub-100ms response times for small prompts when the model is already resident in memory. However, the first request after a model swap incurs a cold-start penalty of several seconds while the model loads into the Ollama daemon.
In contrast, Google Gemini calls traverse the public REST endpoint, with typical response times ranging from 200ms to 1 second depending on network round-trips and backend load. The GeminiProvider class handles these calls in models.py (lines 123-165), though latency spikes occur during API throttling or high-demand periods.
Throughput and Scalability
Local deployments are bounded by the host machine's CPU/GPU capacity. The Ollama server processes requests sequentially unless explicitly configured for parallel workers, requiring additional compute resources or multiple instances to achieve horizontal scaling.
Gemini leverages Google's managed infrastructure, allowing many concurrent requests until you hit the quota limits defined for your API key. This elasticity makes it suitable for high-volume batch processing without hardware provisioning, though the GeminiProvider implements retry logic (lines 35-84 in models.py) to handle transient rate limits.
Cost and Resource Utilization
Running local models incurs zero per-call fees—only hardware-related costs for GPU, CPU, and RAM. Larger models like gemma3:12b require ≥16GB VRAM, and the provider configures a 32K context window via ollama_options["num_ctx"] = 32768 to enable longer conversations.
Gemini operates on a pay-as-you-go model where each generated token incurs charges defined in Google's pricing structure. While this eliminates local hardware constraints, costs scale linearly with usage, and the pipeline must account for potential retry delays when quotas are exceeded.
Reliability and Data Privacy
Local Ollama operates fully offline, eliminating external dependencies and keeping sensitive resume data on-premise. Failures only occur if the local daemon crashes, assuming model files are present on disk via ollama pull.
Gemini requires internet connectivity and depends on Google's service availability. The provider implements exponential backoff with jitter to handle ResourceExhausted errors, adding retry latency during rate limiting. Data privacy requires trusting Google's cloud infrastructure, though the repository warns about API key exposure risks in prompt.py where GEMINI_API_KEY is read.
Implementation in the Codebase
The provider selection logic resides in llm_utils.py, which returns either an OllamaProvider or GeminiProvider based on the DEFAULT_MODEL environment variable and the mapping in MODEL_PROVIDER_MAPPING (defined in prompt.py):
# Initialize the correct LLM provider based on DEFAULT_MODEL
provider = initialize_llm_provider(model_name=DEFAULT_MODEL) # Returns OllamaProvider or GeminiProvider
Both providers implement the chat method defined in the LLMProvider protocol. The OllamaProvider.chat implementation (lines 71-110 in models.py) calls ollama.chat, while GeminiProvider.chat (lines 123-165) wraps Google's SDK and converts responses to the expected format:
response = provider.chat(
model=DEFAULT_MODEL,
messages=[{"role": "user", "content": prompt}],
options={"temperature": 0.7},
stream=False,
)
The Gemini provider includes specific retry logic for rate limiting:
for attempt in range(MAX_RETRIES):
try:
response = gemini_model.generate_content(gemini_messages)
return {"message": {"role": "assistant", "content": response.text}}
except ResourceExhausted as e:
# Exponential back-off with optional API-provided delay
time.sleep(sleep_time)
Configuration happens via environment variables as documented in README.md:
# Use Ollama (default)
export LLM_PROVIDER=ollama
export DEFAULT_MODEL=gemma3:4b
# Or use Gemini
export LLM_PROVIDER=gemini
export DEFAULT_MODEL=gemini-2.5-pro
export GEMINI_API_KEY=your_key_here
The pipeline in pdf.py triggers separate LLM calls for each resume section (Basics, Work, Education, Skills, Projects, Awards), meaning provider latency directly multiplies across the six extraction steps, making the performance implications particularly significant for batch processing.
Summary
- Local Ollama delivers sub-100ms latency and zero per-call costs but requires significant GPU/VRAM resources (≥16GB for larger models) and experiences cold-start delays when swapping models.
- Google Gemini provides managed scalability and access to advanced models without local hardware, but adds network latency (200ms–1s), per-token costs, and potential rate-limit retries with exponential backoff.
- The Hiring Agent abstracts both providers through the
LLMProviderprotocol inmodels.py, enabling seamless switching via theLLM_PROVIDERenvironment variable andinitialize_llm_providerfunction. - Data privacy considerations favor local deployment for confidential resumes, as all data stays on-premise, while Gemini sends prompts to Google's cloud.
- The per-section processing in
pdf.pymeans provider choice directly impacts total pipeline duration, with local models excelling at sustained batch processing on capable hardware.
Frequently Asked Questions
How do I switch between local Ollama and Google Gemini in the Hiring Agent?
Set the LLM_PROVIDER environment variable to ollama or gemini, and ensure the corresponding DEFAULT_MODEL and GEMINI_API_KEY (for Gemini) are configured. The initialize_llm_provider function in llm_utils.py automatically instantiates the correct provider class based on these variables and the MODEL_PROVIDER_MAPPING in prompt.py.
Why does the first request to a local Ollama model take longer than subsequent requests?
Ollama loads models into memory on first use, causing a cold-start delay of several seconds. Once loaded, the model remains resident until swapped, delivering sub-100ms response times for subsequent calls. This behavior is inherent to the OllamaProvider implementation in models.py which calls the local Ollama HTTP API.
How does the Hiring Agent handle Gemini API rate limits?
The GeminiProvider class implements an exponential backoff with jitter strategy (lines 35-84 in models.py). When a ResourceExhausted error occurs, it retries up to 5 times with increasing delays calculated from the error response or a default backoff schedule, shielding the pipeline from transient quota errors while adding latency during throttling periods.
Which provider should I choose for processing sensitive resume data?
Local Ollama is recommended for confidential hiring workflows because all prompt data stays on-premise and never leaves your infrastructure. Google Gemini sends data to Google's cloud, requiring trust in their data handling policies and exposure to potential API key security risks documented in the repository's README.md and prompt.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →