Performance Implications of Using Different LLM Providers with Graphify: A Technical Deep Dive
Graphify’s unified backend abstraction in graphify/llm.py exposes significant performance trade-offs between cloud providers (OpenAI, Gemini, Kimi) and local Ollama inference, particularly regarding concurrency models, retry behaviors, and cost structures.
Graphify is an open-source code intelligence platform that abstracts LLM interactions through a configurable backend registry. Understanding the performance implications of using different LLM providers with Graphify is essential for optimizing throughput, managing operational costs, and ensuring reliable code analysis pipelines. The implementation details in graphify/llm.py reveal distinct architectural approaches to concurrency, error recovery, and token budget management across supported providers.
Backend Architecture and Provider Configuration
The Unified Backend Registry
In graphify/llm.py, Graphify maintains a BACKENDS dictionary that standardizes provider-specific metadata including pricing, token limits, and timeout configurations. This registry enables consistent API calls through the _call_openai_compat function while preserving provider-specific optimizations. Each backend entry defines the concurrency model, retry policies, and cost parameters that directly impact runtime performance.
Cloud Provider Specifications
OpenAI, Gemini, and Kimi backends share a high-concurrency architecture designed for distributed throughput. These providers implement a default retry policy of 7 attempts with exponential backoff over 180 seconds, configurable via GRAPHIFY_MAX_RETRIES. They enforce rate limits through API constraints but support parallel request execution via thread pools. Cost structures vary significantly: OpenAI GPT-4 charges approximately $0.03 per 1K input tokens and $0.06 per 1K output tokens, while Kimi offers a free tier with competitive latency characteristics.
Local Ollama Implementation
The Ollama backend operates at zero cost per token but implements fundamentally different performance characteristics. Unlike cloud providers, Ollama defaults to serial execution to prevent local resource exhaustion, with special handling for context window injection via the num_ctx parameter. The backend definition in BACKENDS["ollama"] explicitly sets pricing to zero and disables automatic retries by default, requiring explicit configuration for production reliability.
Performance Implications by Provider
Concurrency and Throughput
Cloud backends utilize thread pools to process multiple code files simultaneously, maximizing throughput for large ingestion tasks. In contrast, Ollama processing is forced to serial execution unless the environment variable GRAPHIFY_OLLAMA_PARALLEL is set to "1". This serialization creates a bottleneck when processing large corpora, as each request must complete before the next begins. The implementation in graphify/llm.py (around lines 970–1010) explicitly checks this environment variable before spawning threads for Ollama calls.
Retry Logic and Reliability
The _call_openai_compat function implements aggressive retry logic for cloud providers to handle transient network failures and rate limiting. However, Ollama defaults to zero retries unless GRAPHIFY_MAX_RETRIES is explicitly configured. This architectural difference means that local inference failures (due to model loading errors or GPU memory exhaustion) will terminate operations immediately rather than attempting recovery, potentially destabilizing batch processing pipelines.
Cost and Token Economics
Cloud providers incur direct costs proportional to token consumption, making retry logic and context window size critical cost factors. The BACKENDS registry exposes pricing metadata that Graphify uses to estimate operation costs. Ollama eliminates per-token costs entirely, making it economically advantageous for high-volume code analysis, though this benefit must be weighed against the throughput limitations of serial execution.
Context Window Management
For Ollama deployments, Graphify automatically derives the num_ctx parameter from requested max_completion_tokens values (see the extra_body handling around line 1030 in graphify/llm.py). This ensures the local model receives a sufficiently large context window without manual tuning. Cloud providers enforce their own context limits through API responses, requiring Graphify to truncate or split inputs when exceeding provider-specific maximums.
Configuring Graphify for Optimal Performance
Enable parallel Ollama processing to improve throughput on capable hardware:
import os
os.environ["GRAPHIFY_OLLAMA_PARALLEL"] = "1"
from graphify.graph import ingest
# Now processes files in parallel using local Ollama
ingest(paths=["src/"], backend="ollama")
Configure retry policies for local reliability:
import os
# Set aggressive retry logic for Ollama to match cloud behavior
os.environ["GRAPHIFY_MAX_RETRIES"] = "5"
from graphify.llm import _call_openai_compat
Explicitly select backends and tune context windows:
from graphify.llm import detect_backend, BACKENDS
# Auto-detect based on environment variables
backend = detect_backend() # Returns "openai", "ollama", etc.
# Manual Ollama call with specific context size
_call_openai_compat(
base_url="http://localhost:11434/v1",
backend="ollama",
model="qwen2.5-coder:7b",
prompt="Analyze this code structure:",
temperature=0.0,
max_completion_tokens=8192, # Automatically sets num_ctx for Ollama
)
Summary
- Cloud providers (OpenAI, Gemini, Kimi) support high-concurrency execution with 7-attempt retry logic but incur per-token costs that scale with workload size.
- Ollama provides zero-cost inference but defaults to serial execution and zero retries, requiring explicit configuration (
GRAPHIFY_OLLAMA_PARALLEL=1,GRAPHIFY_MAX_RETRIES) for comparable throughput. - Context management differs significantly: Ollama uses auto-derived
num_ctxinjection while cloud providers enforce hard API limits. - Latency characteristics favor cloud providers for distributed workloads, while Ollama excels in offline, privacy-sensitive scenarios where network latency is eliminated.
Frequently Asked Questions
How does Ollama's serial execution affect Graphify throughput?
Ollama defaults to serial processing to prevent local GPU memory exhaustion, which significantly reduces throughput when ingesting multiple files compared to cloud providers. Setting GRAPHIFY_OLLAMA_PARALLEL=1 enables multithreading but requires sufficient local hardware resources to avoid Out-of-Memory errors.
What environment variables control Graphify's retry behavior?
The GRAPHIFY_MAX_RETRIES variable configures the number of exponential backoff attempts for all providers. Cloud defaults default to 7 retries over 180 seconds, while Ollama defaults to 0 retries unless explicitly overridden, making this variable critical for local reliability.
How does Graphify optimize context windows for local Ollama models?
Graphify automatically calculates the num_ctx parameter from your max_completion_tokens request (around line 1030 in graphify/llm.py), passing it via extra_body to the Ollama API. This ensures the model loads with sufficient context without manual configuration, though you can override this by setting Ollama-specific parameters directly.
Are there cost implications when using retry logic with cloud providers?
Yes. Since cloud providers charge per token, each retry attempt consumes additional input tokens (and potentially output tokens if the partial generation fails). The 7-attempt default policy means transient failures can multiply costs by up to 7x for affected requests, though this is mitigated by the high reliability of hosted infrastructure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →