How llmfit Integrates with Ollama for LLM Management: Complete Technical Guide

llmfit integrates with Ollama through the OllamaProvider struct in llmfit-core/src/providers.rs, implementing a unified Provider trait that wraps both the Ollama CLI and REST API for model detection, lifecycle management, and inference.

The llmfit project provides a hardware-aware LLM management and benchmarking suite. Its Ollama integration allows users to discover, pull, delete, and benchmark Ollama models alongside other providers like mlx, llama.cpp, and vLLM. This article examines the complete implementation path from binary detection through UI exposure and API serving.

Provider Architecture and Trait Design

All LLM providers in llmfit implement the Provider trait defined in llmfit-core/src/providers.rs. This trait establishes a consistent interface across different backends.

The trait requires several key methods:

  • is_available() – reports whether the provider can be used
  • installed_models() – returns the set of locally available model tags
  • start_pull(&tag) – initiates an asynchronous model download
  • delete_model(&tag) – removes a model from local storage

OllamaProvider implements this trait specifically for Ollama's dual CLI/API architecture.

// From llmfit-core/src/providers.rs
pub struct OllamaProvider {
    binary_path: Option<PathBuf>,
    api_url: String,
    installed: HashSet<String>,
}

Detecting Ollama: Binary and Server Discovery

Detection occurs during application startup in llmfit-tui/src/tui_app.rs. The OllamaProvider::detect_with_installed() method performs a two-phase discovery:

Phase 1: Binary Detection

// From tui_app.rs - provider initialization
let ollama = OllamaProvider::new();
let (available, installed, count) = ollama.detect_with_installed();

The provider checks command_exists("ollama") to locate the binary on $PATH. If found, it executes ollama list to enumerate installed model tags via installed_models_counted().

Phase 2: Server Probe

If the binary is absent, llmfit attempts to contact a running Ollama server directly. This creates a fallback mode where ollama_binary_available remains false but ollama_available can still become true.

Detection results populate three application state fields:

  • ollama_available – general usability flag
  • ollama_binary_available – CLI tool presence
  • installed.ollama – HashSet of model tags

Model Lifecycle Operations

Listing Installed Models

The installed_models() method returns the cached set built during detection. The TUI displays this at line 4734 of tui_app.rs.

Pulling New Models

start_pull(&tag) spawns ollama pull <tag> as a subprocess and returns a PullHandle for progress monitoring:

// From tui_app.rs line 4332
match self.ollama.start_pull(&tag) {
    Ok(handle) => self.start_background_pull(handle),
    Err(e)    => self.show_error(e),
}

The UI polls the handle to render a progress bar without blocking the main thread.

Deleting Models

delete_model(&model_tag) executes ollama delete <tag> synchronously, called from tui_app.rs line 2318.

Operation Method Call Site
List installed installed_models() tui_app.rs:4734
Pull model start_pull(&tag) tui_app.rs:4332
Delete model delete_model(&tag) tui_app.rs:2318

Inference and Benchmarking Integration

Quality Testing via REST API

The core test suite uses Ollama's HTTP endpoint rather than CLI calls for lower latency. The quality_ollama_generate function in llmfit-core/src/quality.rs (line 175) constructs requests to /api/generate:

// From quality.rs
let result = quality_ollama_generate(
    &ollama_url,        // e.g. "http://localhost:11434"
    "llama3.1:8b",      // model tag
    "Say hello.",       // prompt
    64,                 // max_tokens
    0.3,                // temperature
);

The URL is built via api_url("/api/generate") helper method on the provider.

Throughput Benchmarking

bench_ollama(url, model, …) in llmfit-core/src/bench.rs (line 455) sends small payloads to measure tokens-per-second:

  • Warmup request to stabilize server state
  • Timed generation with max_tokens=128
  • Calculates throughput and latency percentiles

Results feed into llmfit's hardware-fit scoring alongside benchmarks from other providers.

UI and API Exposure

Terminal UI Integration

The status bar in llmfit-tui/src/tui_ui.rs (lines 146-148) renders Ollama state:


Ollama: ✓ (3 installed)

Model selection dialogs populate from installed_models(), with pull-in-progress indicators for active downloads.

HTTP API Endpoints

The Axum router in llmfit-tui/src/serve_api.rs exposes Ollama data as JSON:

// From serve_api.rs
("ollama", p.is_available(), p.installed_models())

Clients can query availability and enumerate models without TUI interaction.

MCP Server Protocol

The Model Context Protocol server (llmfit-tui/src/mcp_server.rs, line 238) includes Ollama status in its capability payload, enabling tool-use integrations to discover Ollama-hosted models.

Fallback Behavior and Remote Servers

The detection logic distinguishes between CLI availability and server reachability. This supports three operational modes:

  1. Full integration – Binary and server both present: all operations available
  2. Server-only – Remote Ollama server, no local binary: generation and benchmarking work; pull/delete unavailable
  3. Unavailable – Neither binary nor server responds: Ollama models excluded from analysis

This flexibility lets teams run llmfit on workstations while targeting Ollama servers on GPU-equipped hosts.

Key Source Files

File Responsibility
llmfit-core/src/providers.rs OllamaProvider struct and Provider trait
llmfit-tui/src/tui_app.rs Detection orchestration and UI wiring
llmfit-tui/src/tui_ui.rs Status bar rendering
llmfit-tui/src/serve_api.rs HTTP endpoint definitions
llmfit-core/src/quality.rs REST-based generation helper
llmfit-core/src/bench.rs Throughput measurement
llmfit-tui/src/mcp_server.rs MCP protocol integration

Summary

  • Unified interface – OllamaProvider implements the Provider trait for consistent treatment alongside other backends
  • Dual detection – Discovers both CLI binary and running server, with fallback to server-only operation
  • Complete lifecycle – Supports list, pull, delete, generate, and benchmark operations
  • Multi-surface exposure – Available in TUI, HTTP API, and MCP server contexts
  • REST optimization – Inference and benchmarking use HTTP API; CLI used for model management

Frequently Asked Questions

How does llmfit detect if Ollama is available?

llmfit runs command_exists("ollama") to check for the binary, then executes ollama list to enumerate models. If the binary is missing, it probes http://localhost:11434 to detect a running server. These checks populate ollama_available and ollama_binary_available flags in the application state.

Can llmfit use a remote Ollama server?

Yes. If the Ollama binary is absent but a server is reachable, ollama_available becomes true while ollama_binary_available remains false. Generation, quality testing, and benchmarking work against the remote server; pull and delete operations require the local CLI.

What Ollama operations are exposed through llmfit's HTTP API?

The Axum router in serve_api.rs exposes availability status and the installed model list. Clients receive a tuple of ("ollama", is_available, installed_models) enabling programmatic discovery of Ollama resources without TUI interaction.

How does llmfit benchmark Ollama models?

The bench_ollama function in llmfit-core/src/bench.rs sends timed generation requests to the Ollama server's /api/generate endpoint. It measures tokens-per-second and latency after a warmup request, integrating results into llmfit's hardware-fit scoring alongside other providers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →