# How LLMFIT Detects Ollama, llama.cpp, MLX, and LM Studio Runtimes

> Discover how LLMFIT efficiently detects Ollama, llama.cpp, MLX, and LM Studio runtimes. Our system uses single-pass startup probes for network, binary, and cache checks, building a unified model catalog.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-08-21

---

**The LLMFIT provider detection system performs single-pass startup probes in [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) to discover locally installed inference runtimes, checking network reachability, binary existence, and filesystem caches to build a unified available-model catalog.**

LLMFIT automatically adapts to local AI infrastructure by implementing a unified discovery mechanism for popular inference engines. The **provider detection system** inspects the host environment at startup to determine which of the four supported backends—**Ollama**, **llama.cpp**, **MLX**, or **LM Studio**—are operational and what models they have loaded. This detection logic is centralized in the `ModelProvider` trait implementation within the AlexsJones/llmfit repository.

## Core Detection Architecture

All provider implementations follow a consistent two-phase discovery pattern defined in [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs). First, the system establishes **runtime availability** through network probes or filesystem checks. Second, it aggregates **installed model sets** by scanning local caches or querying REST endpoints. Each provider exposes `detect_with_installed()` for comprehensive discovery and `is_available()` for lightweight status checks.

## Ollama Detection via REST API Probing

The `OllamaProvider` implementation prioritizes network reachability over filesystem inspection, reflecting Ollama’s daemon-based architecture.

### Dual-Stack Localhost Handling

The detection system initializes with `OllamaProvider::default()` (lines 5‑40), which constructs a primary URL of `http://localhost:11434`. To handle environments where `localhost` resolves only to IPv6, the implementation maintains a fallback to `http://127.0.0.1:1144`. The `detect_with_installed()` method (lines 85‑115) attempts a short-timeout GET request to `/api/tags`; if the primary URL fails, it automatically retries the fallback and adopts the working endpoint for subsequent operations.

### Model Set Construction

Upon successful connection, the system invokes `build_installed_set()` (lines 330‑355) to parse the JSON `TagsResponse`. This function normalizes model identifiers to lowercase and filters out cloud-hosted entries, returning a `HashSet<String>` of locally available model stems. The lightweight `is_available()` function (lines 75‑82) performs a similar GET to `/api/tags` but uses a strict 2 second timeout for rapid status checks.

## llama.cpp Detection via Binary and Cache Scanning

Unlike Ollama’s networked approach, the `LlamaCppProvider` focuses on binary discovery and filesystem scanning to identify both standalone CLI tools and running server instances.

### Binary Discovery and Server Fallback

The constructor `LlamaCppProvider::default()` (lines 80‑92) invokes `find_binary()` to locate `llama-cli` and `llama-server` executables by walking the `PATH` and common installation directories. If neither binary exists, the system falls back to `probe_llama_server()`, which health-checks `http://localhost:<LLAMA_SERVER_PORT>` to detect running instances.

### GGUF Cache Enumeration

The `installed_models()` method (lines 88‑102) constructs the available model catalog by recursively scanning `self.models_dir` for `*.gguf` files and merging results from `scan_hf_cache_for_gguf()`, which inspects the HuggingFace cache directory. The `is_available()` method returns `true` when either a binary is present or the server is detected as running, enabling detection of both interactive CLI and server deployments.

## MLX Detection for macOS (Server and Python)

The `MlxProvider` implements a tiered detection strategy specific to Apple Silicon environments, checking for MLX-compatible servers before falling back to Python package verification.

### HuggingFace Cache Inspection

The system first executes `scan_hf_cache_for_mlx()` (lines 66‑96), which iterates through all HuggingFace cache directories (`dirs_hf_cache_all()`) identifying folders matching MLX naming conventions. This returns an initial `HashSet<String>` of model candidates without requiring network connectivity.

### Network and Python Verification

The `server_candidates()` method (lines 71‑81) generates probe URLs from the `MLX_LM_HOST` environment variable or defaults to `http://127.0.0.1:8000`. The `fetch_candidate_models()` function (lines 84‑95) validates these endpoints by checking for an `omlx` status payload at `/api/status` before calling the OpenAI-compatible `/v1/models` endpoint. If no server responds, the detection system executes `check_mlx_python()` (lines 124‑132), which attempts to import the `mlx_lm` package using a cached `OnceLock` to avoid redundant Python subprocess calls. The final `detect_with_installed()` result (lines 100‑115) represents the union of cache-scanned and server-reported models.

## LM Studio Detection via CLI and REST

The `LmStudioProvider` combines application presence verification with authenticated REST enumeration to discover both the LM Studio application and its loaded models.

### Application Installation Check

Before network probing, `lmstudio_app_installed()` (lines 24‑34) verifies LM Studio presence by checking for the `lms` CLI command via `command_exists("lms")`. If absent, it inspects platform-specific installation paths defined in `lmstudio_install_candidates` using `Path::exists()`.

### Authenticated Model Enumeration

The `LmStudioProvider::default()` constructor (lines 84‑100) normalizes the `LMSTUDIO_HOST` environment variable or defaults to `http://127.0.0.1:1234`. The `detect_with_installed()` method (lines 25‑56) sends a GET request to `/<base_url>/v1/models` with an 800 ms timeout, optionally injecting an `Authorization: Bearer` header if `LMSTUDIO_API_KEY` is configured. The response parser extracts model IDs from the `LmStudioModelList` JSON, storing both the full identifier and the short name suffix (e.g., `qwen3` from `lmstudio-community/Qwen3...`) in the result set.

## Integration with the LLMFIT Ecosystem

The detection results propagate through the codebase to drive both UI and analysis functionality. The [`llmfit-core/src/doctor.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/doctor.rs) module aggregates provider status into system diagnostics, while [`llmfit-core/src/analysis.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/analysis.rs) consumes `ModelProvider::installed_models()` to populate fit-analysis pipelines. In the TUI layer, [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs) renders availability indicators (e.g., “Ollama: ✓”), and [`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs) maintains the centralized `installed` state combining all provider model sets.

## Practical Implementation Examples

The following Rust examples demonstrate direct usage of the provider detection system.

### Checking Provider Availability

```rust
use llmfit_core::providers::{OllamaProvider, LlamaCppProvider, MlxProvider, LmStudioProvider};

fn available_providers() {
    let mut ollama = OllamaProvider::default();
    let (ollama_ok, _, _) = ollama.detect_with_installed();
    println!("Ollama available: {}", ollama_ok);

    let mlx = MlxProvider::default();
    let (mlx_ok, _) = mlx.detect_with_installed();
    println!("MLX available: {}", mlx_ok);

    let mut llama = LlamaCppProvider::default();
    println!("llama.cpp available: {}", llama.is_available());

    let lm = LmStudioProvider::default();
    let (lm_ok, _, _) = lm.detect_with_installed();
    println!("LM Studio available: {}", lm_ok);
}

```

### Listing Installed Models

```rust
let ollama = OllamaProvider::default();
let (ok, models, count) = ollama.detect_with_installed();
if ok {
    println!("Ollama has {} locally installed models:", count);
    for m in models.iter().take(10) {   // show first 10
        println!("  {}", m);
    }
}

```

### Pulling Models via LM Studio API

```rust
let lm = LmStudioProvider::default();
let handle = lm.start_pull("lmstudio-community/Qwen3-1.7B-MLX-4bit")
    .expect("failed to start LM Studio download");

// Simple progress loop
while let Ok(event) = handle.receiver.recv() {
    match event {
        PullEvent::Progress { status, percent } => {
            println!("{} [{:.1}%]", status, percent.unwrap_or(0.0));
        }
        PullEvent::Done => {
            println!("Download completed!");
            break;
        }
        PullEvent::Error(msg) => {
            eprintln!("Error: {}", msg);
            break;
        }
    }
}

```

## Summary

- The **provider detection system** in [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) implements a unified `ModelProvider` trait for runtime discovery.
- **Ollama** detection relies on HTTP probes to `/api/tags` with automatic IPv4 fallback from `localhost` to `127.0.0.1`.
- **llama.cpp** detection locates `llama-cli` and `llama-server` binaries or probes running servers, then scans for local `*.gguf` files and HuggingFace cache entries.
- **MLX** detection on macOS unions HuggingFace cache scans with responsive server checks or Python package import verification via `check_mlx_python()`.
- **LM Studio** detection verifies CLI installation with `lmstudio_app_installed()` and queries the local REST API at `/v1/models` with optional `LMSTUDIO_API_KEY` authentication.

## Frequently Asked Questions

### What timeout values does LLMFIT use for provider health checks?

**Ollama** uses a 2 second timeout in `is_available()`, while **LM Studio** uses an 800 millisecond timeout in `detect_with_installed()`. These values balance responsiveness against the latency of local inference engines.

### How does LLMFIT handle IPv6 localhost resolution issues with Ollama?

The detection system attempts to connect to `http://localhost:11434` first. If that fails, it automatically falls back to `http://127.0.0.1:1144` to handle systems where `localhost` resolves only to IPv6, ensuring reliable detection regardless of network stack configuration.

### Can LLMFIT detect llama.cpp if only the server is running without CLI binaries?

Yes. If `find_binary()` fails to locate `llama-cli` or `llama-server` on the `PATH`, the system executes `probe_llama_server()` to check `http://localhost:<port>` for a running instance. Detection succeeds if either the binary exists or the server responds.

### Why does the MLX provider check for Python package importability?

When no MLX-compatible server responds to network probes, `check_mlx_python()` attempts to import the `mlx_lm` package to verify that the MLX Python libraries are installed locally. This fallback ensures detection on macOS systems where users run MLX models via Python scripts rather than dedicated servers.