# How llmfit-core Discovers Ollama, MLX, llama.cpp, and Other Model Runners

> Discover how llmfit-core finds Ollama, MLX, and llama.cpp. Learn how the ModelProvider trait unifies local LLM discovery by probing binaries, endpoints, and caches.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-11

---

**llmfit-core unifies local LLM discovery through the `ModelProvider` trait, where each backend implements `is_available()`, `installed_models()`, and `start_pull()` to probe binaries, HTTP endpoints, and local caches.**

The `llmfit-core` crate handles runtime detection for multiple local AI inference engines through a single abstraction layer. Located in [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs), the discovery logic probes Ollama servers, MLX Python environments, llama.cpp binaries, and OpenAI-compatible endpoints to build a unified model catalog without manual configuration.

## The ModelProvider Trait Architecture

Every local model runtime in llmfit-core implements the **`ModelProvider`** trait defined at the top of [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) (lines 14‑26). This abstraction allows the TUI and CLI to treat diverse backends uniformly through three core methods:

- **`is_available()`** – Performs a quick probe to determine if the service can be used
- **`installed_models()`** – Returns a `HashSet<String>` containing model name stems the provider currently knows about  
- **`start_pull()`** – Initiates a background download and returns a `PullHandle` for progress polling

Concrete provider implementations handle the specific discovery mechanisms for each runtime, from HTTP health checks to filesystem scans.

## Ollama Discovery via HTTP Probing

The **`OllamaProvider`** structure discovers the popular Ollama runtime through environment configuration and REST API probing.

### Configuration and Endpoint Detection

`OllamaProvider::default()` reads the `OLLAMA_HOST` environment variable and normalizes it, falling back to `"http://localhost:11434"` with an additional IPv6 fallback to `"http://127.0.0.1:11434"` when needed (lines 107‑125). The `detect_with_installed()` method contacts the primary URL first; if the request fails, it automatically retries the fallback address and adopts the successful endpoint for all subsequent calls (lines 85‑122).

### Availability Checks and Model Listing

The **`is_available()`** method performs a lightweight GET request to `/api/tags` with a 2‑second timeout to verify the server is responsive (lines 75‑82). Once connected, the provider parses the JSON response into `TagsResponse` structures and extracts model stems through `build_installed_set()` (lines 33‑56). 

The discovery logic filters out cloud‑hosted models using `OllamaModel::is_cloud()` and performs family‑stem deduplication to avoid false positives (lines 84‑90). Only locally stored model tags are included in the final set.

## MLX (Apple MLX) Discovery Strategy

The **`MlxProvider`** handles discovery for Apple’s MLX framework, combining HTTP server probing with Python environment inspection.

### Server Candidates and Identity Verification

`MlxProvider::default()` checks for the `MLX_LM_HOST` environment variable (requiring `http://` or `https://` prefixes) and defaults to `http://localhost:8080` (lines 90‑104). The `server_candidates()` method yields the configured URL plus an "omlx" fallback at `http://127.0.0.1:8000` when using defaults (lines 60‑68).

Before accepting a server, **`fetch_candidate_models()`** verifies identity through the OMLX status endpoint (`/api/status`) when required, then fetches the OpenAI‑compatible `/v1/models` list. It explicitly rejects servers advertising non‑MLX identities such as llama.cpp, llama‑swap, or vLLM (lines 41‑50).

### HuggingFace Cache and Python Fallback

The **`detect_with_installed()`** method first scans the local HuggingFace cache for MLX‑named repositories using `scan_hf_cache_for_mlx()`. On macOS, it then probes each candidate server; the first responsive server returning a valid model list is marked as available and its models added to the set. If no server responds, the provider falls back to checking whether the `mlx_lm` Python package is importable via `check_mlx_python()` (lines 59‑77).

## llama.cpp (GGUF) Binary and Server Detection

The **`LlamaCppProvider`** discovers llama.cpp through filesystem binary detection and optional server probing.

### Binary Discovery in PATH

The `Default` implementation searches for `llama-cli` and `llama-server` binaries in the system `PATH` using `find_binary()`. If neither binary is found, it attempts to contact a locally‑running server on the port specified by `LLAMA_SERVER_PORT` via `probe_llama_server()` (lines 41‑53).

The **`is_available()`** method stores detection results as `server_running`, while `detection_hint()` provides human‑readable status messages such as `"server detected"` or `"not in PATH, set LLAMA_CPP_PATH"` (lines 45‑52).

### GGUF Cache Scanning

When listing installed models, the provider calls **`scan_hf_cache_for_gguf()`** to search the HuggingFace cache for GGUF repositories, combined with **`list_gguf_files()`** which walks the local GGUF directory respecting depth limits. The union of both sources produces the final model stem set (lines 25‑33).

## OpenAI-Compatible Runtime Classification

For generic OpenAI‑compatible servers (Docker Model Runner, vLLM, Ferrum, LM Studio), llmfit-core uses **`classify_openai_endpoint()`** to distinguish between implementations. The function examines the `owned_by` field in model lists (detecting `"vllm"`, `"docker"`, `"ferrum"`), the `Server` HTTP header (identifying `"llama.cpp"`), and LM‑Studio‑specific fields like `compatibility_type` and `state` (lines 84‑110).

The **`fetch_openai_model_list()`** helper retrieves `/v1/models` along with optional `Server` headers, returning an `OpenAiEndpointIdentity` enum that determines which concrete provider to instantiate.

## Practical Usage Examples

### Listing Models Across All Providers

```rust
use llmfit_core::providers::{OllamaProvider, MlxProvider, LlamaCppProvider, ModelProvider};

fn print_detected_models() {
    let mut ollama = OllamaProvider::new();
    let (ollama_ok, ollama_set, _) = ollama.detect_with_installed();
    if ollama_ok {
        println!("Ollama models: {:?}", ollama_set);
    }

    let mlx = MlxProvider::new();
    let (mlx_ok, mlx_set) = mlx.detect_with_installed();
    if mlx_ok {
        println!("MLX models: {:?}", mlx_set);
    }

    let llama = LlamaCppProvider::new();
    if llama.is_available() {
        let (llama_set, count) = llama.installed_models_counted();
        println!("llama.cpp models ({} files): {:?}", count, llama_set);
    }
}

```

### Pulling an Ollama Model

```rust
let provider = OllamaProvider::new();
let handle = provider.start_pull("llama3.1:8b").expect("failed to start pull");
while let Ok(event) = handle.receiver.recv() {
    println!("{:?}", event);
}

```

### Identifying Docker Model Runner

```rust
if let Some((_list, identity)) = fetch_openai_model_list(
    "http://localhost:8000", 
    std::time::Duration::from_secs(2)
) {
    if identity == OpenAiEndpointIdentity::DockerModelRunner {
        println!("Docker Model Runner detected");
    }
}

```

## Summary

- **Unified Trait**: [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) defines the `ModelProvider` trait (lines 14‑26) that abstracts Ollama, MLX, llama.cpp, and other backends behind consistent `is_available()` and `installed_models()` methods.
- **Ollama Discovery**: Probes `OLLAMA_HOST` or defaults to `localhost:11434`, queries `/api/tags`, and filters cloud models while deduplicating stems (lines 75‑90).
- **MLX Strategy**: Checks `MLX_LM_HOST`, validates server identity via `/api/status`, scans HuggingFace cache, and falls back to Python package detection (lines 41‑104).
- **llama.cpp Detection**: Searches `PATH` for binaries or probes `LLAMA_SERVER_PORT`, then aggregates GGUF files from cache and local directories (lines 25‑53).
- **OpenAI Compatibility**: Classifies endpoints by inspecting `owned_by` fields, `Server` headers, and implementation‑specific markers to distinguish vLLM, Docker Model Runner, and LM Studio (lines 84‑110).

## Frequently Asked Questions

### How does llmfit-core determine if Ollama is running?

The `OllamaProvider::is_available()` method sends a GET request to `/api/tags` with a 2‑second timeout. If the server responds successfully at either the `OLLAMA_HOST` URL or the fallback `127.0.0.1:11434` address, the provider marks Ollama as available and uses that endpoint for subsequent model queries.

### What happens if MLX server detection fails?

If the MLX provider cannot connect to configured or default server candidates (including the `omlx` fallback at port 8000), it falls back to `check_mlx_python()` to verify whether the `mlx_lm` package is importable in the local Python environment. Available models are still discovered through `scan_hf_cache_for_mlx()` regardless of server status.

### How does llmfit-core distinguish between llama.cpp and other OpenAI-compatible servers?

The `classify_openai_endpoint()` function inspects HTTP response headers and JSON fields: it looks for the `Server` header containing `"llama.cpp"`, checks the `owned_by` field for values like `"vllm"` or `"docker"`, and detects LM Studio through unique fields like `compatibility_type`. This classification prevents MLX or Ollama providers from misidentifying foreign endpoints.

### Can llmfit-core discover models without running servers?

Yes. Both the MLX and llama.cpp providers scan the local HuggingFace cache (`scan_hf_cache_for_mlx()` and `scan_hf_cache_for_gguf()`) and filesystem directories to build model sets even when no inference server is currently running. Ollama requires a running server because it queries the `/api/tags` endpoint rather than scanning the Ollama filesystem directly.