How llmfit-core Discovers Ollama, MLX, llama.cpp, and Other Model Runners

llmfit-core unifies local LLM discovery through the ModelProvider trait, where each backend implements is_available(), installed_models(), and start_pull() to probe binaries, HTTP endpoints, and local caches.

The llmfit-core crate handles runtime detection for multiple local AI inference engines through a single abstraction layer. Located in llmfit-core/src/providers.rs, the discovery logic probes Ollama servers, MLX Python environments, llama.cpp binaries, and OpenAI-compatible endpoints to build a unified model catalog without manual configuration.

The ModelProvider Trait Architecture

Every local model runtime in llmfit-core implements the ModelProvider trait defined at the top of llmfit-core/src/providers.rs (lines 14‑26). This abstraction allows the TUI and CLI to treat diverse backends uniformly through three core methods:

  • is_available() – Performs a quick probe to determine if the service can be used
  • installed_models() – Returns a HashSet<String> containing model name stems the provider currently knows about
  • start_pull() – Initiates a background download and returns a PullHandle for progress polling

Concrete provider implementations handle the specific discovery mechanisms for each runtime, from HTTP health checks to filesystem scans.

Ollama Discovery via HTTP Probing

The OllamaProvider structure discovers the popular Ollama runtime through environment configuration and REST API probing.

Configuration and Endpoint Detection

OllamaProvider::default() reads the OLLAMA_HOST environment variable and normalizes it, falling back to "http://localhost:11434" with an additional IPv6 fallback to "http://127.0.0.1:11434" when needed (lines 107‑125). The detect_with_installed() method contacts the primary URL first; if the request fails, it automatically retries the fallback address and adopts the successful endpoint for all subsequent calls (lines 85‑122).

Availability Checks and Model Listing

The is_available() method performs a lightweight GET request to /api/tags with a 2‑second timeout to verify the server is responsive (lines 75‑82). Once connected, the provider parses the JSON response into TagsResponse structures and extracts model stems through build_installed_set() (lines 33‑56).

The discovery logic filters out cloud‑hosted models using OllamaModel::is_cloud() and performs family‑stem deduplication to avoid false positives (lines 84‑90). Only locally stored model tags are included in the final set.

MLX (Apple MLX) Discovery Strategy

The MlxProvider handles discovery for Apple’s MLX framework, combining HTTP server probing with Python environment inspection.

Server Candidates and Identity Verification

MlxProvider::default() checks for the MLX_LM_HOST environment variable (requiring http:// or https:// prefixes) and defaults to http://localhost:8080 (lines 90‑104). The server_candidates() method yields the configured URL plus an "omlx" fallback at http://127.0.0.1:8000 when using defaults (lines 60‑68).

Before accepting a server, fetch_candidate_models() verifies identity through the OMLX status endpoint (/api/status) when required, then fetches the OpenAI‑compatible /v1/models list. It explicitly rejects servers advertising non‑MLX identities such as llama.cpp, llama‑swap, or vLLM (lines 41‑50).

HuggingFace Cache and Python Fallback

The detect_with_installed() method first scans the local HuggingFace cache for MLX‑named repositories using scan_hf_cache_for_mlx(). On macOS, it then probes each candidate server; the first responsive server returning a valid model list is marked as available and its models added to the set. If no server responds, the provider falls back to checking whether the mlx_lm Python package is importable via check_mlx_python() (lines 59‑77).

llama.cpp (GGUF) Binary and Server Detection

The LlamaCppProvider discovers llama.cpp through filesystem binary detection and optional server probing.

Binary Discovery in PATH

The Default implementation searches for llama-cli and llama-server binaries in the system PATH using find_binary(). If neither binary is found, it attempts to contact a locally‑running server on the port specified by LLAMA_SERVER_PORT via probe_llama_server() (lines 41‑53).

The is_available() method stores detection results as server_running, while detection_hint() provides human‑readable status messages such as "server detected" or "not in PATH, set LLAMA_CPP_PATH" (lines 45‑52).

GGUF Cache Scanning

When listing installed models, the provider calls scan_hf_cache_for_gguf() to search the HuggingFace cache for GGUF repositories, combined with list_gguf_files() which walks the local GGUF directory respecting depth limits. The union of both sources produces the final model stem set (lines 25‑33).

OpenAI-Compatible Runtime Classification

For generic OpenAI‑compatible servers (Docker Model Runner, vLLM, Ferrum, LM Studio), llmfit-core uses classify_openai_endpoint() to distinguish between implementations. The function examines the owned_by field in model lists (detecting "vllm", "docker", "ferrum"), the Server HTTP header (identifying "llama.cpp"), and LM‑Studio‑specific fields like compatibility_type and state (lines 84‑110).

The fetch_openai_model_list() helper retrieves /v1/models along with optional Server headers, returning an OpenAiEndpointIdentity enum that determines which concrete provider to instantiate.

Practical Usage Examples

Listing Models Across All Providers

use llmfit_core::providers::{OllamaProvider, MlxProvider, LlamaCppProvider, ModelProvider};

fn print_detected_models() {
    let mut ollama = OllamaProvider::new();
    let (ollama_ok, ollama_set, _) = ollama.detect_with_installed();
    if ollama_ok {
        println!("Ollama models: {:?}", ollama_set);
    }

    let mlx = MlxProvider::new();
    let (mlx_ok, mlx_set) = mlx.detect_with_installed();
    if mlx_ok {
        println!("MLX models: {:?}", mlx_set);
    }

    let llama = LlamaCppProvider::new();
    if llama.is_available() {
        let (llama_set, count) = llama.installed_models_counted();
        println!("llama.cpp models ({} files): {:?}", count, llama_set);
    }
}

Pulling an Ollama Model

let provider = OllamaProvider::new();
let handle = provider.start_pull("llama3.1:8b").expect("failed to start pull");
while let Ok(event) = handle.receiver.recv() {
    println!("{:?}", event);
}

Identifying Docker Model Runner

if let Some((_list, identity)) = fetch_openai_model_list(
    "http://localhost:8000", 
    std::time::Duration::from_secs(2)
) {
    if identity == OpenAiEndpointIdentity::DockerModelRunner {
        println!("Docker Model Runner detected");
    }
}

Summary

  • Unified Trait: llmfit-core/src/providers.rs defines the ModelProvider trait (lines 14‑26) that abstracts Ollama, MLX, llama.cpp, and other backends behind consistent is_available() and installed_models() methods.
  • Ollama Discovery: Probes OLLAMA_HOST or defaults to localhost:11434, queries /api/tags, and filters cloud models while deduplicating stems (lines 75‑90).
  • MLX Strategy: Checks MLX_LM_HOST, validates server identity via /api/status, scans HuggingFace cache, and falls back to Python package detection (lines 41‑104).
  • llama.cpp Detection: Searches PATH for binaries or probes LLAMA_SERVER_PORT, then aggregates GGUF files from cache and local directories (lines 25‑53).
  • OpenAI Compatibility: Classifies endpoints by inspecting owned_by fields, Server headers, and implementation‑specific markers to distinguish vLLM, Docker Model Runner, and LM Studio (lines 84‑110).

Frequently Asked Questions

How does llmfit-core determine if Ollama is running?

The OllamaProvider::is_available() method sends a GET request to /api/tags with a 2‑second timeout. If the server responds successfully at either the OLLAMA_HOST URL or the fallback 127.0.0.1:11434 address, the provider marks Ollama as available and uses that endpoint for subsequent model queries.

What happens if MLX server detection fails?

If the MLX provider cannot connect to configured or default server candidates (including the omlx fallback at port 8000), it falls back to check_mlx_python() to verify whether the mlx_lm package is importable in the local Python environment. Available models are still discovered through scan_hf_cache_for_mlx() regardless of server status.

How does llmfit-core distinguish between llama.cpp and other OpenAI-compatible servers?

The classify_openai_endpoint() function inspects HTTP response headers and JSON fields: it looks for the Server header containing "llama.cpp", checks the owned_by field for values like "vllm" or "docker", and detects LM Studio through unique fields like compatibility_type. This classification prevents MLX or Ollama providers from misidentifying foreign endpoints.

Can llmfit-core discover models without running servers?

Yes. Both the MLX and llama.cpp providers scan the local HuggingFace cache (scan_hf_cache_for_mlx() and scan_hf_cache_for_gguf()) and filesystem directories to build model sets even when no inference server is currently running. Ollama requires a running server because it queries the /api/tags endpoint rather than scanning the Ollama filesystem directly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →