How LLMFIT Detects Ollama, llama.cpp, MLX, and LM Studio Runtimes

The LLMFIT provider detection system performs single-pass startup probes in llmfit-core/src/providers.rs to discover locally installed inference runtimes, checking network reachability, binary existence, and filesystem caches to build a unified available-model catalog.

LLMFIT automatically adapts to local AI infrastructure by implementing a unified discovery mechanism for popular inference engines. The provider detection system inspects the host environment at startup to determine which of the four supported backends—Ollama, llama.cpp, MLX, or LM Studio—are operational and what models they have loaded. This detection logic is centralized in the ModelProvider trait implementation within the AlexsJones/llmfit repository.

Core Detection Architecture

All provider implementations follow a consistent two-phase discovery pattern defined in llmfit-core/src/providers.rs. First, the system establishes runtime availability through network probes or filesystem checks. Second, it aggregates installed model sets by scanning local caches or querying REST endpoints. Each provider exposes detect_with_installed() for comprehensive discovery and is_available() for lightweight status checks.

Ollama Detection via REST API Probing

The OllamaProvider implementation prioritizes network reachability over filesystem inspection, reflecting Ollama’s daemon-based architecture.

Dual-Stack Localhost Handling

The detection system initializes with OllamaProvider::default() (lines 5‑40), which constructs a primary URL of http://localhost:11434. To handle environments where localhost resolves only to IPv6, the implementation maintains a fallback to http://127.0.0.1:1144. The detect_with_installed() method (lines 85‑115) attempts a short-timeout GET request to /api/tags; if the primary URL fails, it automatically retries the fallback and adopts the working endpoint for subsequent operations.

Model Set Construction

Upon successful connection, the system invokes build_installed_set() (lines 330‑355) to parse the JSON TagsResponse. This function normalizes model identifiers to lowercase and filters out cloud-hosted entries, returning a HashSet<String> of locally available model stems. The lightweight is_available() function (lines 75‑82) performs a similar GET to /api/tags but uses a strict 2 second timeout for rapid status checks.

llama.cpp Detection via Binary and Cache Scanning

Unlike Ollama’s networked approach, the LlamaCppProvider focuses on binary discovery and filesystem scanning to identify both standalone CLI tools and running server instances.

Binary Discovery and Server Fallback

The constructor LlamaCppProvider::default() (lines 80‑92) invokes find_binary() to locate llama-cli and llama-server executables by walking the PATH and common installation directories. If neither binary exists, the system falls back to probe_llama_server(), which health-checks http://localhost:<LLAMA_SERVER_PORT> to detect running instances.

GGUF Cache Enumeration

The installed_models() method (lines 88‑102) constructs the available model catalog by recursively scanning self.models_dir for *.gguf files and merging results from scan_hf_cache_for_gguf(), which inspects the HuggingFace cache directory. The is_available() method returns true when either a binary is present or the server is detected as running, enabling detection of both interactive CLI and server deployments.

MLX Detection for macOS (Server and Python)

The MlxProvider implements a tiered detection strategy specific to Apple Silicon environments, checking for MLX-compatible servers before falling back to Python package verification.

HuggingFace Cache Inspection

The system first executes scan_hf_cache_for_mlx() (lines 66‑96), which iterates through all HuggingFace cache directories (dirs_hf_cache_all()) identifying folders matching MLX naming conventions. This returns an initial HashSet<String> of model candidates without requiring network connectivity.

Network and Python Verification

The server_candidates() method (lines 71‑81) generates probe URLs from the MLX_LM_HOST environment variable or defaults to http://127.0.0.1:8000. The fetch_candidate_models() function (lines 84‑95) validates these endpoints by checking for an omlx status payload at /api/status before calling the OpenAI-compatible /v1/models endpoint. If no server responds, the detection system executes check_mlx_python() (lines 124‑132), which attempts to import the mlx_lm package using a cached OnceLock to avoid redundant Python subprocess calls. The final detect_with_installed() result (lines 100‑115) represents the union of cache-scanned and server-reported models.

LM Studio Detection via CLI and REST

The LmStudioProvider combines application presence verification with authenticated REST enumeration to discover both the LM Studio application and its loaded models.

Application Installation Check

Before network probing, lmstudio_app_installed() (lines 24‑34) verifies LM Studio presence by checking for the lms CLI command via command_exists("lms"). If absent, it inspects platform-specific installation paths defined in lmstudio_install_candidates using Path::exists().

Authenticated Model Enumeration

The LmStudioProvider::default() constructor (lines 84‑100) normalizes the LMSTUDIO_HOST environment variable or defaults to http://127.0.0.1:1234. The detect_with_installed() method (lines 25‑56) sends a GET request to /<base_url>/v1/models with an 800 ms timeout, optionally injecting an Authorization: Bearer header if LMSTUDIO_API_KEY is configured. The response parser extracts model IDs from the LmStudioModelList JSON, storing both the full identifier and the short name suffix (e.g., qwen3 from lmstudio-community/Qwen3...) in the result set.

Integration with the LLMFIT Ecosystem

The detection results propagate through the codebase to drive both UI and analysis functionality. The llmfit-core/src/doctor.rs module aggregates provider status into system diagnostics, while llmfit-core/src/analysis.rs consumes ModelProvider::installed_models() to populate fit-analysis pipelines. In the TUI layer, llmfit-tui/src/tui_ui.rs renders availability indicators (e.g., “Ollama: ✓”), and llmfit-tui/src/tui_app.rs maintains the centralized installed state combining all provider model sets.

Practical Implementation Examples

The following Rust examples demonstrate direct usage of the provider detection system.

Checking Provider Availability

use llmfit_core::providers::{OllamaProvider, LlamaCppProvider, MlxProvider, LmStudioProvider};

fn available_providers() {
    let mut ollama = OllamaProvider::default();
    let (ollama_ok, _, _) = ollama.detect_with_installed();
    println!("Ollama available: {}", ollama_ok);

    let mlx = MlxProvider::default();
    let (mlx_ok, _) = mlx.detect_with_installed();
    println!("MLX available: {}", mlx_ok);

    let mut llama = LlamaCppProvider::default();
    println!("llama.cpp available: {}", llama.is_available());

    let lm = LmStudioProvider::default();
    let (lm_ok, _, _) = lm.detect_with_installed();
    println!("LM Studio available: {}", lm_ok);
}

Listing Installed Models

let ollama = OllamaProvider::default();
let (ok, models, count) = ollama.detect_with_installed();
if ok {
    println!("Ollama has {} locally installed models:", count);
    for m in models.iter().take(10) {   // show first 10
        println!("  {}", m);
    }
}

Pulling Models via LM Studio API

let lm = LmStudioProvider::default();
let handle = lm.start_pull("lmstudio-community/Qwen3-1.7B-MLX-4bit")
    .expect("failed to start LM Studio download");

// Simple progress loop
while let Ok(event) = handle.receiver.recv() {
    match event {
        PullEvent::Progress { status, percent } => {
            println!("{} [{:.1}%]", status, percent.unwrap_or(0.0));
        }
        PullEvent::Done => {
            println!("Download completed!");
            break;
        }
        PullEvent::Error(msg) => {
            eprintln!("Error: {}", msg);
            break;
        }
    }
}

Summary

  • The provider detection system in llmfit-core/src/providers.rs implements a unified ModelProvider trait for runtime discovery.
  • Ollama detection relies on HTTP probes to /api/tags with automatic IPv4 fallback from localhost to 127.0.0.1.
  • llama.cpp detection locates llama-cli and llama-server binaries or probes running servers, then scans for local *.gguf files and HuggingFace cache entries.
  • MLX detection on macOS unions HuggingFace cache scans with responsive server checks or Python package import verification via check_mlx_python().
  • LM Studio detection verifies CLI installation with lmstudio_app_installed() and queries the local REST API at /v1/models with optional LMSTUDIO_API_KEY authentication.

Frequently Asked Questions

What timeout values does LLMFIT use for provider health checks?

Ollama uses a 2 second timeout in is_available(), while LM Studio uses an 800 millisecond timeout in detect_with_installed(). These values balance responsiveness against the latency of local inference engines.

How does LLMFIT handle IPv6 localhost resolution issues with Ollama?

The detection system attempts to connect to http://localhost:11434 first. If that fails, it automatically falls back to http://127.0.0.1:1144 to handle systems where localhost resolves only to IPv6, ensuring reliable detection regardless of network stack configuration.

Can LLMFIT detect llama.cpp if only the server is running without CLI binaries?

Yes. If find_binary() fails to locate llama-cli or llama-server on the PATH, the system executes probe_llama_server() to check http://localhost:<port> for a running instance. Detection succeeds if either the binary exists or the server responds.

Why does the MLX provider check for Python package importability?

When no MLX-compatible server responds to network probes, check_mlx_python() attempts to import the mlx_lm package to verify that the MLX Python libraries are installed locally. This fallback ensures detection on macOS systems where users run MLX models via Python scripts rather than dedicated servers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →