Which LLM Runtimes Does llmfit Support? A Complete Guide to Local AI Backends

llmfit integrates with seven major local LLM runtimes—Ollama, MLX, llama.cpp, Docker Model Runner, LM Studio, vLLM, and RamaLama—through a unified ModelProvider trait defined in the Rust core.

The AlexsJones/llmfit repository provides a cross-platform CLI and TUI for discovering, pulling, and managing local large language models. By abstracting runtime specifics behind a common trait interface, the tool unifies access to diverse backends ranging from macOS-optimized frameworks to containerized inference engines. This article examines each supported runtime, the architectural pattern that enables extensibility, and how the codebase detects and interacts with these providers.

The ModelProvider Trait Architecture

At the heart of llmfit lies the ModelProvider trait defined in llmfit-core/src/providers.rs. This contract ensures every runtime integration implements four critical capabilities:

pub trait ModelProvider {
    fn name(&self) -> &str;                // Human-readable name shown in the UI
    fn is_available(&self) -> bool;        // Quick probe to see if the daemon/binary is reachable
    fn installed_models(&self) -> HashSet<String>;
    fn start_pull(&self, model_tag: &str) -> Result<PullHandle, String>;
}

Each concrete provider struct implements these methods to handle runtime-specific discovery protocols. For example, OllamaProvider queries the Ollama HTTP API on port 11434, while MlxProvider checks for MLX binaries in the system PATH. The UI layers in llmfit-tui/src/main.rs instantiate all providers at startup, filter for availability, and aggregate model catalogs into a unified interface.

Supported LLM Runtimes

The codebase currently ships with seven provider implementations, each targeting a distinct local inference stack:

Ollama

OllamaProvider (lines 370-380 in llmfit-core/src/providers.rs) integrates with the Ollama daemon, the most popular tool for running quantized models locally. It handles model pulling via Ollama's REST API and manages the ollama pull lifecycle.

MLX (macOS-only)

MlxProvider (lines 27-37) targets Apple's MLX framework, optimized for Apple Silicon. This provider checks for mlx-lm installations and manages GGUF files through the MLX Python ecosystem.

llama.cpp

LlamaCppProvider (lines 1884-1894) interfaces directly with llama.cpp binaries for GGUF file inference. It detects running llama-server instances or standalone binaries and manages model quantization contexts.

Docker Model Runner

DockerModelRunnerProvider (lines 2152-2162) enables support for Docker's native model management capabilities. This provider queries the Docker daemon for available AI models and orchestrates containerized inference endpoints.

LM Studio

LmStudioProvider (lines 2730-2740) connects to LM Studio's local server mode, allowing llmfit to discover and pull models from LM Studio's curated registry.

vLLM

VllmProvider (lines 3226-3236) supports the vLLM high-throughput inference engine commonly used for serving models in production-like local environments.

RamaLama

RamaLamaProvider (lines 3542-3552) integrates with RamaLama, providing access to additional specialized quantization formats and inference backends.

Runtime Detection and Initialization

When the application starts, llmfit executes a parallel probe of all provider binaries. The analysis.rs module drives this process by calling is_available() on each struct, which performs lightweight health checks without loading model weights into memory.

The following pattern from llmfit-tui/src/main.rs demonstrates how the CLI assembles the provider registry:

use llmfit_core::providers::{self, ModelProvider};

fn list_available_runtimes() -> Vec<String> {
    // Each provider implements `ModelProvider`; we create a default instance.
    let mut runtimes: Vec<Box<dyn ModelProvider>> = vec![
        Box::new(providers::OllamaProvider::new()),
        Box::new(providers::MlxProvider::new()),
        Box::new(providers::LlamaCppProvider::new()),
        Box::new(providers::DockerModelRunnerProvider::new()),
        Box::new(providers::LmStudioProvider::new()),
        Box::new(providers::VllmProvider::new()),
        Box::new(providers::RamaLamaProvider::new()),
    ];

    // Keep only those that report they are reachable.
    runtimes.retain(|p| p.is_available());

    // Return the human-readable names.
    runtimes.iter().map(|p| p.name().to_string()).collect()
}

This approach ensures users see only functional runtimes. If Ollama isn't installed, for instance, its provider is filtered from the UI without error states.

Pulling Models from Specific Runtimes

Beyond discovery, providers handle asynchronous model downloads through the start_pull() method. The following example shows error handling and progress streaming when fetching a model via Ollama:

fn pull_from_ollama(tag: &str) -> Result<(), String> {
    let provider = providers::OllamaProvider::new();
    if !provider.is_available() {
        return Err("Ollama daemon not reachable".into());
    }
    let handle = provider.start_pull(tag)?;
    // Simple progress loop – the UI normally does this in a background thread.
    while let Ok(event) = handle.receiver.recv() {
        match event {
            providers::PullEvent::Progress { status, percent } => {
                println!("{} – {:.0}% ", status, percent.unwrap_or(0.0));
            }
            providers::PullEvent::Done => {
                println!("Pull finished!");
                break;
            }
            providers::PullEvent::Error(err) => return Err(err),
        }
    }
    Ok(())
}

Each runtime implements its own PullHandle to abstract download mechanics, whether that's streaming from Docker registries, Ollama's library, or direct GGUF downloads.

TUI Integration and Filtering

In the terminal interface, providers enable runtime-specific filtering. The tui_app.rs module stores the originating runtime in each ModelFit record, allowing users to isolate models by backend:

// Inside llmfit-tui/src/tui_app.rs
fn apply_runtime_filter(&mut self, runtime_name: &str) {
    self.filtered_models = self
        .all_models
        .iter()
        .filter(|m| m.runtime == runtime_name)
        .cloned()
        .collect();
}

This design keeps the UI decoupled from provider specifics while preserving metadata about which runtime serves each model.

Summary

  • Seven runtimes supported: Ollama, MLX, llama.cpp, Docker Model Runner, LM Studio, vLLM, and RamaLama through dedicated provider structs in llmfit-core/src/providers.rs.
  • Unified trait interface: The ModelProvider contract standardizes discovery via is_available(), naming via name(), and model management via installed_models() and start_pull().
  • Automatic detection: The startup sequence probes all providers and presents only reachable runtimes to the user.
  • Extensible architecture: New backends require only a struct implementing the four-method trait, following the pattern established at lines 27-3552 of the providers module.

Frequently Asked Questions

How do I add support for a new LLM runtime to llmfit?

Implement the ModelProvider trait defined in llmfit-core/src/providers.rs. Your struct must provide name(), is_available(), installed_models(), and start_pull() methods. Add the provider to the initialization vector in llmfit-tui/src/main.rs alongside the existing seven implementations.

Why is MLX support limited to macOS?

MLX is Apple's machine learning framework optimized specifically for Apple Silicon (M1/M2/M3) GPUs and Neural Engines. The MlxProvider checks for macOS-specific binaries and frameworks that do not compile or run on Linux or Windows systems.

Can I use multiple runtimes simultaneously?

Yes. llmfit aggregates models from all available providers at startup. If you have both Ollama and LM Studio running locally, the UI will display models from both sources, and you can filter or select between them dynamically.

Where are the provider implementations located in the source code?

All provider structs reside in llmfit-core/src/providers.rs with specific line ranges: MlxProvider at lines 27-37, OllamaProvider at 370-380, LlamaCppProvider at 1884-1894, DockerModelRunnerProvider at 2152-2162, LmStudioProvider at 2730-2740, VllmProvider at 3226-3236, and RamaLamaProvider at 3542-3552.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →