# How to Use llmfit with a Specific Runtime or Quantization

> Master llmfit by learning to specify inference runtimes and quantization levels. Control your AI model's performance and resource usage with simple CLI flags and API parameters. Optimize your llmfit experience today.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: how-to-guide
- Published: 2026-08-23

---

**llmfit automatically selects the optimal inference runtime based on your hardware, but you can override this choice using the `--force-runtime` CLI flag or the `force_runtime` API parameter, and control quantization with the `--quant` flag.**

The [AlexsJones/llmfit](https://github.com/AlexsJones/llmfit) project is an open-source tool that analyzes which large language models fit your hardware by computing memory usage, throughput, and compatibility across multiple inference runtimes. By default, it detects your system specs and picks the best runtime automatically — MLX on Apple Silicon, llama.cpp for CPU, vLLM for GPU clusters, and so on. But when you need to **use llmfit with a specific runtime or quantization**, the tool gives you explicit overrides at every layer: CLI, REST API, and interactive TUI.

## How Runtime Selection Works in llmfit

Runtime selection in llmfit flows through a single path: the `InferenceRuntime` enum, which lives in the core library and maps to concrete providers like `LlamaCpp`, `Mlx`, and `Vllm`. When you invoke a command without overrides, llmfit probes your hardware via `SystemSpecs::detect()` and auto-selects a runtime. The override mechanism intercepts this step.

### The `--force-runtime` CLI Flag

When you want to use a specific runtime regardless of what llmfit detects, pass `--force-runtime <runtime>` to either the `recommend` or `plan` sub-command. The flag is defined in [`src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/main.rs) using `clap`, and it propagates through the core to `ModelFit::analyze_with_forced_runtime()`.

```sh

# Force llama.cpp for recommendation output

llmfit recommend --force-runtime llamacpp --json --limit 5

```

Internally, the string you provide (e.g., `llamacpp`, `mlx`, `vllm`) is converted into the `InferenceRuntime` enum, and every downstream calculation — memory footprint, per-token latency, and fit score — is computed assuming that specific backend.

### Forcing a Runtime via the REST API

For programmatic control, the embedded Axum server in [`llmfit-tui/src/serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs) accepts a `force_runtime` field in the JSON request body. The handler maps the string to the same `InferenceRuntime` enum and calls `analyze_with_forced_runtime()`. Here's a sample request:

```bash
curl -X POST http://localhost:8787/api/v1/models/top \
     -H "Content-Type: application/json" \
     -d '{"limit":5,"min_fit":"good","force_runtime":"vllm"}'

```

## Controlling Quantization with `--quant`

Quantization is configured separately from the runtime. The `plan` subcommand accepts a `--quant` flag with values like `mlx-4bit` or `q4_k_m`. The quantization string is parsed in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) and stored on the `ModelRequest` structure before the fit analysis runs. This value directly feeds into the memory-usage formula that `ModelFit::analyze()` applies to estimate whether a model will fit within your available VRAM or system memory.

```bash

# Force a runtime and specify quantization together when planning a model

llmfit plan "Qwen/Qwen3-4B-MLX-4bit" --context 8192 \
       --force-runtime llamacpp --quant q4_k_m --json

```

Note that quantization names differ by runtime — `mlx-4bit` applies to Apple MLX, while `q4_k_m` is a llama.cpp quant. llmfit does not convert between these names; you must supply the one that matches your chosen runtime.

## Filtering Runtimes in the Interactive TUI

The TUI gives you a visual way to narrow results. Press **`R`** to open the Runtime popup, then toggle the runtimes you want to allow (e.g., only MLX and llama.cpp, disabling vLLM). This state lives in [`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs), which tracks `runtimes`, `selected_runtimes`, and `runtime_cursor`. When you close the popup, `apply_filters()` restricts the displayed model list to entries compatible with the selected runtimes. The popup itself is rendered in [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs).

## End-to-End Flow When Override Is Applied

Here's the full pipeline that runs when you use a forced runtime or quantization:

1. **Detect hardware** via `SystemSpecs::detect()`.
2. **Load the model catalogue** through `ModelDatabase::new()`.
3. **Apply overrides** — `--force-runtime` and `--quant` propagate to `ModelFit::analyze_with_forced_runtime()`.
4. **Run fit analysis** — compute memory, throughput, and fit level against the selected runtime.
5. **Render results** — CLI table, JSON output, TUI view, or API JSON response.

Both front-ends (CLI and TUI) and the API share the same core logic in `llmfit-core`, so you get identical results regardless of how you invoke it.

## Summary

- **To use llmfit with a specific runtime**, pass `--force-runtime <name>` on the CLI, include `"force_runtime"` in an API request, or toggle runtimes in the TUI with `R`.
- **Quantization** is set independently with `--quant <format>` on the `plan` subcommand and stored on the `ModelRequest` object.
- The core logic lives in [[`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (`analyze_with_forced_runtime()`), with CLI wiring in [`src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/src/main.rs) and API handling in [`llmfit-tui/src/serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs).
- Runtime names differ by backend — e.g., `q4_k_m` for llama.cpp and `mlx-4bit` for MLX — and llmfit does not translate between them.

## Frequently Asked Questions

### What runtimes does llmfit support with `--force-runtime`?

llmfit supports the runtimes enumerated in the `InferenceRuntime` enum in `llmfit-core`, including `llamacpp`, `mlx`, and `vllm`. The exact list may grow as the project evolves, so check `llmfit recommend --help` to see which values are accepted on your installed version.

### Does `--force-runtime` also force a quantization format?

No. Runtime and quantization are independent options. `--force-runtime` selects the inference backend, while `--quant` sets the model's quantization format. You must supply a quantization format that is compatible with the runtime you chose.

### Can I filter runtimes in the TUI without restarting?

Yes. Press the `R` key while in the interactive model list to open the Runtime popup. Toggle runtimes on or off, then press Enter; the model list updates immediately using the filter state stored in [`tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/tui_app.rs).

### Where is the `force_runtime` API parameter documented?

The API request structure is defined in [[`llmfit-tui/src/serve_api.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs)](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs), and the endpoint parameters are documented in the repository's [`docs/api.md`](https://github.com/AlexsJones/llmfit/blob/main/docs/api.md) file. The `force_runtime` field accepts the same string values as the CLI flag.