How to Use llmfit with a Specific Runtime or Quantization
llmfit automatically selects the optimal inference runtime based on your hardware, but you can override this choice using the --force-runtime CLI flag or the force_runtime API parameter, and control quantization with the --quant flag.
The AlexsJones/llmfit project is an open-source tool that analyzes which large language models fit your hardware by computing memory usage, throughput, and compatibility across multiple inference runtimes. By default, it detects your system specs and picks the best runtime automatically — MLX on Apple Silicon, llama.cpp for CPU, vLLM for GPU clusters, and so on. But when you need to use llmfit with a specific runtime or quantization, the tool gives you explicit overrides at every layer: CLI, REST API, and interactive TUI.
How Runtime Selection Works in llmfit
Runtime selection in llmfit flows through a single path: the InferenceRuntime enum, which lives in the core library and maps to concrete providers like LlamaCpp, Mlx, and Vllm. When you invoke a command without overrides, llmfit probes your hardware via SystemSpecs::detect() and auto-selects a runtime. The override mechanism intercepts this step.
The --force-runtime CLI Flag
When you want to use a specific runtime regardless of what llmfit detects, pass --force-runtime <runtime> to either the recommend or plan sub-command. The flag is defined in src/main.rs using clap, and it propagates through the core to ModelFit::analyze_with_forced_runtime().
# Force llama.cpp for recommendation output
llmfit recommend --force-runtime llamacpp --json --limit 5
Internally, the string you provide (e.g., llamacpp, mlx, vllm) is converted into the InferenceRuntime enum, and every downstream calculation — memory footprint, per-token latency, and fit score — is computed assuming that specific backend.
Forcing a Runtime via the REST API
For programmatic control, the embedded Axum server in llmfit-tui/src/serve_api.rs accepts a force_runtime field in the JSON request body. The handler maps the string to the same InferenceRuntime enum and calls analyze_with_forced_runtime(). Here's a sample request:
curl -X POST http://localhost:8787/api/v1/models/top \
-H "Content-Type: application/json" \
-d '{"limit":5,"min_fit":"good","force_runtime":"vllm"}'
Controlling Quantization with --quant
Quantization is configured separately from the runtime. The plan subcommand accepts a --quant flag with values like mlx-4bit or q4_k_m. The quantization string is parsed in llmfit-core/src/models.rs and stored on the ModelRequest structure before the fit analysis runs. This value directly feeds into the memory-usage formula that ModelFit::analyze() applies to estimate whether a model will fit within your available VRAM or system memory.
# Force a runtime and specify quantization together when planning a model
llmfit plan "Qwen/Qwen3-4B-MLX-4bit" --context 8192 \
--force-runtime llamacpp --quant q4_k_m --json
Note that quantization names differ by runtime — mlx-4bit applies to Apple MLX, while q4_k_m is a llama.cpp quant. llmfit does not convert between these names; you must supply the one that matches your chosen runtime.
Filtering Runtimes in the Interactive TUI
The TUI gives you a visual way to narrow results. Press R to open the Runtime popup, then toggle the runtimes you want to allow (e.g., only MLX and llama.cpp, disabling vLLM). This state lives in llmfit-tui/src/tui_app.rs, which tracks runtimes, selected_runtimes, and runtime_cursor. When you close the popup, apply_filters() restricts the displayed model list to entries compatible with the selected runtimes. The popup itself is rendered in llmfit-tui/src/tui_ui.rs.
End-to-End Flow When Override Is Applied
Here's the full pipeline that runs when you use a forced runtime or quantization:
- Detect hardware via
SystemSpecs::detect(). - Load the model catalogue through
ModelDatabase::new(). - Apply overrides —
--force-runtimeand--quantpropagate toModelFit::analyze_with_forced_runtime(). - Run fit analysis — compute memory, throughput, and fit level against the selected runtime.
- Render results — CLI table, JSON output, TUI view, or API JSON response.
Both front-ends (CLI and TUI) and the API share the same core logic in llmfit-core, so you get identical results regardless of how you invoke it.
Summary
- To use llmfit with a specific runtime, pass
--force-runtime <name>on the CLI, include"force_runtime"in an API request, or toggle runtimes in the TUI withR. - Quantization is set independently with
--quant <format>on theplansubcommand and stored on theModelRequestobject. - The core logic lives in [
llmfit-core/src/fit.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (analyze_with_forced_runtime()), with CLI wiring insrc/main.rsand API handling inllmfit-tui/src/serve_api.rs. - Runtime names differ by backend — e.g.,
q4_k_mfor llama.cpp andmlx-4bitfor MLX — and llmfit does not translate between them.
Frequently Asked Questions
What runtimes does llmfit support with --force-runtime?
llmfit supports the runtimes enumerated in the InferenceRuntime enum in llmfit-core, including llamacpp, mlx, and vllm. The exact list may grow as the project evolves, so check llmfit recommend --help to see which values are accepted on your installed version.
Does --force-runtime also force a quantization format?
No. Runtime and quantization are independent options. --force-runtime selects the inference backend, while --quant sets the model's quantization format. You must supply a quantization format that is compatible with the runtime you chose.
Can I filter runtimes in the TUI without restarting?
Yes. Press the R key while in the interactive model list to open the Runtime popup. Toggle runtimes on or off, then press Enter; the model list updates immediately using the filter state stored in tui_app.rs.
Where is the force_runtime API parameter documented?
The API request structure is defined in [llmfit-tui/src/serve_api.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs), and the endpoint parameters are documented in the repository's docs/api.md file. The force_runtime field accepts the same string values as the CLI flag.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →