How to Use llmfit with a Specific Runtime or Quantization

llmfit automatically selects the optimal inference runtime based on your hardware, but you can override this choice using the --force-runtime CLI flag or the force_runtime API parameter, and control quantization with the --quant flag.

The AlexsJones/llmfit project is an open-source tool that analyzes which large language models fit your hardware by computing memory usage, throughput, and compatibility across multiple inference runtimes. By default, it detects your system specs and picks the best runtime automatically — MLX on Apple Silicon, llama.cpp for CPU, vLLM for GPU clusters, and so on. But when you need to use llmfit with a specific runtime or quantization, the tool gives you explicit overrides at every layer: CLI, REST API, and interactive TUI.

How Runtime Selection Works in llmfit

Runtime selection in llmfit flows through a single path: the InferenceRuntime enum, which lives in the core library and maps to concrete providers like LlamaCpp, Mlx, and Vllm. When you invoke a command without overrides, llmfit probes your hardware via SystemSpecs::detect() and auto-selects a runtime. The override mechanism intercepts this step.

The --force-runtime CLI Flag

When you want to use a specific runtime regardless of what llmfit detects, pass --force-runtime <runtime> to either the recommend or plan sub-command. The flag is defined in src/main.rs using clap, and it propagates through the core to ModelFit::analyze_with_forced_runtime().


# Force llama.cpp for recommendation output

llmfit recommend --force-runtime llamacpp --json --limit 5

Internally, the string you provide (e.g., llamacpp, mlx, vllm) is converted into the InferenceRuntime enum, and every downstream calculation — memory footprint, per-token latency, and fit score — is computed assuming that specific backend.

Forcing a Runtime via the REST API

For programmatic control, the embedded Axum server in llmfit-tui/src/serve_api.rs accepts a force_runtime field in the JSON request body. The handler maps the string to the same InferenceRuntime enum and calls analyze_with_forced_runtime(). Here's a sample request:

curl -X POST http://localhost:8787/api/v1/models/top \
     -H "Content-Type: application/json" \
     -d '{"limit":5,"min_fit":"good","force_runtime":"vllm"}'

Controlling Quantization with --quant

Quantization is configured separately from the runtime. The plan subcommand accepts a --quant flag with values like mlx-4bit or q4_k_m. The quantization string is parsed in llmfit-core/src/models.rs and stored on the ModelRequest structure before the fit analysis runs. This value directly feeds into the memory-usage formula that ModelFit::analyze() applies to estimate whether a model will fit within your available VRAM or system memory.


# Force a runtime and specify quantization together when planning a model

llmfit plan "Qwen/Qwen3-4B-MLX-4bit" --context 8192 \
       --force-runtime llamacpp --quant q4_k_m --json

Note that quantization names differ by runtime — mlx-4bit applies to Apple MLX, while q4_k_m is a llama.cpp quant. llmfit does not convert between these names; you must supply the one that matches your chosen runtime.

Filtering Runtimes in the Interactive TUI

The TUI gives you a visual way to narrow results. Press R to open the Runtime popup, then toggle the runtimes you want to allow (e.g., only MLX and llama.cpp, disabling vLLM). This state lives in llmfit-tui/src/tui_app.rs, which tracks runtimes, selected_runtimes, and runtime_cursor. When you close the popup, apply_filters() restricts the displayed model list to entries compatible with the selected runtimes. The popup itself is rendered in llmfit-tui/src/tui_ui.rs.

End-to-End Flow When Override Is Applied

Here's the full pipeline that runs when you use a forced runtime or quantization:

  1. Detect hardware via SystemSpecs::detect().
  2. Load the model catalogue through ModelDatabase::new().
  3. Apply overrides — --force-runtime and --quant propagate to ModelFit::analyze_with_forced_runtime().
  4. Run fit analysis — compute memory, throughput, and fit level against the selected runtime.
  5. Render results — CLI table, JSON output, TUI view, or API JSON response.

Both front-ends (CLI and TUI) and the API share the same core logic in llmfit-core, so you get identical results regardless of how you invoke it.

Summary

Frequently Asked Questions

What runtimes does llmfit support with --force-runtime?

llmfit supports the runtimes enumerated in the InferenceRuntime enum in llmfit-core, including llamacpp, mlx, and vllm. The exact list may grow as the project evolves, so check llmfit recommend --help to see which values are accepted on your installed version.

Does --force-runtime also force a quantization format?

No. Runtime and quantization are independent options. --force-runtime selects the inference backend, while --quant sets the model's quantization format. You must supply a quantization format that is compatible with the runtime you chose.

Can I filter runtimes in the TUI without restarting?

Yes. Press the R key while in the interactive model list to open the Runtime popup. Toggle runtimes on or off, then press Enter; the model list updates immediately using the filter state stored in tui_app.rs.

Where is the force_runtime API parameter documented?

The API request structure is defined in [llmfit-tui/src/serve_api.rs](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/serve_api.rs), and the endpoint parameters are documented in the repository's docs/api.md file. The force_runtime field accepts the same string values as the CLI flag.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →