# LLMFIT Quantization Hierarchies for GGUF, MLX, and ONNX Models

> Discover LLMFIT quantization hierarchies for GGUF, MLX, and ONNX. Explore native GGUF, MLX 4bit/8bit mappings, and ONNX KV-quant levels with TurboQuant for efficient model compression.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-21

---

**LLMFIT supports format-specific quantization hierarchies: native GGUF quantizations (Q4_K_M, Q8_0, Q4_0, Q5_0, Q6_K, etc.) with experimental TurboQuant for KV-compression; MLX-specific mlx-4bit and mlx-8bit mappings; and ONNX KV-quant levels (fp16, fp8, q8_0, q4_0) plus TurboQuant for CUDA backends.**

The AlexsJones/llmfit repository implements distinct quantization hierarchies across three major model-weight formats to optimize memory usage during inference. These hierarchies are encoded in the `ModelFormat` enum and associated quantization logic within [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs), allowing the system to automatically select appropriate compression strategies based on model type and hardware constraints.

## GGUF Quantization Hierarchy

GGUF models utilize the most extensive quantization hierarchy in LLMFIT, supporting both native GGUF quantization strings and experimental compression methods.

### Native GGUF Quantizations

For GGUF models, LLMFIT treats every distinct quantization string appearing in the filename as a valid quantization option. The system extracts these directly from `*.gguf` files in the model catalog, as seen in the TUI table rendering code in [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs) around line 820.

**Supported native quantizations include:**

- **Q4_K_M** – Balanced 4-bit quantization with medium compression
- **Q8_0** – 8-bit quantization with higher precision
- **Q4_0** – Standard 4-bit quantization
- **Q5_0** – 5-bit intermediate precision
- **Q6_K** – 6-bit quantization variant

Any quant name that appears in the GGUF filename is automatically recognized and made available for selection via the `--quant` CLI flag, as documented in [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs) around line 495.

### TurboQuant Support

GGUF models also support **TurboQuant** (`tq`), an experimental KV-compression method designed to make "Too-Tight" models fit within available system memory. The `FitFilter::TurboQuantFit` variant in [`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs) (line ~641) enables this fallback mechanism when standard quantizations exceed memory constraints.

## MLX Quantization Hierarchy

MLX models follow a simplified hierarchy with only two quantization levels, mapped internally to KV-quant variants.

**Supported MLX quantizations:**

- **mlx-4bit** – Maps to `KvQuant::Q4_0` in the internal representation
- **mlx-8bit** – Maps to `KvQuant::Q8_0` for higher precision

The provider logic in [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) (line ~1057) handles this mapping, ensuring MLX backends receive only these two specific quant strings regardless of the broader quantization options available for other formats.

## ONNX Quantization Hierarchy

ONNX models utilize a dedicated KV-quant hierarchy defined by the `KvQuant` enum in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs) around line 600.

### KV-Quant Levels

The ONNX quantization hierarchy supports five distinct levels:

1. **fp16** – Half-precision floating point
2. **fp8** – 8-bit floating point quantization
3. **q8_0** – 8-bit integer quantization
4. **q4_0** – 4-bit integer quantization
5. **TurboQuant** – Experimental KV-compression (CUDA only)

### Runtime Quantization and CUDA Gating

ONNX models are quantized at runtime based on the selected KV-quant level. The `KvQuant::parse` function, used in both the TUI ([`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs) line ~3096) and CLI, validates quantization strings before application.

TurboQuant for ONNX is gated to CUDA backends only, as implemented in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) (lines ~658-666), ensuring compatibility checks prevent deployment on unsupported hardware.

## Quantization Selection Workflow

LLMFIT implements a four-stage hierarchy for selecting and applying quantizations:

1. **Model Discovery** – The system reads the `model_format` field (GGUF, MLX, or ONNX) from the model catalog in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs).

2. **Quant Collection** – For GGUF models, all quant strings from `gguf_sources` are collected; for MLX and ONNX, the user-provided `--quant` flag values are processed via `model_quants` in [`llmfit-tui/src/tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_app.rs).

3. **Preference Ordering** – [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) defines a quality-first preference order: GGUF quants follow lexical filename order, while ONNX quants follow the `KvQuant::all()` list sequence.

4. **Turboquant Fallback** – When `FitFilter::TurboQuantFit` is triggered (line ~641 in [`tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/tui_app.rs)), the system attempts TurboQuant compression if standard quantizations result in "Too-Tight" memory estimates.

## Practical Usage Examples

Override default quantizations via CLI or programmatically select KV-quant levels in Rust.

### CLI Quantization Selection

Request specific GGUF or MLX quantizations using the `--quant` flag:

```bash

# Select a specific GGUF quantization

llmfit fit --model gemma-2b --quant Q4_K_M

# Use MLX 4-bit quantization

llmfit fit --model phi-2-mlx --quant mlx-4bit

```

### Programmatic ONNX Quantization

Select KV-quant levels programmatically for memory estimation:

```rust
use llmfit_core::models::{KvQuant, LlmModel};

let model: LlmModel = /* load model metadata */;
let kv = KvQuant::parse("q4_0").expect("unsupported KV quant");
let mem_gb = model.estimate_memory_gb(kv, &system_specs);

```

### Interactive Quantization Filtering

The TUI provides a `QuantPopup` interface (`draw_quant_popup` in [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs) line ~3270) for interactively toggling which quantizations appear in the model selection table.

## Summary

- **GGUF models** support native filename-based quantizations (Q4_K_M, Q8_0, etc.) plus experimental TurboQuant for extreme compression scenarios.
- **MLX models** are limited to mlx-4bit and mlx-8bit, internally mapped to Q4_0 and Q8_0 variants.
- **ONNX models** use a structured KV-quant hierarchy (fp16, fp8, q8_0, q4_0, TurboQuant) with runtime quantization and CUDA-specific TurboQuant gating.
- **Quantization selection** follows a preference order defined in [`providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/providers.rs), with fallback mechanisms in [`tui_app.rs`](https://github.com/AlexsJones/llmfit/blob/main/tui_app.rs) for memory-constrained deployments.
- All quantization validation occurs through `KvQuant::parse` and model format detection in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs).

## Frequently Asked Questions

### What GGUF quantization levels does LLMFIT support?

LLMFIT supports any quantization string appearing in a GGUF filename, including Q4_K_M, Q8_0, Q4_0, Q5_0, and Q6_K. The system dynamically discovers these from `*.gguf` filenames in the model catalog rather than maintaining a hardcoded list, as implemented in [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs) around line 820.

### How does LLMFIT handle MLX quantization differently from GGUF?

MLX quantization is restricted to `mlx-4bit` and `mlx-8bit` strings only, which the provider logic in [`llmfit-core/src/providers.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/providers.rs) maps to internal `KvQuant::Q4_0` and `KvQuant::Q8_0` variants. Unlike GGUF, MLX does not support arbitrary quantization strings from filenames.

### What is TurboQuant and when should I use it?

TurboQuant is an experimental KV-compression method available for both GGUF and ONNX models when standard quantizations result in "Too-Tight" memory fits. For ONNX, it is restricted to CUDA backends via runtime gating in [`llmfit-core/src/plan.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/plan.rs) (line ~658). It should be used as a fallback when models barely exceed available VRAM.

### How do I specify a quantization level from the command line?

Use the `--quant` flag followed by the specific quantization string. For GGUF models, use the exact quant name from the filename (e.g., `--quant Q4_K_M`). For MLX models, use `--quant mlx-4bit` or `--quant mlx-8bit`. The CLI parser in [`llmfit-tui/src/main.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/main.rs) (line ~495) validates these inputs against the supported hierarchies for each format.