LLMFIT Quantization Hierarchies for GGUF, MLX, and ONNX Models

LLMFIT supports format-specific quantization hierarchies: native GGUF quantizations (Q4_K_M, Q8_0, Q4_0, Q5_0, Q6_K, etc.) with experimental TurboQuant for KV-compression; MLX-specific mlx-4bit and mlx-8bit mappings; and ONNX KV-quant levels (fp16, fp8, q8_0, q4_0) plus TurboQuant for CUDA backends.

The AlexsJones/llmfit repository implements distinct quantization hierarchies across three major model-weight formats to optimize memory usage during inference. These hierarchies are encoded in the ModelFormat enum and associated quantization logic within llmfit-core/src/models.rs, allowing the system to automatically select appropriate compression strategies based on model type and hardware constraints.

GGUF Quantization Hierarchy

GGUF models utilize the most extensive quantization hierarchy in LLMFIT, supporting both native GGUF quantization strings and experimental compression methods.

Native GGUF Quantizations

For GGUF models, LLMFIT treats every distinct quantization string appearing in the filename as a valid quantization option. The system extracts these directly from *.gguf files in the model catalog, as seen in the TUI table rendering code in llmfit-tui/src/tui_ui.rs around line 820.

Supported native quantizations include:

  • Q4_K_M – Balanced 4-bit quantization with medium compression
  • Q8_0 – 8-bit quantization with higher precision
  • Q4_0 – Standard 4-bit quantization
  • Q5_0 – 5-bit intermediate precision
  • Q6_K – 6-bit quantization variant

Any quant name that appears in the GGUF filename is automatically recognized and made available for selection via the --quant CLI flag, as documented in llmfit-tui/src/main.rs around line 495.

TurboQuant Support

GGUF models also support TurboQuant (tq), an experimental KV-compression method designed to make "Too-Tight" models fit within available system memory. The FitFilter::TurboQuantFit variant in llmfit-tui/src/tui_app.rs (line ~641) enables this fallback mechanism when standard quantizations exceed memory constraints.

MLX Quantization Hierarchy

MLX models follow a simplified hierarchy with only two quantization levels, mapped internally to KV-quant variants.

Supported MLX quantizations:

  • mlx-4bit – Maps to KvQuant::Q4_0 in the internal representation
  • mlx-8bit – Maps to KvQuant::Q8_0 for higher precision

The provider logic in llmfit-core/src/providers.rs (line ~1057) handles this mapping, ensuring MLX backends receive only these two specific quant strings regardless of the broader quantization options available for other formats.

ONNX Quantization Hierarchy

ONNX models utilize a dedicated KV-quant hierarchy defined by the KvQuant enum in llmfit-core/src/models.rs around line 600.

KV-Quant Levels

The ONNX quantization hierarchy supports five distinct levels:

  1. fp16 – Half-precision floating point
  2. fp8 – 8-bit floating point quantization
  3. q8_0 – 8-bit integer quantization
  4. q4_0 – 4-bit integer quantization
  5. TurboQuant – Experimental KV-compression (CUDA only)

Runtime Quantization and CUDA Gating

ONNX models are quantized at runtime based on the selected KV-quant level. The KvQuant::parse function, used in both the TUI (llmfit-tui/src/tui_app.rs line ~3096) and CLI, validates quantization strings before application.

TurboQuant for ONNX is gated to CUDA backends only, as implemented in llmfit-core/src/plan.rs (lines ~658-666), ensuring compatibility checks prevent deployment on unsupported hardware.

Quantization Selection Workflow

LLMFIT implements a four-stage hierarchy for selecting and applying quantizations:

  1. Model Discovery – The system reads the model_format field (GGUF, MLX, or ONNX) from the model catalog in llmfit-core/src/models.rs.

  2. Quant Collection – For GGUF models, all quant strings from gguf_sources are collected; for MLX and ONNX, the user-provided --quant flag values are processed via model_quants in llmfit-tui/src/tui_app.rs.

  3. Preference Ordering – llmfit-core/src/providers.rs defines a quality-first preference order: GGUF quants follow lexical filename order, while ONNX quants follow the KvQuant::all() list sequence.

  4. Turboquant Fallback – When FitFilter::TurboQuantFit is triggered (line ~641 in tui_app.rs), the system attempts TurboQuant compression if standard quantizations result in "Too-Tight" memory estimates.

Practical Usage Examples

Override default quantizations via CLI or programmatically select KV-quant levels in Rust.

CLI Quantization Selection

Request specific GGUF or MLX quantizations using the --quant flag:


# Select a specific GGUF quantization

llmfit fit --model gemma-2b --quant Q4_K_M

# Use MLX 4-bit quantization

llmfit fit --model phi-2-mlx --quant mlx-4bit

Programmatic ONNX Quantization

Select KV-quant levels programmatically for memory estimation:

use llmfit_core::models::{KvQuant, LlmModel};

let model: LlmModel = /* load model metadata */;
let kv = KvQuant::parse("q4_0").expect("unsupported KV quant");
let mem_gb = model.estimate_memory_gb(kv, &system_specs);

Interactive Quantization Filtering

The TUI provides a QuantPopup interface (draw_quant_popup in llmfit-tui/src/tui_ui.rs line ~3270) for interactively toggling which quantizations appear in the model selection table.

Summary

  • GGUF models support native filename-based quantizations (Q4_K_M, Q8_0, etc.) plus experimental TurboQuant for extreme compression scenarios.
  • MLX models are limited to mlx-4bit and mlx-8bit, internally mapped to Q4_0 and Q8_0 variants.
  • ONNX models use a structured KV-quant hierarchy (fp16, fp8, q8_0, q4_0, TurboQuant) with runtime quantization and CUDA-specific TurboQuant gating.
  • Quantization selection follows a preference order defined in providers.rs, with fallback mechanisms in tui_app.rs for memory-constrained deployments.
  • All quantization validation occurs through KvQuant::parse and model format detection in llmfit-core/src/models.rs.

Frequently Asked Questions

What GGUF quantization levels does LLMFIT support?

LLMFIT supports any quantization string appearing in a GGUF filename, including Q4_K_M, Q8_0, Q4_0, Q5_0, and Q6_K. The system dynamically discovers these from *.gguf filenames in the model catalog rather than maintaining a hardcoded list, as implemented in llmfit-tui/src/tui_ui.rs around line 820.

How does LLMFIT handle MLX quantization differently from GGUF?

MLX quantization is restricted to mlx-4bit and mlx-8bit strings only, which the provider logic in llmfit-core/src/providers.rs maps to internal KvQuant::Q4_0 and KvQuant::Q8_0 variants. Unlike GGUF, MLX does not support arbitrary quantization strings from filenames.

What is TurboQuant and when should I use it?

TurboQuant is an experimental KV-compression method available for both GGUF and ONNX models when standard quantizations result in "Too-Tight" memory fits. For ONNX, it is restricted to CUDA backends via runtime gating in llmfit-core/src/plan.rs (line ~658). It should be used as a fallback when models barely exceed available VRAM.

How do I specify a quantization level from the command line?

Use the --quant flag followed by the specific quantization string. For GGUF models, use the exact quant name from the filename (e.g., --quant Q4_K_M). For MLX models, use --quant mlx-4bit or --quant mlx-8bit. The CLI parser in llmfit-tui/src/main.rs (line ~495) validates these inputs against the supported hierarchies for each format.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →