How Quantization Levels Are Defined and Ordered in llmfit: A Complete Developer Guide
Quantization levels in llmfit are defined through three ordered hierarchies that rank formats from highest fidelity to most compressed, with automatic selection powered by helper functions for bytes-per-parameter, speed multipliers, and quality penalties.
llmfit provides a systematic approach to model quantization, defining explicit hierarchies that guide automatic format selection based on hardware constraints. The library balances memory efficiency, inference speed, and output quality through ranked quantization formats implemented in llmfit-core/src/models.rs.
The Three Quantization Hierarchies
llmfit maintains separate hierarchies for different execution backends. Each hierarchy orders formats from best quality (largest memory footprint) to most compressed (fastest inference).
Core Quantization Hierarchy
The primary hierarchy supports most backends and provides six granular levels:
// llmfit-core/src/models.rs, lines 3-6
pub const QUANT_HIERARCHY: &[&str] = &[
"Q8_0", // 8-bit, highest fidelity
"Q6_K", // 6-bit K-quant
"Q5_K_M", // 5-bit K-quant medium
"Q4_K_M", // 4-bit K-quant medium
"Q3_K_M", // 3-bit K-quant medium
"Q2_K", // 2-bit K-quant, most compressed
];
MLX-Native Quantization (Apple Silicon)
Apple Silicon devices use a simplified MLX-specific hierarchy:
// llmfit-core/src/models.rs, lines 7-9
pub const MLX_QUANT_HIERARCHY: &[&str] = &[
"mlx-8bit", // 8-bit MLX native
"mlx-4bit", // 4-bit MLX native, more compressed
];
ONNX Catalog Quantization
ONNX model compatibility uses a minimal two-level hierarchy:
// llmfit-core/src/models.rs, lines 10-12
pub const ONNX_QUANT_HIERARCHY: &[&str] = &[
"Q8_0", // 8-bit ONNX
"Q4_0", // 4-bit ONNX
];
Quantization Evaluation Functions
The models.rs module provides four helper functions that quantify the trade-offs of each quantization level. These enable the fit-analysis pipeline to compute optimal selections programmatically.
quant_bpp: Bytes Per Parameter
Returns precise memory consumption for size estimation:
// llmfit-core/src/models.rs, lines 13-38
pub fn quant_bpp(quant: &str) -> f64 {
match quant {
"Q8_0" => 1.05, // 1.05 bytes per parameter
"Q6_K" => 0.80,
"Q5_K_M" => 0.68,
"Q4_K_M" => 0.58, // 58% of original fp16 size
"Q3_K_M" => 0.48,
"Q2_K" => 0.38,
"mlx-8bit" => 1.0,
"mlx-4bit" => 0.55, // 55% of original, slightly better than Q4_K_M
"Q4_0" => 0.50,
_ => 2.0, // default fallback (fp16)
}
}
quant_speed_multiplier: Inference Speed Factor
Lower values indicate faster inference relative to baseline:
// llmfit-core/src/models.rs, lines 40-62
pub fn quant_speed_multiplier(quant: &str) -> f64 {
match quant {
"Q8_0" => 0.8, // 20% slower than baseline (heavier compute)
"Q6_K" => 0.95,
"Q5_K_M" => 1.05,
"Q4_K_M" => 1.15, // 15% faster (reduced memory bandwidth)
"Q3_K_M" => 1.25,
"Q2_K" => 1.35, // 35% fastest (most aggressive compression)
"mlx-8bit" => 0.85, // MLX optimized
"mlx-4bit" => 1.15,
_ => 1.0,
}
}
quant_quality_penalty: Output Fidelity Score
Numerical penalty applied to fit scores when selecting lower-quality formats:
// llmfit-core/src/models.rs, lines 89-115
pub fn quant_quality_penalty(quant: &str) -> i32 {
match quant {
"Q8_0" => 0, // no penalty (reference quality)
"Q6_K" => -2, // minimal degradation
"Q5_K_M" => -3,
"Q4_K_M" => -5, // noticeable but acceptable trade-off
"Q3_K_M" => -10, // significant quality loss
"Q2_K" => -20, // substantial degradation
"mlx-8bit" => -1, // Apple Silicon optimizations reduce penalty
"mlx-4bit" => -4,
"Q4_0" => -6,
_ => 0,
}
}
Practical Usage: Selecting Quantization by Memory Constraint
The following example demonstrates how llmfit's hierarchy and helper functions combine for automatic quantization selection:
use llmfit_core::models::{
QUANT_HIERARCHY,
MLX_QUANT_HIERARCHY,
quant_bpp,
quant_speed_multiplier,
quant_quality_penalty,
};
fn select_quantization_for_vram(
model_params: u64,
max_vram_bytes: f64,
) -> Option<&'static str> {
// Iterate hierarchy from best quality downward
for quant in QUANT_HIERARCHY {
let required_bytes = quant_bpp(quant) * model_params as f64;
if required_bytes <= max_vram_bytes {
println!("Selected {}: {:.2} GB required, speed multiplier {:.2}, quality penalty {}",
quant,
required_bytes / 1e9,
quant_speed_multiplier(quant),
quant_quality_penalty(quant)
);
return Some(quant);
}
}
None // No compatible quantization found
}
// Example: 7B parameter model with 4GB VRAM budget
fn main() {
let model_params = 7_000_000_000u64;
let vram_budget = 4.0 * 1024.0 * 1024.0 * 1024.0; // 4 GiB
let chosen = select_quantization_for_vram(model_params, vram_budget);
// Output: Selected Q4_K_M: 4.06 GB → rounds to fit, or falls through
// Actually computes: 0.58 * 7e9 = 4.06e9 bytes → slightly over
// Next iteration finds Q3_K_M: 0.48 * 7e9 = 3.36e9 ✓
}
Integration with the Fit Analysis Pipeline
The quantization hierarchies and helper functions serve the broader fit-analysis system in llmfit-core/src/fit.rs. This pipeline:
- Queries hardware specifications (available VRAM, memory bandwidth, compute capability)
- Enumerates valid quantizations from the appropriate hierarchy for the target backend
- Scores each candidate using
quant_bpp,quant_speed_multiplier, andquant_quality_penalty - Selects the optimal balance based on user-configurable priorities (speed vs. quality vs. memory)
The selected quantization level propagates to TUI rendering components in llmfit-tui/src/display.rs and llmfit-tui/src/tui_ui.rs for user visibility.
Summary
- Three ordered hierarchies in
llmfit-core/src/models.rsdefine quantization levels:QUANT_HIERARCHYfor general use,MLX_QUANT_HIERARCHYfor Apple Silicon, andONNX_QUANT_HIERARCHYfor ONNX compatibility - Ordering principle: highest fidelity first, most compressed last—enabling sequential fallback selection
- Four evaluation functions quantify trade-offs:
quant_bpp(memory),quant_speed_multiplier(throughput),quant_quality_penalty(accuracy), andquant_bytes_per_param(bandwidth estimates) - Automatic selection occurs in
fit.rsby iterating hierarchies and scoring candidates against hardware constraints - Quality-compression spectrum:
Q8_0(1.05 B/param, 0 penalty) toQ2_K(0.38 B/param, -20 penalty) for core formats; MLX offers optimizedmlx-4bitat 0.55 B/param with only -4 penalty
Frequently Asked Questions
What determines which quantization hierarchy llmfit uses?
The backend target determines hierarchy selection. General CUDA/CPU inference uses QUANT_HIERARCHY, Apple Silicon devices automatically select MLX_QUANT_HIERARCHY, and ONNX model paths trigger ONNX_QUANT_HIERARCHY. This mapping occurs during device detection in the fit-analysis initialization.
How does the quality penalty affect model selection?
The quality penalty feeds into a composite fit score where higher penalties reduce a quantization format's ranking. According to the llmfit source code, Q8_0 carries zero penalty as the reference, while aggressively compressed formats like Q2_K incur a -20 penalty that typically requires significant memory constraints to justify selection.
Can I force a specific quantization level instead of automatic selection?
Yes. While fit.rs provides automatic selection, you can bypass the hierarchy iteration by directly specifying a quantization string. The helper functions will still validate the format and return appropriate metrics for your custom choice, though manual selection bypasses the quality penalty scoring system.
Why does MLX quantization use different penalty values than core formats?
Apple Silicon's unified memory architecture and dedicated neural engine optimizations reduce the quality degradation typically associated with 4-bit quantization. The llmfit source code assigns mlx-4bit a -4 penalty compared to Q4_K_M's -5, reflecting measurable output quality improvements on MLX-compatible hardware despite similar compression ratios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →