How llmfit Performs Dynamic Quantization Selection for LLM Deployment

llmfit automatically selects the optimal model quantization at runtime by scanning ranked hierarchies of compression formats and choosing the first option that fits within available hardware memory budgets.

AlexsJones/llmfit is a Rust-based framework that analyzes large language model compatibility with local hardware. Its dynamic quantization selection system eliminates manual configuration by programmatically evaluating memory requirements across multiple quantization levels and matching them to GPU constraints.

Quantization Hierarchies by Runtime

The foundation of llmfit's selection process resides in three ordered constants defined in llmfit-core/src/models.rs. These hierarchies rank quantization formats from highest fidelity to most aggressive compression:

  • General GGUF hierarchy (QUANT_HIERARCHY): ["Q8_0","Q6_K","Q5_K_M","Q4_K_M","Q3_K_M","Q2_K"]
  • MLX-specific hierarchy (MLX_QUANT_HIERARCHY): ["mlx-8bit","mlx-4bit"] for Apple Silicon runtimes
  • ONNX hierarchy (ONNX_QUANT_HIERARCHY): ["Q8_0","Q4_0"]

Each hierarchy represents the complete spectrum of quantization options for its respective runtime backend.

Memory Estimation and Budget Constraints

Before scanning hierarchies, llmfit calculates precise memory requirements using LlmModel::estimate_memory_gb. This method aggregates model weights, KV-cache overhead, and runtime buffers to produce a per-quantization memory footprint.

The estimation logic calls kv_cache_gb to account for KV-cache quantization levels. When detailed metadata is available, the function applies a per-layer formula to generate accurate size predictions rather than rough approximations.

The Selection Algorithm

The core selection logic lives in best_quant_for_budget and its flexible variant best_quant_for_budget_with, both implemented in llmfit-core/src/models.rs. The algorithm executes the following steps:

  1. Iterate the hierarchy: Scan the quantization list from highest to lowest quality
  2. Estimate memory: Call estimate_memory_gb for each quantization candidate
  3. Compare to budget: Return the first quantization where memory ≤ available budget
  4. Fallback reduction: If no quantization fits, halve the context length once to reduce KV-cache size, then retry the scan

Pre-quantized formats including AWQ, GPTQ, and AutoRound bypass this dynamic selection entirely. When llmfit detects these fixed formats, it preserves the model's existing compression level and skips the hierarchy scan.

Runtime Integration and Hardware Validation

During model analysis in llmfit-core/src/fit.rs, the selection process integrates with the broader fit pipeline. After determining the available memory budget, the system selects the appropriate hierarchy based on the target runtime:

  • ONNX: Uses ONNX_QUANT_HIERARCHY
  • MLX: Uses MLX_QUANT_HIERARCHY with fallback to the generic GGUF hierarchy if selection fails
  • Generic GGUF: Uses QUANT_HIERARCHY

The chosen quantization persists in ModelFit.best_quant for downstream speed estimation and UI reporting.

For GPU-based deployments, hardware.rs validates selections against compute capability constraints via hardware::quant_min_compute_capability. This ensures the selected quantization format is executable on the target hardware architecture.

Programmatic Usage Examples

Select the best quantization for a 12 GB VRAM budget with 4K context:

use llmfit_core::models::LlmModel;
use llmfit_core::models::QUANT_HIERARCHY;

// Assume `model` is a loaded LlmModel
let budget_gb = 12.0;
let ctx = 4096;
if let Some((quant, mem)) = model.best_quant_for_budget_with(budget_gb, ctx, QUANT_HIERARCHY) {
    println!("Best quant: {} (needs {:.2} GB)", quant, mem);
} else {
    println!("No quant fits the budget");
}

Command-line usage automatically triggers selection:

cargo run -- fit --model "Llama-3.1-8B"

# Output includes:

# Best quantization for hardware: Q4_K_M (model default: Q8_0)

Retrieve the selected quantization after analysis:

let model_fit = llmfit_core::fit::analyze_model(&model, &system, &config);
println!("Chosen quantization: {}", model_fit.best_quant);

Summary

  • Hierarchy-based ranking: llmfit maintains runtime-specific ordered lists (GGUF, MLX, ONNX) that prioritize quality before compression
  • Accurate memory prediction: The estimate_memory_gb method calculates total VRAM requirements including KV-cache overhead using per-layer formulas when available
  • Greedy selection: best_quant_for_budget returns the first quantization in the hierarchy that satisfies the memory budget, with optional context-length reduction as a fallback
  • Format awareness: Pre-quantized models (AWQ, GPTQ, AutoRound) skip dynamic selection to preserve their fixed compression
  • Hardware validation: GPU compute capabilities are verified in hardware.rs before finalizing the selection

Frequently Asked Questions

How does llmfit handle models that are already quantized?

When analyzing pre-quantized formats like AWQ, GPTQ, or AutoRound, llmfit detects the fixed quantization level in llmfit-core/src/fit.rs and bypasses the dynamic selection algorithm entirely. The system respects the model's existing compression and validates only that the fixed memory footprint fits within the available budget.

What happens if no quantization level fits the memory budget?

If best_quant_for_budget exhausts the hierarchy without finding a suitable quantization, llmfit performs a single context-length reduction (halving the requested context window) to decrease KV-cache requirements. It then rescans the hierarchy once. If the budget still cannot accommodate any quantization, the function returns None and the fit analysis marks the model as incompatible with the hardware.

Can I use a custom quantization hierarchy instead of the defaults?

Yes. While best_quant_for_budget uses the default hierarchies, the best_quant_for_budget_with method accepts a custom hierarchy slice as its third parameter. This allows you to specify exact quantization priorities or limit the search to specific formats by passing a custom array such as &["Q6_K", "Q5_K_M"].

How does llmfit differentiate between MLX and standard GGUF runtimes?

The framework checks the target runtime type during the fit analysis phase in llmfit-core/src/fit.rs. For MLX Apple Silicon targets, it selects MLX_QUANT_HIERARCHY containing ["mlx-8bit","mlx-4bit"]. If the MLX-specific selection fails, it automatically falls back to the standard GGUF hierarchy to maximize compatibility options.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →