# How llmfit Performs Dynamic Quantization Selection for LLM Deployment

> llmfit automatically selects optimal LLM quantization at runtime by scanning compression formats and choosing the first fit for your hardware memory budget.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-08-22

---

**llmfit automatically selects the optimal model quantization at runtime by scanning ranked hierarchies of compression formats and choosing the first option that fits within available hardware memory budgets.**

AlexsJones/llmfit is a Rust-based framework that analyzes large language model compatibility with local hardware. Its **dynamic quantization selection** system eliminates manual configuration by programmatically evaluating memory requirements across multiple quantization levels and matching them to GPU constraints.

## Quantization Hierarchies by Runtime

The foundation of llmfit's selection process resides in three ordered constants defined in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs). These hierarchies rank quantization formats from highest fidelity to most aggressive compression:

- **General GGUF hierarchy** (`QUANT_HIERARCHY`): `["Q8_0","Q6_K","Q5_K_M","Q4_K_M","Q3_K_M","Q2_K"]`
- **MLX-specific hierarchy** (`MLX_QUANT_HIERARCHY`): `["mlx-8bit","mlx-4bit"]` for Apple Silicon runtimes
- **ONNX hierarchy** (`ONNX_QUANT_HIERARCHY`): `["Q8_0","Q4_0"]`

Each hierarchy represents the complete spectrum of quantization options for its respective runtime backend.

## Memory Estimation and Budget Constraints

Before scanning hierarchies, llmfit calculates precise memory requirements using `LlmModel::estimate_memory_gb`. This method aggregates model weights, KV-cache overhead, and runtime buffers to produce a per-quantization memory footprint.

The estimation logic calls `kv_cache_gb` to account for KV-cache quantization levels. When detailed metadata is available, the function applies a per-layer formula to generate accurate size predictions rather than rough approximations.

## The Selection Algorithm

The core selection logic lives in `best_quant_for_budget` and its flexible variant `best_quant_for_budget_with`, both implemented in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs). The algorithm executes the following steps:

1. **Iterate the hierarchy**: Scan the quantization list from highest to lowest quality
2. **Estimate memory**: Call `estimate_memory_gb` for each quantization candidate
3. **Compare to budget**: Return the first quantization where memory ≤ available budget
4. **Fallback reduction**: If no quantization fits, halve the context length once to reduce KV-cache size, then retry the scan

Pre-quantized formats including **AWQ**, **GPTQ**, and **AutoRound** bypass this dynamic selection entirely. When `llmfit` detects these fixed formats, it preserves the model's existing compression level and skips the hierarchy scan.

## Runtime Integration and Hardware Validation

During model analysis in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), the selection process integrates with the broader fit pipeline. After determining the available memory budget, the system selects the appropriate hierarchy based on the target runtime:

- **ONNX**: Uses `ONNX_QUANT_HIERARCHY`
- **MLX**: Uses `MLX_QUANT_HIERARCHY` with fallback to the generic GGUF hierarchy if selection fails
- **Generic GGUF**: Uses `QUANT_HIERARCHY`

The chosen quantization persists in `ModelFit.best_quant` for downstream speed estimation and UI reporting.

For GPU-based deployments, [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) validates selections against compute capability constraints via `hardware::quant_min_compute_capability`. This ensures the selected quantization format is executable on the target hardware architecture.

## Programmatic Usage Examples

Select the best quantization for a 12 GB VRAM budget with 4K context:

```rust
use llmfit_core::models::LlmModel;
use llmfit_core::models::QUANT_HIERARCHY;

// Assume `model` is a loaded LlmModel
let budget_gb = 12.0;
let ctx = 4096;
if let Some((quant, mem)) = model.best_quant_for_budget_with(budget_gb, ctx, QUANT_HIERARCHY) {
    println!("Best quant: {} (needs {:.2} GB)", quant, mem);
} else {
    println!("No quant fits the budget");
}

```

Command-line usage automatically triggers selection:

```bash
cargo run -- fit --model "Llama-3.1-8B"

# Output includes:

# Best quantization for hardware: Q4_K_M (model default: Q8_0)

```

Retrieve the selected quantization after analysis:

```rust
let model_fit = llmfit_core::fit::analyze_model(&model, &system, &config);
println!("Chosen quantization: {}", model_fit.best_quant);

```

## Summary

- **Hierarchy-based ranking**: llmfit maintains runtime-specific ordered lists (GGUF, MLX, ONNX) that prioritize quality before compression
- **Accurate memory prediction**: The `estimate_memory_gb` method calculates total VRAM requirements including KV-cache overhead using per-layer formulas when available
- **Greedy selection**: `best_quant_for_budget` returns the first quantization in the hierarchy that satisfies the memory budget, with optional context-length reduction as a fallback
- **Format awareness**: Pre-quantized models (AWQ, GPTQ, AutoRound) skip dynamic selection to preserve their fixed compression
- **Hardware validation**: GPU compute capabilities are verified in [`hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/hardware.rs) before finalizing the selection

## Frequently Asked Questions

### How does llmfit handle models that are already quantized?

When analyzing pre-quantized formats like AWQ, GPTQ, or AutoRound, llmfit detects the fixed quantization level in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) and bypasses the dynamic selection algorithm entirely. The system respects the model's existing compression and validates only that the fixed memory footprint fits within the available budget.

### What happens if no quantization level fits the memory budget?

If `best_quant_for_budget` exhausts the hierarchy without finding a suitable quantization, llmfit performs a single context-length reduction (halving the requested context window) to decrease KV-cache requirements. It then rescans the hierarchy once. If the budget still cannot accommodate any quantization, the function returns `None` and the fit analysis marks the model as incompatible with the hardware.

### Can I use a custom quantization hierarchy instead of the defaults?

Yes. While `best_quant_for_budget` uses the default hierarchies, the `best_quant_for_budget_with` method accepts a custom hierarchy slice as its third parameter. This allows you to specify exact quantization priorities or limit the search to specific formats by passing a custom array such as `&["Q6_K", "Q5_K_M"]`.

### How does llmfit differentiate between MLX and standard GGUF runtimes?

The framework checks the target runtime type during the fit analysis phase in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs). For MLX Apple Silicon targets, it selects `MLX_QUANT_HIERARCHY` containing `["mlx-8bit","mlx-4bit"]`. If the MLX-specific selection fails, it automatically falls back to the standard GGUF hierarchy to maximize compatibility options.