# Dynamic Quantization Selection in LLMFIT: How It Automatically Chooses Optimal Model Precision

> Discover dynamic quantization selection in LLMFIT. Learn how it automatically finds the best model precision for your hardware, optimizing performance and memory usage.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: deep-dive
- Published: 2026-08-21

---

**Dynamic quantization selection is an automated process that evaluates hardware memory budgets against ordered backend-specific hierarchies to select the highest-precision quantization level that will fit on the target system.**

LLMFIT is an open-source framework designed to optimize large language model deployment by programmatically matching model configurations to available hardware. The project's **dynamic quantization selection** system eliminates manual trial-and-error by analyzing GPU VRAM, system RAM, and context length requirements to determine the optimal quantization automatically.

## How Dynamic Quantization Selection Works

The selection algorithm follows a three-stage pipeline: defining precision hierarchies, calculating memory budgets, and matching runtime-specific constraints.

### Quantization Hierarchies by Backend

LLMFIT maintains ordered arrays of supported quantization formats in [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs), arranged from highest to lowest memory consumption. The order ensures the system prioritizes precision while searching for a viable option:

- **General hierarchy** (Ollama, VLLM, Llama.cpp): `["Q8_0","Q6_K","Q5_K_M","Q4_K_M","Q3_K_M","Q2_K"]` ([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs#L5-L8))
- **MLX-specific**: `["mlx-8bit","mlx-4bit"]` ([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs#L8-L9))
- **ONNX-specific**: `["Q8_0","Q4_0"]` ([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs#L9-L10))

### Budget-Driven Search Strategy

The core selection logic resides in `Model::best_quant_for_budget` within [`llmfit-core/src/models.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs). This method accepts a memory budget in gigabytes and a context length, then delegates to `best_quant_for_budget_with` to traverse the hierarchy:

```rust
pub fn best_quant_for_budget(&self, budget_gb: f64, ctx: u32) -> Option<(&'static str, f64)> {
    self.best_quant_for_budget_with(budget_gb, ctx, QUANT_HIERARCHY)
}

```

The helper function iterates through the supplied hierarchy, calling `estimate_memory_gb` for each candidate until it finds the first quantization where the estimated usage is less than or equal to the budget ([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs#L950-L966)).

### Runtime-Aware Selection Logic

During fit analysis in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs), the system determines which hierarchy applies based on the target inference runtime. The `best_quant_for_runtime_budget` function (around line 922) selects the appropriate constant before invoking the budget check:

```rust
fn best_quant_for_runtime_budget(
    model: &Model,
    runtime: InferenceRuntime,
    budget: f64,
    ctx: u32,
) -> Option<String> {
    let hierarchy = match runtime {
        InferenceRuntime::Mlx => models::MLX_QUANT_HIERARCHY,
        InferenceRuntime::Onnx => models::ONNX_QUANT_HIERARCHY,
        _ => models::QUANT_HIERARCHY,
    };
    model.best_quant_for_budget_with(budget, ctx, hierarchy).map(|(q, _)| q.to_string())
}

```

([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L922-L949))

## Implementation in the Codebase

### Core Selection Logic in [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs)

The [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs) file defines the quantization hierarchies and implements `best_quant_for_budget_with` (lines 950-966). This method performs the actual memory estimation comparison, returning a tuple containing the quantization string and its estimated memory footprint.

### Runtime Integration in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs)

The [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) file orchestrates end-to-end quantization selection during model analysis. When `analyze_with_forced_runtime` executes, it:

1. Calculates the available hardware budget (GPU VRAM for GPU runtimes, system RAM for CPU)
2. Invokes `best_quant_for_runtime_budget` to obtain the candidate quantization string
3. Stores the result in `ModelFit.best_quant` at line 677 ([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L677-L682))

### UI Representation

The terminal UI in [`llmfit-tui/src/tui_ui.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs) displays the selected quantization to users around line 1005, rendering the `best_quant` field from the `ModelFit` structure in the model selection interface ([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs#L1000-L1010)).

## Practical Usage Examples

### Querying Optimal Quantization Directly

You can invoke the selection logic manually to determine which quantization fits a specific hardware configuration:

```rust
use llmfit_core::models::Model;

// Query against an 8 GiB GPU with 4096 token context
let budget_gb = 8.0;
let ctx = 4096;

if let Some((quant, mem_used)) = model.best_quant_for_budget(budget_gb, ctx) {
    println!("Best quant: {} (uses {:.2} GiB)", quant, mem_used);
} else {
    println!("No supported quantization fits within budget");
}

```

([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/models.rs#L950-L966))

### Full Fit Analysis with Auto-Selection

For complete analysis that automatically populates the best quantization:

```rust
use llmfit_core::fit::FitConfig;
use llmfit_core::providers::InferenceRuntime;
use llmfit_core::hardware::SystemSpecs;

let system = SystemSpecs::detect()?;
let fit = llmfit_core::analysis::analyze_model(
    &model,
    &system,
    InferenceRuntime::Vllm,
    FitConfig::default(),
)?;

println!("Auto-selected quant: {}", fit.best_quant);

```

([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs#L677-L682))

### Displaying Selection in Custom Interfaces

When building custom UIs, access the selected quantization through the `ModelFit` structure:

```rust
// Accessing the dynamically selected value
let quant_label = format!("Quant: {}", fit.best_quant);

```

([source](https://github.com/AlexsJones/llmfit/blob/main/llmfit-tui/src/tui_ui.rs#L1000-L1010))

## Summary

- **Dynamic quantization selection** automatically matches model precision to available hardware constraints without user intervention
- **Backend-specific hierarchies** in [`models.rs`](https://github.com/AlexsJones/llmfit/blob/main/models.rs) define the search order, ensuring MLX and ONNX runtimes use compatible quantization formats
- **Budget-driven search** via `best_quant_for_budget` selects the first quantization in the hierarchy that satisfies memory constraints
- **Context-aware calculations** incorporate the requested context length into memory estimates, ensuring sufficient working space remains available
- **Integration points** in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) store results in `ModelFit.best_quant`, while the TUI renders this value for user visibility

## Frequently Asked Questions

### What triggers dynamic quantization selection?

The process triggers automatically during `analyze_model` or `analyze_with_forced_runtime` execution when no specific quantization is manually forced. The system detects available hardware resources and invokes `best_quant_for_runtime_budget` to determine the highest viable precision.

### How does context length affect quantization choice?

The `ctx` parameter feeds into `estimate_memory_gb` calculations within `best_quant_for_budget_with`. Longer context windows increase the memory required for the KV-cache, potentially forcing the selection of lower-precision quantizations (such as Q4_K_M instead of Q8_0) to fit within the same hardware budget.

### Can I override the automatically selected quantization?

While the system defaults to automatic selection, you can bypass the `best_quant_for_budget` logic by manually specifying a quantization format in the model configuration. However, overriding the selection risks out-of-memory errors if the hardware cannot support the chosen precision.

### Which inference runtimes support dynamic selection?

LLMFIT implements specialized hierarchies for **MLX** and **ONNX** runtimes, while using a general quantization hierarchy for **Ollama**, **VLLM**, and other standard backends. The runtime detection logic in [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) ensures the appropriate hierarchy is matched to each target environment.