Can llmfit Analyze Models with Different Quantization Levels?

Yes, llmfit analyzes models with different quantization levels by evaluating memory footprints, inference speed, and quality degradation for every supported compression format from full-precision down to 2-bit representations.

AlexsJones/llmfit is a Rust-based tool that determines how well large language models fit your hardware constraints. Unlike tools that treat model compression as a static property, llmfit's core engine is quantization-aware throughout the entire fit-analysis pipeline, allowing it to compare feasibility across Q8_0, Q4_K_M, mlx-4bit, and other formats algorithmically.

Quantization Hierarchy and Metric Conversion

The foundation of llmfit's analysis capability rests in llmfit-core/src/models.rs, where the engine maintains a comprehensive understanding of compression trade-offs.

The Quality Spectrum

At lines 3-6 of models.rs, llmfit defines a static quantization hierarchy ordered from highest fidelity to maximum compression. This ordered list enables the system to iterate through quality levels when searching for feasible configurations.

Converting Labels to Concrete Metrics

Between lines 13-61, three critical helper functions translate quantization labels into numerical values used for calculation:

  • quant_bpp: Returns bytes-per-parameter for formats like Q4_K_M or mlx-4bit
  • quant_speed_multiplier: Estimates inference speed relative to full precision
  • quant_quality_penalty: Calculates the accuracy degradation coefficient applied to fit scores

These conversions allow llmfit to compare disparate quantization schemes using standardized units of memory, throughput, and quality.

Memory Estimation and Automatic Selection

Once metrics are normalized, llmfit calculates hardware requirements and selects optimal compression levels based on user constraints.

Calculating Memory Requirements

The estimate_memory_gb function (lines 777-785) and its KV-cache-aware variant compute total RAM or VRAM consumption by combining the quantization's bytes-per-parameter with the model's parameter count and requested context length. This accounts for both weight storage and activation memory.

Budget-Conscious Quantization Selection

The best_quant_for_budget function (lines 55-86) implements automatic selection logic. It walks the quantization hierarchy from highest to lowest quality, testing each level against the supplied memory budget using estimate_memory_gb, and returns the first quantization that fits along with its estimated usage.

This logic is exposed through the CLI via the --budget flag and through the TUI's quantization filter interface.

Quantization Impact on Fit Analysis

During fit analysis in llmfit-core/src/fit.rs, the selected quantization directly influences three critical dimensions:

  • Model Weight Size: Determines memory bandwidth requirements and loading constraints
  • Speed Multiplier: Affects throughput estimates and latency calculations
  • Quality Penalty: Modifies the final fit score to reflect the trade-off between compression and output fidelity

The Plan struct in llmfit-core/src/plan.rs (lines 122-149) stores the selected quantization throughout this pipeline, ensuring downstream calculations remain consistent with the chosen compression level.

Practical Usage Examples

Command-Line Interface

Request automatic selection for an 8 GB budget:

llmfit fit --budget 8 --context 4096

Force a specific quantization level:

llmfit fit --model "Meta-Llama-3.1-8B" --quantization Q4_K_M --context 4096

The argument parsing for these flags is implemented in llmfit-tui/src/main.rs at lines 267-289.

Programmatic Rust API

use llmfit_core::models::LlmModel;

let model: LlmModel = /* load from ModelDatabase */;
let budget_gb = 8.0;
let ctx = 4096;

// Returns ("Q4_K_M", 7.9) – quant name and estimated memory usage
if let Some((quant, mem)) = model.best_quant_for_budget(budget_gb, ctx) {
    println!("Best quant: {quant} (≈{mem:.2} GB)");
}

Interactive TUI

In the terminal interface, press q to activate the quantization filter, then type a format like Q4_K_M. The "Quantization" column rendered at lines 1820-1822 in llmfit-tui/src/tui_ui.rs displays the current compression level for each model in the database.

Summary

  • llmfit maintains a quantization hierarchy in models.rs that orders compression levels by quality
  • Helper functions convert labels like Q4_K_M into concrete bytes-per-parameter, speed multipliers, and quality penalties
  • estimate_memory_gb calculates exact RAM/VRAM requirements for any quantization and context length combination
  • best_quant_for_budget automatically selects the highest-quality compression that fits your hardware constraints
  • The fit analysis pipeline in fit.rs incorporates quantization into final feasibility scores

Frequently Asked Questions

What quantization formats does llmfit support?

llmfit supports a hierarchy of formats defined in llmfit-core/src/models.rs, including GGUF variants like Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q3_K_S, and MLX-specific formats such as mlx-4bit and mlx-8bit. The system is designed to accommodate new quantization schemes by extending the hierarchy and corresponding helper functions.

How does llmfit calculate memory for quantized models?

The engine uses the estimate_memory_gb function to multiply the model's parameter count by the quantization's bytes-per-parameter (determined via quant_bpp), then adds overhead for the KV cache based on context length. This produces an accurate RAM or VRAM prediction for inference.

Can I force a specific quantization level instead of auto-selection?

Yes. Use the --quantization CLI flag followed by the specific format (e.g., --quantization Q4_K_M) to override automatic selection. Programmatically, you can set the quantization field directly on the Plan struct before running the fit analysis, bypassing the best_quant_for_budget logic.

Does quantization affect the final fit score?

Yes. The quant_quality_penalty function applies a degradation coefficient to the fit score based on the compression level. Higher quantization (more compression) increases the penalty, reflecting the trade-off between resource efficiency and model output quality in the final feasibility assessment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →