How LLMFIT Validates CUDA Compute Capability for Pre-Quantized Models (AWQ/GPTQ)

LLMFIT validates pre-quantized model compatibility by comparing the detected NVIDIA GPU compute capability against the minimum required compute capability for AWQ and GPTQ formats, filtering out incompatible models before inference begins.

The backend_compatible function in the AlexsJones/llmfit repository serves as the gatekeeper for CUDA compute capability validation. When loading AWQ-4bit, AWQ-8bit, GPTQ-Int4, or GPTQ-Int8 models, LLMFIT ensures the underlying hardware supports the Tensor Core instructions introduced with NVIDIA's Turing architecture.

Detecting Pre-Quantized Models

The compatibility check begins in llmfit-core/src/fit.rs by determining whether a model requires quantization-specific hardware features.

The function first invokes model.is_prequantized() to identify if the model uses pre-quantized weights. This check distinguishes standard floating-point models from those requiring specific integer tensor core support.

If the model is not pre-quantized, LLMFIT immediately returns compatibility without inspecting GPU capabilities, as standard CUDA operations function across all compute capabilities.

Mapping Quantization Formats to Minimum Requirements

For pre-quantized models, LLMFIT queries quant_min_compute_capability in llmfit-core/src/hardware.rs (lines 89-93) to retrieve the minimum compute capability tuple.

AWQ and GPTQ formats require compute capability 7.5 (Turing or newer). Specifically:

  • AWQ-4bit and AWQ-8bit → (7, 5)
  • GPTQ-Int4 and GPTQ-Int8 → (7, 5)

Other quantization formats return None, indicating no specific compute capability requirement.

Resolving GPU Compute Capability

After determining the minimum requirement, LLMFIT resolves the actual hardware capability using gpu_compute_capability in llmfit-core/src/hardware.rs (starting at line 95).

This function parses the GPU model string or PCI ID and maps it to the corresponding compute capability tuple. The implementation includes comprehensive tables covering:

  • Blackwell and Hopper architectures
  • Ada Lovelace (8.9)
  • Ampere (8.0, 8.6)
  • Turing (7.5)
  • Volta and Pascal (older generations)

For example, an "NVIDIA RTX 3090" resolves to (8, 6), while an "NVIDIA GTX 1080" resolves to (6, 1).

The Compatibility Comparison Logic

The final validation occurs in llmfit-core/src/fit.rs (lines 70-75) with a direct tuple comparison:

if sys.backend != GpuBackend::Cuda {
    return true; // ROCm assumes compatibility
}
let min_cc = quant_min_compute_capability(&model.quantization);
let gpu_cc = sys.gpu_name.as_ref().and_then(|n| gpu_compute_capability(n));

match (min_cc, gpu_cc) {
    (Some(min), Some(gpu)) => gpu >= min,
    _ => true, // Unknown GPU or non-CUDA backend
}

The comparison gpu_cc >= min_cc ensures the detected GPU meets or exceeds the Turing architecture requirement. If the GPU is older than compute capability 7.5, the pre-quantized model is filtered from available options.

Practical Implementation Examples

Manual Compatibility Checking

You can replicate the validation logic in your own Rust code:

use llmfit_core::{
    hardware::{gpu_compute_capability, quant_min_compute_capability},
    models::LlmModel,
    SystemSpecs, GpuBackend,
};

fn validate_gpu_support(model: &LlmModel, sys: &SystemSpecs) -> bool {
    if !model.is_prequantized() {
        return true;
    }
    
    if sys.backend != GpuBackend::Cuda {
        return true; // ROCm or other backends
    }
    
    let min_cc = quant_min_compute_capability(&model.quantization);
    let gpu_cc = sys.gpu_name.as_ref()
        .and_then(|name| gpu_compute_capability(name));
    
    match (min_cc, gpu_cc) {
        (Some(min), Some(gpu)) => gpu >= min,
        _ => true,
    }
}

CLI Integration

When using the LLMFIT CLI, this validation runs automatically:

cargo run -- fit --budget 8GB

The command constructs a SystemSpecs object, loads the model catalog, and applies backend_compatible to each entry. Incompatible AWQ or GPTQ models on pre-Turing GPUs are silently omitted from results.

Extending for New Formats

To add support for a new quantization format requiring compute capability 8.6 (Ampere):

// In llmfit-core/src/hardware.rs
pub fn quant_min_compute_capability(quant: &Quantization) -> Option<(u32, u32)> {
    match quant.format.as_str() {
        "AWQ-4bit" | "AWQ-8bit" | "GPTQ-Int4" | "GPTQ-Int8" => Some((7, 5)),
        "NEW-FORMAT" => Some((8, 6)), // Ampere requirement
        _ => None,
    }
}

The existing backend_compatible logic automatically enforces the new constraint.

Summary

  • Primary gate: The backend_compatible function in llmfit-core/src/fit.rs filters pre-quantized models based on CUDA compute capability.
  • Hardware minimum: AWQ and GPTQ formats require compute capability 7.5 (Turing), defined in llmfit-core/src/hardware.rs lines 89-93.
  • GPU detection: gpu_compute_capability parses device names into capability tuples using comprehensive architecture tables.
  • Validation logic: Simple tuple comparison gpu_cc >= min_cc at lines 70-75 of fit.rs determines compatibility.
  • Fallback behavior: Non-CUDA backends (ROCm) and unknown GPUs default to compatible to avoid false negatives.

Frequently Asked Questions

What is the minimum NVIDIA compute capability for running AWQ models in LLMFIT?

AWQ-4bit and AWQ-8bit models require CUDA compute capability 7.5, corresponding to NVIDIA Turing architecture (RTX 20 series, GTX 16 series, or newer). This requirement exists because AWQ relies on Tensor Core integer operations introduced in Turing.

How does LLMFIT handle pre-quantized models on AMD ROCm GPUs?

LLMFIT assumes compatibility for pre-quantized models when the system uses GpuBackend::Rocm or any non-CUDA backend. The backend_compatible function returns true immediately for these cases because ROCm does not expose a compute capability number equivalent to NVIDIA's scheme.

Can I manually check if my GPU supports a specific GPTQ model before loading?

Yes. Import gpu_compute_capability from llmfit-core/src/hardware.rs and quant_min_compute_capability to compare capabilities manually. Pass your GPU name string to gpu_compute_capability and compare the resulting tuple against the quantization format's minimum requirement using standard ordering operators (>=).

Where does LLMFIT define the compute capability requirements for different quantization formats?

The requirements are defined in the quant_min_compute_capability function within llmfit-core/src/hardware.rs at lines 89-93. Current mappings include AWQ and GPTQ variants requiring (7, 5), with additional formats added by extending this match statement.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →