# How LLMFIT Validates CUDA Compute Capability for Pre-Quantized Models (AWQ/GPTQ)

> LLMFIT validates pre-quantized AWQ/GPTQ model compatibility by checking CUDA compute capability, ensuring seamless inference. Discover how LLMFIT ensures GPU readiness.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: how-to-guide
- Published: 2026-08-21

---

**LLMFIT validates pre-quantized model compatibility by comparing the detected NVIDIA GPU compute capability against the minimum required compute capability for AWQ and GPTQ formats, filtering out incompatible models before inference begins.**

The `backend_compatible` function in the AlexsJones/llmfit repository serves as the gatekeeper for CUDA compute capability validation. When loading AWQ-4bit, AWQ-8bit, GPTQ-Int4, or GPTQ-Int8 models, LLMFIT ensures the underlying hardware supports the Tensor Core instructions introduced with NVIDIA's Turing architecture.

## Detecting Pre-Quantized Models

The compatibility check begins in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) by determining whether a model requires quantization-specific hardware features.

The function first invokes `model.is_prequantized()` to identify if the model uses pre-quantized weights. This check distinguishes standard floating-point models from those requiring specific integer tensor core support.

If the model is not pre-quantized, LLMFIT immediately returns compatibility without inspecting GPU capabilities, as standard CUDA operations function across all compute capabilities.

## Mapping Quantization Formats to Minimum Requirements

For pre-quantized models, LLMFIT queries `quant_min_compute_capability` in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) (lines 89-93) to retrieve the minimum compute capability tuple.

**AWQ and GPTQ** formats require compute capability **7.5** (Turing or newer). Specifically:

- **AWQ-4bit** and **AWQ-8bit** → `(7, 5)`
- **GPTQ-Int4** and **GPTQ-Int8** → `(7, 5)`

Other quantization formats return `None`, indicating no specific compute capability requirement.

## Resolving GPU Compute Capability

After determining the minimum requirement, LLMFIT resolves the actual hardware capability using `gpu_compute_capability` in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) (starting at line 95).

This function parses the GPU model string or PCI ID and maps it to the corresponding compute capability tuple. The implementation includes comprehensive tables covering:

- **Blackwell** and **Hopper** architectures
- **Ada Lovelace** (8.9)
- **Ampere** (8.0, 8.6)
- **Turing** (7.5)
- **Volta** and **Pascal** (older generations)

For example, an "NVIDIA RTX 3090" resolves to `(8, 6)`, while an "NVIDIA GTX 1080" resolves to `(6, 1)`.

## The Compatibility Comparison Logic

The final validation occurs in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) (lines 70-75) with a direct tuple comparison:

```rust
if sys.backend != GpuBackend::Cuda {
    return true; // ROCm assumes compatibility
}
let min_cc = quant_min_compute_capability(&model.quantization);
let gpu_cc = sys.gpu_name.as_ref().and_then(|n| gpu_compute_capability(n));

match (min_cc, gpu_cc) {
    (Some(min), Some(gpu)) => gpu >= min,
    _ => true, // Unknown GPU or non-CUDA backend
}

```

**The comparison `gpu_cc >= min_cc`** ensures the detected GPU meets or exceeds the Turing architecture requirement. If the GPU is older than compute capability 7.5, the pre-quantized model is filtered from available options.

## Practical Implementation Examples

### Manual Compatibility Checking

You can replicate the validation logic in your own Rust code:

```rust
use llmfit_core::{
    hardware::{gpu_compute_capability, quant_min_compute_capability},
    models::LlmModel,
    SystemSpecs, GpuBackend,
};

fn validate_gpu_support(model: &LlmModel, sys: &SystemSpecs) -> bool {
    if !model.is_prequantized() {
        return true;
    }
    
    if sys.backend != GpuBackend::Cuda {
        return true; // ROCm or other backends
    }
    
    let min_cc = quant_min_compute_capability(&model.quantization);
    let gpu_cc = sys.gpu_name.as_ref()
        .and_then(|name| gpu_compute_capability(name));
    
    match (min_cc, gpu_cc) {
        (Some(min), Some(gpu)) => gpu >= min,
        _ => true,
    }
}

```

### CLI Integration

When using the LLMFIT CLI, this validation runs automatically:

```bash
cargo run -- fit --budget 8GB

```

The command constructs a `SystemSpecs` object, loads the model catalog, and applies `backend_compatible` to each entry. Incompatible AWQ or GPTQ models on pre-Turing GPUs are silently omitted from results.

### Extending for New Formats

To add support for a new quantization format requiring compute capability 8.6 (Ampere):

```rust
// In llmfit-core/src/hardware.rs
pub fn quant_min_compute_capability(quant: &Quantization) -> Option<(u32, u32)> {
    match quant.format.as_str() {
        "AWQ-4bit" | "AWQ-8bit" | "GPTQ-Int4" | "GPTQ-Int8" => Some((7, 5)),
        "NEW-FORMAT" => Some((8, 6)), // Ampere requirement
        _ => None,
    }
}

```

The existing `backend_compatible` logic automatically enforces the new constraint.

## Summary

- **Primary gate**: The `backend_compatible` function in [`llmfit-core/src/fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/fit.rs) filters pre-quantized models based on CUDA compute capability.
- **Hardware minimum**: AWQ and GPTQ formats require compute capability 7.5 (Turing), defined in [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) lines 89-93.
- **GPU detection**: `gpu_compute_capability` parses device names into capability tuples using comprehensive architecture tables.
- **Validation logic**: Simple tuple comparison `gpu_cc >= min_cc` at lines 70-75 of [`fit.rs`](https://github.com/AlexsJones/llmfit/blob/main/fit.rs) determines compatibility.
- **Fallback behavior**: Non-CUDA backends (ROCm) and unknown GPUs default to compatible to avoid false negatives.

## Frequently Asked Questions

### What is the minimum NVIDIA compute capability for running AWQ models in LLMFIT?

AWQ-4bit and AWQ-8bit models require CUDA compute capability **7.5**, corresponding to NVIDIA Turing architecture (RTX 20 series, GTX 16 series, or newer). This requirement exists because AWQ relies on Tensor Core integer operations introduced in Turing.

### How does LLMFIT handle pre-quantized models on AMD ROCm GPUs?

LLMFIT assumes compatibility for pre-quantized models when the system uses `GpuBackend::Rocm` or any non-CUDA backend. The `backend_compatible` function returns `true` immediately for these cases because ROCm does not expose a compute capability number equivalent to NVIDIA's scheme.

### Can I manually check if my GPU supports a specific GPTQ model before loading?

Yes. Import `gpu_compute_capability` from [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) and `quant_min_compute_capability` to compare capabilities manually. Pass your GPU name string to `gpu_compute_capability` and compare the resulting tuple against the quantization format's minimum requirement using standard ordering operators (`>=`).

### Where does LLMFIT define the compute capability requirements for different quantization formats?

The requirements are defined in the `quant_min_compute_capability` function within [`llmfit-core/src/hardware.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/hardware.rs) at lines 89-93. Current mappings include AWQ and GPTQ variants requiring `(7, 5)`, with additional formats added by extending this match statement.