# Soup CLI Quantization Methods: Complete Guide to GPTQ, AWQ, HQQ, and More

> Explore Soup CLI's ten quantization methods including GPTQ, AWQ, HQQ, and more. Optimize your models efficiently with this complete guide to supported quantization techniques.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Soup CLI supports ten distinct quantization formats including GPTQ, AWQ, HQQ, AQLM, EETQ, MXFP4, FP8, BNB 4-bit (Unsloth), Torch-AO schemes, and vLLM runtime quantizations, all configurable via the `--quantization` flag or `--auto-quant` option.**

The Soup CLI from the MakazhanAlpamys/Soup repository provides comprehensive support for quantized large language models. Whether you are fine-tuning pre-quantized checkpoints or serving models with reduced precision, understanding which quantization methods supported by Soup CLI are available is essential for optimizing memory usage and inference speed.

## Core Quantization Registry in [`quant_menu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/quant_menu.py)

The central registry for training-time quantization formats resides in [`src/soup_cli/utils/quant_menu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/quant_menu.py). This module defines the **`PREQUANTIZED_FORMATS`** constant and helper functions like **`is_quant_menu_format`** that validate quantization strings passed to the CLI.

### Supported Pre-Quantized Formats

According to the source code, the following methods are implemented via dedicated builder functions in [`quant_menu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/quant_menu.py):

- **GPTQ**: Implemented via **`build_gptq_config()`**, expects a [`quantize_config.json`](https://github.com/MakazhanAlpamys/Soup/blob/main/quantize_config.json) file. Compatible with LoRA fine-tuning on pre-quantized checkpoints.
- **AWQ**: Handled by **`build_awq_config()`**, looks for [`quant_config.json`](https://github.com/MakazhanAlpamys/Soup/blob/main/quant_config.json) in the model directory.
- **HQQ**: Supports flexible bit widths through **`build_hqq_config()`** and **`parse_hqq_bits()`**; accepts strings like `hqq:4` or `hqq:8`.
- **AQLM**: Available via **`build_aqlm_config()`**, invoked with `quantization="aqlm"`.
- **EETQ**: Enabled through **`build_eetq_config()`** for efficient inference.
- **MXFP4**: Managed by **`build_mxfp4_config()`**, handling the MXFP4 scheme as a variant of FP8.

## Runtime and Inference Quantization

### FP8 De-Quantization

For inference-time loading of FP8 models, [`src/soup_cli/utils/v028_features.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/v028_features.py) detects `quantization_aware="fp8"` and automatically de-quantizes models stored in FP8 to FP16/32 for computation.

### vLLM Auto-Quantization

The [`src/soup_cli/utils/vllm.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/vllm.py) module supports runtime quantization through the `--auto-quant` flag, accepting the strings **`awq`**, **`gptq`**, or **`fp8`** when launching a vLLM engine. If automatic detection fails, [`src/soup_cli/utils/auto_quant.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/auto_quant.py) provides fallback logic to handle incompatible configurations.

## Export and Specialized Formats

### BNB 4-bit via Unsloth

In [`src/soup_cli/utils/unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/unsloth.py), the wrapper automatically applies 4-bit BNB quantization when `quantization="4bit"` is specified, managing double-quant flags and memory optimization automatically.

### Torch-AO Export Workflows

The [`src/soup_cli/utils/save_formats.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/save_formats.py) module provides **`torch_ao_quantize`** workflows supporting schemes like **`Int4WeightOnly`** and **`NVFP4`** via YAML configuration files passed with `--quant-config`.

## Practical Usage Examples

The following commands demonstrate how to exercise these quantization methods supported by Soup CLI in real workflows:

### Export a GPTQ Model with Torch-AO

```bash
soup export \
    --model my-org/gptq-model \
    --format torchao \
    --quant-config gptq_config.yaml

```

*The YAML configuration is built using `build_gptq_config()` as defined in [`quant_menu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/quant_menu.py).*

### Serve a Model with vLLM Auto-Quant

```bash
soup serve \
    --model my-org/awq-model \
    --auto-quant

```

*Behind the scenes, [`src/soup_cli/utils/vllm.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/vllm.py) validates that the chosen method is one of `awq`, `gptq`, or `fp8`.*

### Fine-Tune a Pre-Quantized HQQ Checkpoint

```bash
soup train \
    --model my-org/hqq-8bit \
    --quantization hqq:8 \
    --lora-r 8

```

*The `hqq:8` string is parsed by `parse_hqq_bits()` in [`quant_menu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/quant_menu.py) to configure the 8-bit HQQ loader.*

### Export a 4-bit BNB Checkpoint with Unsloth

```bash
soup export \
    --model my-org/unsloth-model \
    --quantization 4bit \
    --format merged_4bit

```

*[`unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/unsloth.py) forces a 4-bit BNB quantization and handles double-quant settings automatically.*

## Summary

- Soup CLI supports **ten quantization formats** ranging from GPTQ and AWQ to specialized formats like MXFP4 and Torch-AO schemes.
- The central registry in [`src/soup_cli/utils/quant_menu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/quant_menu.py) defines **`PREQUANTIZED_FORMATS`** and validation helpers like **`is_quant_menu_format()`**.
- Use `--quantization` with specific format strings (e.g., `hqq:8`, `aqlm`, `4bit`) for training and export, or `--auto-quant` with vLLM for runtime serving.
- Export workflows support BNB 4-bit via Unsloth ([`unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/unsloth.py)) and Torch-AO via YAML configurations in [`save_formats.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/save_formats.py).
- FP8 inference support is handled by detection logic in [`v028_features.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/v028_features.py), while [`auto_quant.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/auto_quant.py) provides fallback handling for vLLM.

## Frequently Asked Questions

### Does Soup CLI support fine-tuning on pre-quantized models?

Yes. According to the source code in [`src/soup_cli/utils/quant_menu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/quant_menu.py), you can fine-tune pre-quantized checkpoints using LoRA with GPTQ, AWQ, and HQQ formats. The `build_gptq_config()` and `build_awq_config()` functions specifically handle the loading of [`quantize_config.json`](https://github.com/MakazhanAlpamys/Soup/blob/main/quantize_config.json) and [`quant_config.json`](https://github.com/MakazhanAlpamys/Soup/blob/main/quant_config.json) files, allowing parameter-efficient fine-tuning without full precision reconstruction.

### What is the difference between `--quantization` and `--auto-quant` flags?

The `--quantization` flag specifies the exact format when loading or exporting models (e.g., `hqq:4`, `gptq`, `aqlm`), as parsed by functions in [`quant_menu.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/quant_menu.py). The `--auto-quant` flag, implemented in [`src/soup_cli/utils/vllm.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/vllm.py), automatically detects and applies the appropriate runtime quantization (`awq`, `gptq`, or `fp8`) when serving models with the vLLM engine, simplifying deployment decisions when the exact format is already embedded in the model metadata.

### How do I export a model using Torch-AO quantization schemes?

Use the `--format torchao` flag combined with a `--quant-config` YAML file. The `torch_ao_quantize` workflow in [`src/soup_cli/utils/save_formats.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/save_formats.py) supports various schemes including `Int4WeightOnly` and `NVFP4`. This allows exporting to compressed formats that maintain hardware-specific optimizations for inference acceleration on modern accelerators.

### Which quantization method should I use for 4-bit inference with Unsloth?

Specify `--quantization 4bit` when working with Unsloth-compatible models. The [`src/soup_cli/utils/unsloth.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/unsloth.py) module automatically applies BNB 4-bit quantization and manages double-quantization settings internally, making it the recommended path for Unsloth-based 4-bit model export and training workflows.