Soup CLI Quantization Methods: Complete Guide to GPTQ, AWQ, HQQ, and More

Soup CLI supports ten distinct quantization formats including GPTQ, AWQ, HQQ, AQLM, EETQ, MXFP4, FP8, BNB 4-bit (Unsloth), Torch-AO schemes, and vLLM runtime quantizations, all configurable via the --quantization flag or --auto-quant option.

The Soup CLI from the MakazhanAlpamys/Soup repository provides comprehensive support for quantized large language models. Whether you are fine-tuning pre-quantized checkpoints or serving models with reduced precision, understanding which quantization methods supported by Soup CLI are available is essential for optimizing memory usage and inference speed.

Core Quantization Registry in quant_menu.py

The central registry for training-time quantization formats resides in src/soup_cli/utils/quant_menu.py. This module defines the PREQUANTIZED_FORMATS constant and helper functions like is_quant_menu_format that validate quantization strings passed to the CLI.

Supported Pre-Quantized Formats

According to the source code, the following methods are implemented via dedicated builder functions in quant_menu.py:

  • GPTQ: Implemented via build_gptq_config(), expects a quantize_config.json file. Compatible with LoRA fine-tuning on pre-quantized checkpoints.
  • AWQ: Handled by build_awq_config(), looks for quant_config.json in the model directory.
  • HQQ: Supports flexible bit widths through build_hqq_config() and parse_hqq_bits(); accepts strings like hqq:4 or hqq:8.
  • AQLM: Available via build_aqlm_config(), invoked with quantization="aqlm".
  • EETQ: Enabled through build_eetq_config() for efficient inference.
  • MXFP4: Managed by build_mxfp4_config(), handling the MXFP4 scheme as a variant of FP8.

Runtime and Inference Quantization

FP8 De-Quantization

For inference-time loading of FP8 models, src/soup_cli/utils/v028_features.py detects quantization_aware="fp8" and automatically de-quantizes models stored in FP8 to FP16/32 for computation.

vLLM Auto-Quantization

The src/soup_cli/utils/vllm.py module supports runtime quantization through the --auto-quant flag, accepting the strings awq, gptq, or fp8 when launching a vLLM engine. If automatic detection fails, src/soup_cli/utils/auto_quant.py provides fallback logic to handle incompatible configurations.

Export and Specialized Formats

BNB 4-bit via Unsloth

In src/soup_cli/utils/unsloth.py, the wrapper automatically applies 4-bit BNB quantization when quantization="4bit" is specified, managing double-quant flags and memory optimization automatically.

Torch-AO Export Workflows

The src/soup_cli/utils/save_formats.py module provides torch_ao_quantize workflows supporting schemes like Int4WeightOnly and NVFP4 via YAML configuration files passed with --quant-config.

Practical Usage Examples

The following commands demonstrate how to exercise these quantization methods supported by Soup CLI in real workflows:

Export a GPTQ Model with Torch-AO

soup export \
    --model my-org/gptq-model \
    --format torchao \
    --quant-config gptq_config.yaml

The YAML configuration is built using build_gptq_config() as defined in quant_menu.py.

Serve a Model with vLLM Auto-Quant

soup serve \
    --model my-org/awq-model \
    --auto-quant

Behind the scenes, src/soup_cli/utils/vllm.py validates that the chosen method is one of awq, gptq, or fp8.

Fine-Tune a Pre-Quantized HQQ Checkpoint

soup train \
    --model my-org/hqq-8bit \
    --quantization hqq:8 \
    --lora-r 8

The hqq:8 string is parsed by parse_hqq_bits() in quant_menu.py to configure the 8-bit HQQ loader.

Export a 4-bit BNB Checkpoint with Unsloth

soup export \
    --model my-org/unsloth-model \
    --quantization 4bit \
    --format merged_4bit

unsloth.py forces a 4-bit BNB quantization and handles double-quant settings automatically.

Summary

  • Soup CLI supports ten quantization formats ranging from GPTQ and AWQ to specialized formats like MXFP4 and Torch-AO schemes.
  • The central registry in src/soup_cli/utils/quant_menu.py defines PREQUANTIZED_FORMATS and validation helpers like is_quant_menu_format().
  • Use --quantization with specific format strings (e.g., hqq:8, aqlm, 4bit) for training and export, or --auto-quant with vLLM for runtime serving.
  • Export workflows support BNB 4-bit via Unsloth (unsloth.py) and Torch-AO via YAML configurations in save_formats.py.
  • FP8 inference support is handled by detection logic in v028_features.py, while auto_quant.py provides fallback handling for vLLM.

Frequently Asked Questions

Does Soup CLI support fine-tuning on pre-quantized models?

Yes. According to the source code in src/soup_cli/utils/quant_menu.py, you can fine-tune pre-quantized checkpoints using LoRA with GPTQ, AWQ, and HQQ formats. The build_gptq_config() and build_awq_config() functions specifically handle the loading of quantize_config.json and quant_config.json files, allowing parameter-efficient fine-tuning without full precision reconstruction.

What is the difference between --quantization and --auto-quant flags?

The --quantization flag specifies the exact format when loading or exporting models (e.g., hqq:4, gptq, aqlm), as parsed by functions in quant_menu.py. The --auto-quant flag, implemented in src/soup_cli/utils/vllm.py, automatically detects and applies the appropriate runtime quantization (awq, gptq, or fp8) when serving models with the vLLM engine, simplifying deployment decisions when the exact format is already embedded in the model metadata.

How do I export a model using Torch-AO quantization schemes?

Use the --format torchao flag combined with a --quant-config YAML file. The torch_ao_quantize workflow in src/soup_cli/utils/save_formats.py supports various schemes including Int4WeightOnly and NVFP4. This allows exporting to compressed formats that maintain hardware-specific optimizations for inference acceleration on modern accelerators.

Which quantization method should I use for 4-bit inference with Unsloth?

Specify --quantization 4bit when working with Unsloth-compatible models. The src/soup_cli/utils/unsloth.py module automatically applies BNB 4-bit quantization and manages double-quantization settings internally, making it the recommended path for Unsloth-based 4-bit model export and training workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →