Soup CLI Quantization Methods: Complete Guide to GPTQ, AWQ, HQQ, and More
Soup CLI supports ten distinct quantization formats including GPTQ, AWQ, HQQ, AQLM, EETQ, MXFP4, FP8, BNB 4-bit (Unsloth), Torch-AO schemes, and vLLM runtime quantizations, all configurable via the --quantization flag or --auto-quant option.
The Soup CLI from the MakazhanAlpamys/Soup repository provides comprehensive support for quantized large language models. Whether you are fine-tuning pre-quantized checkpoints or serving models with reduced precision, understanding which quantization methods supported by Soup CLI are available is essential for optimizing memory usage and inference speed.
Core Quantization Registry in quant_menu.py
The central registry for training-time quantization formats resides in src/soup_cli/utils/quant_menu.py. This module defines the PREQUANTIZED_FORMATS constant and helper functions like is_quant_menu_format that validate quantization strings passed to the CLI.
Supported Pre-Quantized Formats
According to the source code, the following methods are implemented via dedicated builder functions in quant_menu.py:
- GPTQ: Implemented via
build_gptq_config(), expects aquantize_config.jsonfile. Compatible with LoRA fine-tuning on pre-quantized checkpoints. - AWQ: Handled by
build_awq_config(), looks forquant_config.jsonin the model directory. - HQQ: Supports flexible bit widths through
build_hqq_config()andparse_hqq_bits(); accepts strings likehqq:4orhqq:8. - AQLM: Available via
build_aqlm_config(), invoked withquantization="aqlm". - EETQ: Enabled through
build_eetq_config()for efficient inference. - MXFP4: Managed by
build_mxfp4_config(), handling the MXFP4 scheme as a variant of FP8.
Runtime and Inference Quantization
FP8 De-Quantization
For inference-time loading of FP8 models, src/soup_cli/utils/v028_features.py detects quantization_aware="fp8" and automatically de-quantizes models stored in FP8 to FP16/32 for computation.
vLLM Auto-Quantization
The src/soup_cli/utils/vllm.py module supports runtime quantization through the --auto-quant flag, accepting the strings awq, gptq, or fp8 when launching a vLLM engine. If automatic detection fails, src/soup_cli/utils/auto_quant.py provides fallback logic to handle incompatible configurations.
Export and Specialized Formats
BNB 4-bit via Unsloth
In src/soup_cli/utils/unsloth.py, the wrapper automatically applies 4-bit BNB quantization when quantization="4bit" is specified, managing double-quant flags and memory optimization automatically.
Torch-AO Export Workflows
The src/soup_cli/utils/save_formats.py module provides torch_ao_quantize workflows supporting schemes like Int4WeightOnly and NVFP4 via YAML configuration files passed with --quant-config.
Practical Usage Examples
The following commands demonstrate how to exercise these quantization methods supported by Soup CLI in real workflows:
Export a GPTQ Model with Torch-AO
soup export \
--model my-org/gptq-model \
--format torchao \
--quant-config gptq_config.yaml
The YAML configuration is built using build_gptq_config() as defined in quant_menu.py.
Serve a Model with vLLM Auto-Quant
soup serve \
--model my-org/awq-model \
--auto-quant
Behind the scenes, src/soup_cli/utils/vllm.py validates that the chosen method is one of awq, gptq, or fp8.
Fine-Tune a Pre-Quantized HQQ Checkpoint
soup train \
--model my-org/hqq-8bit \
--quantization hqq:8 \
--lora-r 8
The hqq:8 string is parsed by parse_hqq_bits() in quant_menu.py to configure the 8-bit HQQ loader.
Export a 4-bit BNB Checkpoint with Unsloth
soup export \
--model my-org/unsloth-model \
--quantization 4bit \
--format merged_4bit
unsloth.py forces a 4-bit BNB quantization and handles double-quant settings automatically.
Summary
- Soup CLI supports ten quantization formats ranging from GPTQ and AWQ to specialized formats like MXFP4 and Torch-AO schemes.
- The central registry in
src/soup_cli/utils/quant_menu.pydefinesPREQUANTIZED_FORMATSand validation helpers likeis_quant_menu_format(). - Use
--quantizationwith specific format strings (e.g.,hqq:8,aqlm,4bit) for training and export, or--auto-quantwith vLLM for runtime serving. - Export workflows support BNB 4-bit via Unsloth (
unsloth.py) and Torch-AO via YAML configurations insave_formats.py. - FP8 inference support is handled by detection logic in
v028_features.py, whileauto_quant.pyprovides fallback handling for vLLM.
Frequently Asked Questions
Does Soup CLI support fine-tuning on pre-quantized models?
Yes. According to the source code in src/soup_cli/utils/quant_menu.py, you can fine-tune pre-quantized checkpoints using LoRA with GPTQ, AWQ, and HQQ formats. The build_gptq_config() and build_awq_config() functions specifically handle the loading of quantize_config.json and quant_config.json files, allowing parameter-efficient fine-tuning without full precision reconstruction.
What is the difference between --quantization and --auto-quant flags?
The --quantization flag specifies the exact format when loading or exporting models (e.g., hqq:4, gptq, aqlm), as parsed by functions in quant_menu.py. The --auto-quant flag, implemented in src/soup_cli/utils/vllm.py, automatically detects and applies the appropriate runtime quantization (awq, gptq, or fp8) when serving models with the vLLM engine, simplifying deployment decisions when the exact format is already embedded in the model metadata.
How do I export a model using Torch-AO quantization schemes?
Use the --format torchao flag combined with a --quant-config YAML file. The torch_ao_quantize workflow in src/soup_cli/utils/save_formats.py supports various schemes including Int4WeightOnly and NVFP4. This allows exporting to compressed formats that maintain hardware-specific optimizations for inference acceleration on modern accelerators.
Which quantization method should I use for 4-bit inference with Unsloth?
Specify --quantization 4bit when working with Unsloth-compatible models. The src/soup_cli/utils/unsloth.py module automatically applies BNB 4-bit quantization and manages double-quantization settings internally, making it the recommended path for Unsloth-based 4-bit model export and training workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →