Miles Low‑Precision Formats: MXFP8, NVFP4, FP8, and INT4 QAT Stability Guide
Miles supports MXFP8, NVFP4, blockwise FP8, INT4 Fake‑QAT, and FP8 QAT simulation, with stability ranging from production‑ready (MXFP8/FP8 on Blackwell) to experimental (NVFP4, INT4).
Miles (Model‑in‑the‑Loop Efficient Serving) implements multiple low‑precision numeric formats through dedicated quantizer modules in miles.utils and conversion tools in the tools/ directory. Each format targets specific hardware generations and use cases, with varying stability guarantees documented in the source code and low‑precision training README. This guide examines the implementation details, activation methods, and production readiness of each format based on the radixark/miles repository.
MXFP8 (Matrix‑wise FP8)
Implementation and Hardware Requirements
The MXFP8 format is implemented in miles/utils/mxfp8.py through the MXFP8Quantizer class, which wraps Transformer Engine's MXFP8 quantizer. The code enforces strict dimensional requirements: MXFP8_GROUP_SIZE = 32 must divide the last tensor dimension, with explicit validation at lines 15‑18 raising a clear ValueError if this constraint is violated.
# From miles/utils/mxfp8.py
MXFP8_GROUP_SIZE = 32
def _check_dimensions(tensor):
if tensor.shape[-1] % MXFP8_GROUP_SIZE != 0:
raise ValueError(
f"Last dimension {tensor.shape[-1]} must be divisible by "
f"MXFP8_GROUP_SIZE ({MXFP8_GROUP_SIZE})"
)
Enabling MXFP8 in Training
MXFP8 requires Blackwell/SM100+ GPUs. Activate via command‑line flags and environment variables:
# Convert checkpoint to MXFP8
python tools/convert_hf_to_mxfp8.py \
--hf-checkpoint /path/to/bf16_hf \
--out-dir /path/to/mxfp8_hf
# Run with MXFP8 (automatically selected on Blackwell)
export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1
python scripts/run_qwen3_30b_a3b.py \
--model-dir /path/to/models \
--model-name qwen3-30b-a3b \
--fp8-format e4m3 \
--fp8-recipe blockwise
Stability Guarantee
- Production‑ready on Blackwell hardware
- The low‑precision README states MXFP8 "can achieve more efficient inference throughput and lower training‑inference mismatch, resulting in more stable training" (
examples/infra_features/low_precision/README.mdlines 3‑4) - Dimensional validation prevents silent failures
- No known fatal bugs in current
mainbranch
NVFP4 (NVIDIA FP4 E2M1)
Implementation Details
The NVFP4 format resides in miles/utils/nvfp4.py via the NVFP4Quantizer wrapper. Like MXFP8, it enforces alignment constraints: NVFP4_GROUP_SIZE = 16 must divide weight dimensions, with validation at lines 7‑9.
# From miles/utils/nvfp4.py
NVFP4_GROUP_SIZE = 16
def _validate_weights(weight):
if weight.shape[-1] % NVFP4_GROUP_SIZE != 0:
raise RuntimeError(
f"Weight dimension {weight.shape[-1]} not aligned to "
f"NVFP4_GROUP_SIZE ({NVFP4_GROUP_SIZE})"
)
Activation Pipeline
# Convert to NVFP4
python tools/convert_hf_to_nvfp4.py \
--hf-checkpoint /path/to/bf16_hf \
--out-dir /path/to/nvfp4_hf
# Configure environment and run
export NVTE_NVFP4_4OVER6=all
python scripts/run_qwen3_4b.py \
--model-dir /path/to/models \
--model-name qwen3-4b \
--fp4-format e2m1
Stability Assessment
- Experimental but generally stable
- Extensive test coverage in
tests/fast-gpu/test_nvfp4_quantizer.py(marked "requires Blackwell/B200 CI runner") - Functional on supported hardware with no stability warnings in documentation
- Suitable for MoE models seeking maximum memory reduction
Standard FP8 (Blockwise)
Distinction from MXFP8
While sharing activation flags with MXFP8, standard blockwise FP8 uses pure torch.float8_e4m3fn conversion via tools/convert_hf_to_fp8.py. The format scales per‑block rather than per‑matrix, offering broader hardware compatibility.
Configuration and Execution
# Convert checkpoint
python tools/convert_hf_to_fp8.py \
--hf-checkpoint /path/to/bf16_hf \
--out-dir /path/to/fp8_hf
# Run with blockwise scaling
export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1
python scripts/run_qwen3_4b.py \
--model-dir /path/to/models \
--model-name qwen3-4b \
--fp8-format e4m3 \
--fp8-recipe blockwise
Stability Profile
- Stable for inference and training on supported GPUs
- Same README stability claims as MXFP8 regarding throughput and training‑inference mismatch reduction
- Known limitation: FP8 weights (
--fp8-param-gather) conflict with CPU offloading techniques, requiring the fused Adam optimizer (examples/infra_features/low_precision/README.mdlines 62‑66) - Safe for production deployment when optimizer constraints are respected
INT4 (4‑bit Integer with Fake QAT)
Fake‑QAT Architecture
INT4 in Miles operates through a Fake Quantization‑Aware Training pipeline using straight‑through estimation (STE). The conversion tools tools/convert_hf_to_int4_direct.py and tools/convert_hf_to_hf_int4.py implement NVIDIA's PTQ pipeline with trainable quantization simulation.
Environment and Execution
# Required environment configuration
export OPEN_TRAINING_INT4_FAKE_QAT_FLAG=1
export OPEN_TRAINING_INT4_GROUP_SIZE=128 # Model‑specific: 128 for Qwen‑3‑30B‑A3B
# Convert and run
python tools/convert_hf_to_int4_direct.py \
--input-dir /path/to/bf16_hf \
--output-dir /path/to/int4_hf
python scripts/run_qwen3_30b_a3b.py \
--model-dir /path/to/models \
--model-name qwen3-30b-a3b \
--int4-format
Stability Classification
- Experimentally stable — functional with caveats
- The README explicitly identifies "INT4 STE training" as providing "significant throughput improvement" but does not guarantee full numerical stability
- Documented bugs remain; the same optimizer offloading conflict affecting FP8 weights applies (
README.mdlines 62‑66) - Suitable for experimentation and throughput‑critical scenarios where occasional precision loss is acceptable
FP8 QAT (DeepSeek‑V4 Quantization‑Aware Training)
Simulation‑Based Implementation
The QAT format in Miles is implemented in miles_plugins/models/deepseek_v4/ops/qat.py as a custom autograd function that simulates FP8 quantization without modifying underlying BF16 weights. The DeepSeekV4LinearQATFunc.apply forward/backward pass mimics FP8 behavior during training.
# From miles_plugins/models/deepseek_v4/ops/qat.py
class DeepSeekV4LinearQATFunc(torch.autograd.Function):
@staticmethod
def forward(ctx, input, weight, ...):
# Simulate FP8 quantization without altering BF16 storage
fp8_simulated = simulate_fp8_quant(input, weight)
return fp8_simulated
Automatic Activation
QAT requires no explicit flags. When quantization_config specifies "qat" in a DeepSeek‑V4 checkpoint, the simulation layer inserts automatically:
# Standard FP8 conversion sufficient
python tools/convert_hf_to_fp8.py \
--hf-checkpoint /path/to/deepseek_v4_bf16 \
--out-dir /path/to/deepseek_v4_fp8
# QAT simulates FP8 behavior at inference time
Stability Characteristics
- Inherently stable — simulation preserves original weights
- No runtime exceptions or dimensional constraints
- No impact on training stability (inference‑time only)
- Test suite includes dedicated QAT validation
Format Comparison and Selection Guide
| Format | Module Path | Target Hardware | Stability Level | Best Use Case |
|---|---|---|---|---|
| MXFP8 | miles/utils/mxfp8.py |
Blackwell/SM100+ | Production | Maximum efficiency on latest GPUs |
| FP8 (blockwise) | tools/convert_hf_to_fp8.py |
Ampere+ | Production | Broad compatibility, stable training |
| NVFP4 | miles/utils/nvfp4.py |
Blackwell/B200 | Experimental | Extreme compression for MoE models |
| INT4 QAT | tools/convert_hf_to_int4_direct.py |
Various | Experimental | Throughput‑critical, precision‑tolerant |
| FP8 QAT | miles_plugins/models/deepseek_v4/ops/qat.py |
Any (simulation) | Production | Inference‑time FP8 behavior testing |
Summary
- MXFP8 and FP8 blockwise are the most stable low‑precision formats in Miles, recommended for production training and inference on supported NVIDIA GPUs
- NVFP4 delivers maximum compression but remains experimental; verify on Blackwell/B200 hardware before deployment
- INT4 Fake‑QAT provides substantial throughput gains through STE training, yet exhibits documented instability and optimizer conflicts
- FP8 QAT offers risk‑free FP8 simulation for DeepSeek‑V4 models without modifying base weights
- All formats enforce dimensional alignment through explicit validation in their respective quantizer implementations
Frequently Asked Questions
What GPU hardware is required for Miles MXFP8 support?
MXFP8 requires Blackwell architecture (SM100+) GPUs such as B100 or B200. The format automatically activates on compatible hardware when using --fp8-recipe blockwise with NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1. Older GPUs fall back to standard blockwise FP8.
Why does Miles INT4 training have stability warnings?
INT4 uses Fake QAT with straight‑through estimation, which introduces approximation errors during gradient computation. The README explicitly notes remaining bugs and warns that the technique conflicts with common optimizer configurations like Adam CPU offloading—unlike FP8 formats which maintain full gradient precision.
Can NVFP4 and MXFP8 be used together in the same model?
No—Miles quantizers are mutually exclusive per layer. NVFP4 (--fp4-format e2m1) and MXFP8 (--fp8-format e4m3) target different bit widths and scaling schemes. Model‑wise mixing would require manual layer‑by‑layer configuration not currently supported by the conversion scripts.
What is the optimizer conflict mentioned for FP8 and INT4 weights?
Both FP8 weights (--fp8-param-gather) and INT4 training currently require the fused Adam optimizer implementation. This conflicts with CPU offloading techniques that move optimizer states to host memory. The low‑precision README documents this at lines 62‑66 as a known limitation awaiting resolution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →