Miles Low‑Precision Formats: MXFP8, NVFP4, FP8, and INT4 QAT Stability Guide

Miles supports MXFP8, NVFP4, blockwise FP8, INT4 Fake‑QAT, and FP8 QAT simulation, with stability ranging from production‑ready (MXFP8/FP8 on Blackwell) to experimental (NVFP4, INT4).

Miles (Model‑in‑the‑Loop Efficient Serving) implements multiple low‑precision numeric formats through dedicated quantizer modules in miles.utils and conversion tools in the tools/ directory. Each format targets specific hardware generations and use cases, with varying stability guarantees documented in the source code and low‑precision training README. This guide examines the implementation details, activation methods, and production readiness of each format based on the radixark/miles repository.

MXFP8 (Matrix‑wise FP8)

Implementation and Hardware Requirements

The MXFP8 format is implemented in miles/utils/mxfp8.py through the MXFP8Quantizer class, which wraps Transformer Engine's MXFP8 quantizer. The code enforces strict dimensional requirements: MXFP8_GROUP_SIZE = 32 must divide the last tensor dimension, with explicit validation at lines 15‑18 raising a clear ValueError if this constraint is violated.


# From miles/utils/mxfp8.py

MXFP8_GROUP_SIZE = 32

def _check_dimensions(tensor):
    if tensor.shape[-1] % MXFP8_GROUP_SIZE != 0:
        raise ValueError(
            f"Last dimension {tensor.shape[-1]} must be divisible by "
            f"MXFP8_GROUP_SIZE ({MXFP8_GROUP_SIZE})"
        )

Enabling MXFP8 in Training

MXFP8 requires Blackwell/SM100+ GPUs. Activate via command‑line flags and environment variables:


# Convert checkpoint to MXFP8

python tools/convert_hf_to_mxfp8.py \
  --hf-checkpoint /path/to/bf16_hf \
  --out-dir      /path/to/mxfp8_hf

# Run with MXFP8 (automatically selected on Blackwell)

export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1
python scripts/run_qwen3_30b_a3b.py \
  --model-dir /path/to/models \
  --model-name qwen3-30b-a3b \
  --fp8-format e4m3 \
  --fp8-recipe blockwise

Stability Guarantee

  • Production‑ready on Blackwell hardware
  • The low‑precision README states MXFP8 "can achieve more efficient inference throughput and lower training‑inference mismatch, resulting in more stable training" (examples/infra_features/low_precision/README.md lines 3‑4)
  • Dimensional validation prevents silent failures
  • No known fatal bugs in current main branch

NVFP4 (NVIDIA FP4 E2M1)

Implementation Details

The NVFP4 format resides in miles/utils/nvfp4.py via the NVFP4Quantizer wrapper. Like MXFP8, it enforces alignment constraints: NVFP4_GROUP_SIZE = 16 must divide weight dimensions, with validation at lines 7‑9.


# From miles/utils/nvfp4.py

NVFP4_GROUP_SIZE = 16

def _validate_weights(weight):
    if weight.shape[-1] % NVFP4_GROUP_SIZE != 0:
        raise RuntimeError(
            f"Weight dimension {weight.shape[-1]} not aligned to "
            f"NVFP4_GROUP_SIZE ({NVFP4_GROUP_SIZE})"
        )

Activation Pipeline


# Convert to NVFP4

python tools/convert_hf_to_nvfp4.py \
  --hf-checkpoint /path/to/bf16_hf \
  --out-dir      /path/to/nvfp4_hf

# Configure environment and run

export NVTE_NVFP4_4OVER6=all
python scripts/run_qwen3_4b.py \
  --model-dir /path/to/models \
  --model-name qwen3-4b \
  --fp4-format e2m1

Stability Assessment

  • Experimental but generally stable
  • Extensive test coverage in tests/fast-gpu/test_nvfp4_quantizer.py (marked "requires Blackwell/B200 CI runner")
  • Functional on supported hardware with no stability warnings in documentation
  • Suitable for MoE models seeking maximum memory reduction

Standard FP8 (Blockwise)

Distinction from MXFP8

While sharing activation flags with MXFP8, standard blockwise FP8 uses pure torch.float8_e4m3fn conversion via tools/convert_hf_to_fp8.py. The format scales per‑block rather than per‑matrix, offering broader hardware compatibility.

Configuration and Execution


# Convert checkpoint

python tools/convert_hf_to_fp8.py \
  --hf-checkpoint /path/to/bf16_hf \
  --out-dir      /path/to/fp8_hf

# Run with blockwise scaling

export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1
python scripts/run_qwen3_4b.py \
  --model-dir /path/to/models \
  --model-name qwen3-4b \
  --fp8-format e4m3 \
  --fp8-recipe blockwise

Stability Profile

  • Stable for inference and training on supported GPUs
  • Same README stability claims as MXFP8 regarding throughput and training‑inference mismatch reduction
  • Known limitation: FP8 weights (--fp8-param-gather) conflict with CPU offloading techniques, requiring the fused Adam optimizer (examples/infra_features/low_precision/README.md lines 62‑66)
  • Safe for production deployment when optimizer constraints are respected

INT4 (4‑bit Integer with Fake QAT)

Fake‑QAT Architecture

INT4 in Miles operates through a Fake Quantization‑Aware Training pipeline using straight‑through estimation (STE). The conversion tools tools/convert_hf_to_int4_direct.py and tools/convert_hf_to_hf_int4.py implement NVIDIA's PTQ pipeline with trainable quantization simulation.

Environment and Execution


# Required environment configuration

export OPEN_TRAINING_INT4_FAKE_QAT_FLAG=1
export OPEN_TRAINING_INT4_GROUP_SIZE=128  # Model‑specific: 128 for Qwen‑3‑30B‑A3B

# Convert and run

python tools/convert_hf_to_int4_direct.py \
  --input-dir  /path/to/bf16_hf \
  --output-dir /path/to/int4_hf

python scripts/run_qwen3_30b_a3b.py \
  --model-dir /path/to/models \
  --model-name qwen3-30b-a3b \
  --int4-format

Stability Classification

  • Experimentally stable — functional with caveats
  • The README explicitly identifies "INT4 STE training" as providing "significant throughput improvement" but does not guarantee full numerical stability
  • Documented bugs remain; the same optimizer offloading conflict affecting FP8 weights applies (README.md lines 62‑66)
  • Suitable for experimentation and throughput‑critical scenarios where occasional precision loss is acceptable

FP8 QAT (DeepSeek‑V4 Quantization‑Aware Training)

Simulation‑Based Implementation

The QAT format in Miles is implemented in miles_plugins/models/deepseek_v4/ops/qat.py as a custom autograd function that simulates FP8 quantization without modifying underlying BF16 weights. The DeepSeekV4LinearQATFunc.apply forward/backward pass mimics FP8 behavior during training.


# From miles_plugins/models/deepseek_v4/ops/qat.py

class DeepSeekV4LinearQATFunc(torch.autograd.Function):
    @staticmethod
    def forward(ctx, input, weight, ...):
        # Simulate FP8 quantization without altering BF16 storage

        fp8_simulated = simulate_fp8_quant(input, weight)
        return fp8_simulated

Automatic Activation

QAT requires no explicit flags. When quantization_config specifies "qat" in a DeepSeek‑V4 checkpoint, the simulation layer inserts automatically:


# Standard FP8 conversion sufficient

python tools/convert_hf_to_fp8.py \
  --hf-checkpoint /path/to/deepseek_v4_bf16 \
  --out-dir      /path/to/deepseek_v4_fp8

# QAT simulates FP8 behavior at inference time

Stability Characteristics

  • Inherently stable — simulation preserves original weights
  • No runtime exceptions or dimensional constraints
  • No impact on training stability (inference‑time only)
  • Test suite includes dedicated QAT validation

Format Comparison and Selection Guide

Format Module Path Target Hardware Stability Level Best Use Case
MXFP8 miles/utils/mxfp8.py Blackwell/SM100+ Production Maximum efficiency on latest GPUs
FP8 (blockwise) tools/convert_hf_to_fp8.py Ampere+ Production Broad compatibility, stable training
NVFP4 miles/utils/nvfp4.py Blackwell/B200 Experimental Extreme compression for MoE models
INT4 QAT tools/convert_hf_to_int4_direct.py Various Experimental Throughput‑critical, precision‑tolerant
FP8 QAT miles_plugins/models/deepseek_v4/ops/qat.py Any (simulation) Production Inference‑time FP8 behavior testing

Summary

  • MXFP8 and FP8 blockwise are the most stable low‑precision formats in Miles, recommended for production training and inference on supported NVIDIA GPUs
  • NVFP4 delivers maximum compression but remains experimental; verify on Blackwell/B200 hardware before deployment
  • INT4 Fake‑QAT provides substantial throughput gains through STE training, yet exhibits documented instability and optimizer conflicts
  • FP8 QAT offers risk‑free FP8 simulation for DeepSeek‑V4 models without modifying base weights
  • All formats enforce dimensional alignment through explicit validation in their respective quantizer implementations

Frequently Asked Questions

What GPU hardware is required for Miles MXFP8 support?

MXFP8 requires Blackwell architecture (SM100+) GPUs such as B100 or B200. The format automatically activates on compatible hardware when using --fp8-recipe blockwise with NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1. Older GPUs fall back to standard blockwise FP8.

Why does Miles INT4 training have stability warnings?

INT4 uses Fake QAT with straight‑through estimation, which introduces approximation errors during gradient computation. The README explicitly notes remaining bugs and warns that the technique conflicts with common optimizer configurations like Adam CPU offloading—unlike FP8 formats which maintain full gradient precision.

Can NVFP4 and MXFP8 be used together in the same model?

No—Miles quantizers are mutually exclusive per layer. NVFP4 (--fp4-format e2m1) and MXFP8 (--fp8-format e4m3) target different bit widths and scaling schemes. Model‑wise mixing would require manual layer‑by‑layer configuration not currently supported by the conversion scripts.

What is the optimizer conflict mentioned for FP8 and INT4 weights?

Both FP8 weights (--fp8-param-gather) and INT4 training currently require the fused Adam optimizer implementation. This conflicts with CPU offloading techniques that move optimizer states to host memory. The low‑precision README documents this at lines 62‑66 as a known limitation awaiting resolution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →