# Miles Low‑Precision Formats: MXFP8, NVFP4, FP8, and INT4 QAT Stability Guide

> Explore Miles low-precision formats MXFP8 NVFP4 FP8 INT4. Understand stability guarantees from production-ready to experimental for efficient AI.

- Repository: [RadixArk/miles](https://github.com/radixark/miles)
- Tags: guide
- Published: 2026-09-06

---

**Miles supports MXFP8, NVFP4, blockwise FP8, INT4 Fake‑QAT, and FP8 QAT simulation, with stability ranging from production‑ready (MXFP8/FP8 on Blackwell) to experimental (NVFP4, INT4).**

Miles (Model‑in‑the‑Loop Efficient Serving) implements multiple low‑precision numeric formats through dedicated quantizer modules in `miles.utils` and conversion tools in the `tools/` directory. Each format targets specific hardware generations and use cases, with varying stability guarantees documented in the source code and low‑precision training README. This guide examines the implementation details, activation methods, and production readiness of each format based on the radixark/miles repository.

## MXFP8 (Matrix‑wise FP8)

### Implementation and Hardware Requirements

The **MXFP8** format is implemented in [`miles/utils/mxfp8.py`](https://github.com/radixark/miles/blob/main/miles/utils/mxfp8.py) through the `MXFP8Quantizer` class, which wraps Transformer Engine's MXFP8 quantizer. The code enforces strict dimensional requirements: `MXFP8_GROUP_SIZE = 32` must divide the last tensor dimension, with explicit validation at lines 15‑18 raising a clear `ValueError` if this constraint is violated.

```python

# From miles/utils/mxfp8.py

MXFP8_GROUP_SIZE = 32

def _check_dimensions(tensor):
    if tensor.shape[-1] % MXFP8_GROUP_SIZE != 0:
        raise ValueError(
            f"Last dimension {tensor.shape[-1]} must be divisible by "
            f"MXFP8_GROUP_SIZE ({MXFP8_GROUP_SIZE})"
        )

```

### Enabling MXFP8 in Training

MXFP8 requires **Blackwell/SM100+ GPUs**. Activate via command‑line flags and environment variables:

```bash

# Convert checkpoint to MXFP8

python tools/convert_hf_to_mxfp8.py \
  --hf-checkpoint /path/to/bf16_hf \
  --out-dir      /path/to/mxfp8_hf

# Run with MXFP8 (automatically selected on Blackwell)

export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1
python scripts/run_qwen3_30b_a3b.py \
  --model-dir /path/to/models \
  --model-name qwen3-30b-a3b \
  --fp8-format e4m3 \
  --fp8-recipe blockwise

```

### Stability Guarantee

- **Production‑ready on Blackwell hardware**
- The low‑precision README states MXFP8 "can achieve more efficient inference throughput and lower training‑inference mismatch, resulting in more stable training" ([`examples/infra_features/low_precision/README.md`](https://github.com/radixark/miles/blob/main/examples/infra_features/low_precision/README.md) lines 3‑4)
- Dimensional validation prevents silent failures
- No known fatal bugs in current `main` branch

## NVFP4 (NVIDIA FP4 E2M1)

### Implementation Details

The **NVFP4** format resides in [`miles/utils/nvfp4.py`](https://github.com/radixark/miles/blob/main/miles/utils/nvfp4.py) via the `NVFP4Quantizer` wrapper. Like MXFP8, it enforces alignment constraints: `NVFP4_GROUP_SIZE = 16` must divide weight dimensions, with validation at lines 7‑9.

```python

# From miles/utils/nvfp4.py

NVFP4_GROUP_SIZE = 16

def _validate_weights(weight):
    if weight.shape[-1] % NVFP4_GROUP_SIZE != 0:
        raise RuntimeError(
            f"Weight dimension {weight.shape[-1]} not aligned to "
            f"NVFP4_GROUP_SIZE ({NVFP4_GROUP_SIZE})"
        )

```

### Activation Pipeline

```bash

# Convert to NVFP4

python tools/convert_hf_to_nvfp4.py \
  --hf-checkpoint /path/to/bf16_hf \
  --out-dir      /path/to/nvfp4_hf

# Configure environment and run

export NVTE_NVFP4_4OVER6=all
python scripts/run_qwen3_4b.py \
  --model-dir /path/to/models \
  --model-name qwen3-4b \
  --fp4-format e2m1

```

### Stability Assessment

- **Experimental but generally stable**
- Extensive test coverage in [`tests/fast-gpu/test_nvfp4_quantizer.py`](https://github.com/radixark/miles/blob/main/tests/fast-gpu/test_nvfp4_quantizer.py) (marked "requires Blackwell/B200 CI runner")
- Functional on supported hardware with no stability warnings in documentation
- Suitable for MoE models seeking maximum memory reduction

## Standard FP8 (Blockwise)

### Distinction from MXFP8

While sharing activation flags with MXFP8, **standard blockwise FP8** uses pure `torch.float8_e4m3fn` conversion via [`tools/convert_hf_to_fp8.py`](https://github.com/radixark/miles/blob/main/tools/convert_hf_to_fp8.py). The format scales per‑block rather than per‑matrix, offering broader hardware compatibility.

### Configuration and Execution

```bash

# Convert checkpoint

python tools/convert_hf_to_fp8.py \
  --hf-checkpoint /path/to/bf16_hf \
  --out-dir      /path/to/fp8_hf

# Run with blockwise scaling

export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1
python scripts/run_qwen3_4b.py \
  --model-dir /path/to/models \
  --model-name qwen3-4b \
  --fp8-format e4m3 \
  --fp8-recipe blockwise

```

### Stability Profile

- **Stable for inference and training** on supported GPUs
- Same README stability claims as MXFP8 regarding throughput and training‑inference mismatch reduction
- **Known limitation**: FP8 weights (`--fp8-param-gather`) conflict with CPU offloading techniques, requiring the fused Adam optimizer ([`examples/infra_features/low_precision/README.md`](https://github.com/radixark/miles/blob/main/examples/infra_features/low_precision/README.md) lines 62‑66)
- Safe for production deployment when optimizer constraints are respected

## INT4 (4‑bit Integer with Fake QAT)

### Fake‑QAT Architecture

**INT4** in Miles operates through a **Fake Quantization‑Aware Training** pipeline using straight‑through estimation (STE). The conversion tools [`tools/convert_hf_to_int4_direct.py`](https://github.com/radixark/miles/blob/main/tools/convert_hf_to_int4_direct.py) and [`tools/convert_hf_to_hf_int4.py`](https://github.com/radixark/miles/blob/main/tools/convert_hf_to_hf_int4.py) implement NVIDIA's PTQ pipeline with trainable quantization simulation.

### Environment and Execution

```bash

# Required environment configuration

export OPEN_TRAINING_INT4_FAKE_QAT_FLAG=1
export OPEN_TRAINING_INT4_GROUP_SIZE=128  # Model‑specific: 128 for Qwen‑3‑30B‑A3B

# Convert and run

python tools/convert_hf_to_int4_direct.py \
  --input-dir  /path/to/bf16_hf \
  --output-dir /path/to/int4_hf

python scripts/run_qwen3_30b_a3b.py \
  --model-dir /path/to/models \
  --model-name qwen3-30b-a3b \
  --int4-format

```

### Stability Classification

- **Experimentally stable** — functional with caveats
- The README explicitly identifies "INT4 STE training" as providing "significant throughput improvement" but **does not guarantee full numerical stability**
- Documented bugs remain; the same optimizer offloading conflict affecting FP8 weights applies ([`README.md`](https://github.com/radixark/miles/blob/main/README.md) lines 62‑66)
- Suitable for experimentation and throughput‑critical scenarios where occasional precision loss is acceptable

## FP8 QAT (DeepSeek‑V4 Quantization‑Aware Training)

### Simulation‑Based Implementation

The **QAT** format in Miles is implemented in [`miles_plugins/models/deepseek_v4/ops/qat.py`](https://github.com/radixark/miles/blob/main/miles_plugins/models/deepseek_v4/ops/qat.py) as a custom autograd function that **simulates** FP8 quantization without modifying underlying BF16 weights. The `DeepSeekV4LinearQATFunc.apply` forward/backward pass mimics FP8 behavior during training.

```python

# From miles_plugins/models/deepseek_v4/ops/qat.py

class DeepSeekV4LinearQATFunc(torch.autograd.Function):
    @staticmethod
    def forward(ctx, input, weight, ...):
        # Simulate FP8 quantization without altering BF16 storage

        fp8_simulated = simulate_fp8_quant(input, weight)
        return fp8_simulated

```

### Automatic Activation

QAT requires no explicit flags. When `quantization_config` specifies `"qat"` in a DeepSeek‑V4 checkpoint, the simulation layer inserts automatically:

```bash

# Standard FP8 conversion sufficient

python tools/convert_hf_to_fp8.py \
  --hf-checkpoint /path/to/deepseek_v4_bf16 \
  --out-dir      /path/to/deepseek_v4_fp8

# QAT simulates FP8 behavior at inference time

```

### Stability Characteristics

- **Inherently stable** — simulation preserves original weights
- No runtime exceptions or dimensional constraints
- No impact on training stability (inference‑time only)
- Test suite includes dedicated QAT validation

## Format Comparison and Selection Guide

| Format | Module Path | Target Hardware | Stability Level | Best Use Case |
|--------|-------------|-----------------|-----------------|---------------|
| MXFP8 | [`miles/utils/mxfp8.py`](https://github.com/radixark/miles/blob/main/miles/utils/mxfp8.py) | Blackwell/SM100+ | **Production** | Maximum efficiency on latest GPUs |
| FP8 (blockwise) | [`tools/convert_hf_to_fp8.py`](https://github.com/radixark/miles/blob/main/tools/convert_hf_to_fp8.py) | Ampere+ | **Production** | Broad compatibility, stable training |
| NVFP4 | [`miles/utils/nvfp4.py`](https://github.com/radixark/miles/blob/main/miles/utils/nvfp4.py) | Blackwell/B200 | **Experimental** | Extreme compression for MoE models |
| INT4 QAT | [`tools/convert_hf_to_int4_direct.py`](https://github.com/radixark/miles/blob/main/tools/convert_hf_to_int4_direct.py) | Various | **Experimental** | Throughput‑critical, precision‑tolerant |
| FP8 QAT | [`miles_plugins/models/deepseek_v4/ops/qat.py`](https://github.com/radixark/miles/blob/main/miles_plugins/models/deepseek_v4/ops/qat.py) | Any (simulation) | **Production** | Inference‑time FP8 behavior testing |

## Summary

- **MXFP8** and **FP8 blockwise** are the **most stable low‑precision formats** in Miles, recommended for production training and inference on supported NVIDIA GPUs
- **NVFP4** delivers maximum compression but remains experimental; verify on Blackwell/B200 hardware before deployment
- **INT4 Fake‑QAT** provides substantial throughput gains through STE training, yet exhibits documented instability and optimizer conflicts
- **FP8 QAT** offers risk‑free FP8 simulation for DeepSeek‑V4 models without modifying base weights
- All formats enforce dimensional alignment through explicit validation in their respective quantizer implementations

## Frequently Asked Questions

### What GPU hardware is required for Miles MXFP8 support?

MXFP8 requires **Blackwell architecture (SM100+) GPUs** such as B100 or B200. The format automatically activates on compatible hardware when using `--fp8-recipe blockwise` with `NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1`. Older GPUs fall back to standard blockwise FP8.

### Why does Miles INT4 training have stability warnings?

INT4 uses **Fake QAT with straight‑through estimation**, which introduces approximation errors during gradient computation. The README explicitly notes remaining bugs and warns that the technique conflicts with common optimizer configurations like Adam CPU offloading—unlike FP8 formats which maintain full gradient precision.

### Can NVFP4 and MXFP8 be used together in the same model?

No—Miles quantizers are **mutually exclusive** per layer. NVFP4 (`--fp4-format e2m1`) and MXFP8 (`--fp8-format e4m3`) target different bit widths and scaling schemes. Model‑wise mixing would require manual layer‑by‑layer configuration not currently supported by the conversion scripts.

### What is the optimizer conflict mentioned for FP8 and INT4 weights?

Both **FP8 weights (`--fp8-param-gather`)** and **INT4 training** currently require the fused Adam optimizer implementation. This conflicts with **CPU offloading techniques** that move optimizer states to host memory. The low‑precision README documents this at lines 62‑66 as a known limitation awaiting resolution.