# Benefits of Using FP8 Precision with MegaDLMs and Required Hardware

> Unlock up to 50% faster MegaDLM training with FP8 precision. Reduce memory usage and accelerate inference on NVIDIA Hopper Ada or Blackwell GPUs. Discover hardware requirements and benefits.

- Repository: [Jinjie Ni/megadlms](https://github.com/jinjieni/megadlms)
- Tags: deep-dive
- Published: 2026-03-04

---

**Using FP8 precision with MegaDLMs delivers up to 50% faster forward passes and 84% faster backward propagation while reducing memory usage through 8-bit activation storage, but requires NVIDIA Hopper, Ada, or Blackwell GPUs with native Tensor Core FP8 support.**

MegaDLMs leverage NVIDIA's Transformer Engine to implement optional **FP8 (8-bit floating point)** training modes, enabling significant performance optimizations for large-scale diffusion language models. According to the `jinjieni/megadlms` source code, this first-generation low-precision format integrates seamlessly with existing optimizations like FlashAttention while maintaining model quality. Understanding the specific performance gains and strict hardware requirements ensures developers can effectively utilize this mixed-precision capability.

## Performance Benefits of FP8 Precision

FP8 acceleration in MegaDLMs provides measurable speedups and resource savings across multiple training phases. The implementation focuses on computational efficiency, memory footprint reduction, and I/O optimization.

### Training Speed Improvements

The Transformer Engine integration yields substantial kernel-level optimizations for tensor operations. When running on compatible hardware, MegaDLMs achieve:

- **Up to 50% speedup on forward passes** through optimized FP8 kernels for attention and MLP operations
- **Up to 84% speedup on backward propagation** via reduced-precision gradient calculations
- **Compounded acceleration** when combined with FlashAttention and cuDNN-based fused kernels

According to the repository's [`README.md`](https://github.com/jinjieni/megadlms/blob/main/README.md), these gains are realized through the `--fp8-hybrid` flag, which enables FP8 training while preserving numerical stability for critical operations.

### Memory and Storage Optimization

Beyond raw computation speed, FP8 precision significantly reduces memory pressure during training. In [`megatron/core/transformer/transformer_config.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/transformer/transformer_config.py), the configuration field `activation_func_fp8_input_store` enables storing MLP activation inputs in FP8 format specifically for backpropagation, cutting activation memory footprints substantially.

The benefits extend to checkpointing operations as well. The distributed checkpointing system in [`megatron/core/dist_checkpointing/strategies/torch.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/dist_checkpointing/strategies/torch.py) contains specialized handling for FP8 tensors, resulting in smaller checkpoint files and faster optimizer state exchanges across distributed nodes.

## Required Hardware for FP8 Support

FP8 acceleration requires modern NVIDIA GPUs with native Tensor Core FP8 arithmetic units. The `jinjieni/megadlms` repository explicitly supports three GPU architectures:

- **Hopper** (e.g., NVIDIA H100)
- **Ada** (e.g., NVIDIA Ada series)
- **Blackwell** (e.g., NVIDIA B200)

These architectures implement the Tensor Core FP8 support necessary for Transformer Engine operations. Older Tensor Core generations such as Volta and Turing cannot execute MegaDLMs in FP8 mode, as documented in the hardware requirements section of [`README.md`](https://github.com/jinjieni/megadlms/blob/main/README.md).

## Enabling FP8 in MegaDLMs

Activation requires both command-line flags and optional configuration adjustments depending on your specific memory and precision requirements.

### Command-Line Configuration

The training script accepts several FP8-related arguments defined in [`megatron/training/arguments.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/arguments.py):

```bash
python pretrain_difflm.py \
    --model-path megatron-13b \
    --data-path /data/tokenized \
    --fp8-hybrid \
    --fp8-format e4m3 \
    --fp8-margin 0

```

Key flags include:

- `--fp8-hybrid`: Enables hybrid FP8 training mode with automatic optimizer adjustments
- `--fp8-format`: Selects the FP8 exponent-mantissa layout (`e4m3` or `hybrid`)
- `--no-fp8-wgrad`: Optionally maintains weight-gradient computation in higher precision when numerical stability is critical

### Python API Configuration

For programmatic control, instantiate `TransformerConfig` with FP8-specific parameters:

```python
from megatron.core.transformer.transformer_config import TransformerConfig

cfg = TransformerConfig(
    hidden_size=4096,
    num_attention_heads=32,
    activation_func_fp8_input_store=True,  # Enable FP8 activation storage

    fp8=True,
    fp8_format="e4m3",
)

model = build_model(cfg)

```

Setting `activation_func_fp8_input_store=True` stores MLP activation inputs in FP8 format, maximizing memory savings during the backward pass.

### Runtime Verification

Confirm FP8 is active during execution by checking the Transformer Engine global state:

```python
from transformer_engine.pytorch.fp8 import FP8GlobalStateManager

if FP8GlobalStateManager.is_fp8_enabled():
    print("Running with FP8!")
    print("FP8 recipe:", FP8GlobalStateManager.get_fp8_recipe())
else:
    print("FP8 not enabled")

```

This verification method is integrated into MegaDLMs through [`megatron/core/cuda_graphs.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/cuda_graphs.py), which manages FP8 state during CUDA graph capture and replay operations.

## Summary

Using FP8 precision with MegaDLMs offers significant advantages for large-scale model training when deployed on compatible hardware:

- **Up to 50% faster forward passes** and **84% faster backward passes** via Transformer Engine kernels
- **Reduced memory consumption** through 8-bit activation storage and smaller checkpoint files
- **Compounded performance gains** when combined with FlashAttention and cuDNN optimizations
- **Strict hardware requirements** limiting deployment to NVIDIA Hopper, Ada, and Blackwell GPUs
- **Simple activation** through CLI flags (`--fp8-hybrid`) or Python configuration (`activation_func_fp8_input_store`)

## Frequently Asked Questions

### What GPUs are required to run MegaDLMs with FP8 precision?

FP8 training requires NVIDIA GPUs with native Tensor Core FP8 support, specifically the **Hopper** (H100), **Ada**, or **Blackwell** (B200) architectures. Older GPUs like Volta or Turing lack the necessary FP8 arithmetic units and cannot execute MegaDLMs in this mode, as documented in the repository's hardware requirements section.

### How much performance improvement does FP8 provide compared to FP16?

According to the MegaDLMs README, FP8 precision delivers **up to 50% speedup on forward passes** and **up to 84% speedup on backward propagation** compared to traditional FP16 or BF16 training. These gains come from optimized kernels in the Transformer Engine that handle attention and MLP operations more efficiently at 8-bit precision.

### Does using FP8 reduce model accuracy or training stability?

The repository implements **hybrid FP8 training** modes that maintain numerical stability for critical operations. The `--fp8-hybrid` flag automatically handles precision-sensitive calculations while keeping high-throughput operations in FP8. Additionally, the `--no-fp8-wgrad` option allows weight gradients to remain in higher precision if specific training scenarios require enhanced stability.

### How do I enable memory-saving FP8 activation storage?

Set `activation_func_fp8_input_store=True` in your `TransformerConfig` initialization (found in [`megatron/core/transformer/transformer_config.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/transformer/transformer_config.py)). This configuration stores MLP activation inputs in FP8 format specifically for backpropagation, significantly reducing memory usage during training without requiring changes to the model architecture.