MegaDLMs: GPU-Optimized Training Framework for Diffusion Language Models

MegaDLMs is a GPU-optimized training framework designed specifically for Diffusion Language Models (DLMs) that leverages Megatron-LM and NVIDIA Transformer-Engine to enable scalable, high-throughput training from billions to hundreds of billions of parameters.

Available in the jinjieni/megadlms repository, MegaDLMs provides a production-ready backend for diffusion-based language model research. The framework combines advanced parallelism strategies with modern precision formats like FP8 and BF16 to achieve superior Model FLOP Utilisation (MFU) across thousands of GPUs.

Architecture and Core Components

Megatron Core Integration

At the foundation of MegaDLMs lies Megatron Core, which supplies low-level building blocks including optimized kernels, tensor parallelism, pipeline parallelism, and distributed optimizers. These components reside in the megatron/core/ directory, specifically within core/models, core/transformer, core/tensor_parallel, and core/pipeline_parallel.

Diffusion Model Stack

The framework implements the DLM architecture through specialized Python classes defined in tools/weights_conversion/hf_configs/gptneox_1.7b_dlm/modeling_dlm.py. Key implementations include:

  • DLMMLP – handles feed-forward layers with RMSNorm
  • DLMAttention – implements rotary embeddings and attention mechanisms
  • DLMDecoderLayer – composes attention and MLP blocks
  • DLMModel – the complete model wrapper integrating all components

Parallelism Strategies

MegaDLMs combines six distinct parallelism approaches to maximize hardware utilization: Data Parallelism (DP), Tensor Parallelism (TP), Pipeline Parallelism (PP), Context Parallelism (CP), Expert Parallelism (EP), and Fully-Sharded Data Parallel (FSDP). Users configure these strategies via command-line flags defined in megatron/training/arguments.py, including --tensor-model-parallel-size, --pipeline-model-parallel-size, and --expert-model-parallel-size.

Training Pipeline and Inference

The end-to-end workflow encompasses data preprocessing, training orchestration, checkpoint management, and production inference. Tokenization runs through tools/preprocess_data.py, training launches via scripts like examples/dlm_training/dlm_pretrain_1.7b.sh, and deployment utilizes the HTTP/REST server in megatron/inference/text_generation_server.py.

Primary Purpose and Performance Goals

The primary purpose of MegaDLMs is to provide a high-throughput, scalable backend for training diffusion language models at any scale. By tightly integrating with Megatron-LM's parallelism primitives and NVIDIA's Transformer-Engine, the framework achieves up to 3× faster training speeds compared to competing implementations.

Key optimization features include:

  • Flash-attention for memory-efficient attention computation
  • Activation checkpointing to reduce memory footprint
  • Communication overlap to hide latency in distributed settings
  • Precision flexibility supporting FP8, BF16, and FP16 formats

These capabilities are centrally configured through megatron/training/arguments.py, allowing researchers to toggle optimizations without modifying source code.

Practical Implementation Guide

Data Preprocessing

Tokenize JSONL corpora using the provided preprocessing utility:

python tools/preprocess_data.py \
    --input data.jsonl \
    --output-prefix /tmp/processed \
    --tokenizer-type HuggingFaceTokenizer \
    --tokenizer-model /path/to/tokenizer.model \
    --workers 8 \
    --append-eod

Training Configuration

Launch a 1.7B-parameter DLM pre-training run:

source envs/.env
bash examples/dlm_training/dlm_pretrain_1.7b.sh

Checkpoint Conversion

Convert HuggingFace checkpoints to Megatron-compatible format:

python tools/weights_conversion/hf_to_megatron_te.py \
    --hf-checkpoint /path/to/hf/model \
    --output-dir /path/to/megatron_ckpt \
    --model-type gptneox_1.7b_dlm

Validate conversion results using tools/weights_conversion/utils/verify_correctness_dlm.py to ensure numerical precision and architecture alignment.

Inference Deployment

Deploy the text-generation server for production inference:

python megatron/inference/text_generation_server.py \
    --model-path /path/to/megatron_ckpt \
    --port 8000 \
    --tensor-model-parallel-size 4 \
    --pipeline-model-parallel-size 2

Summary

  • MegaDLMs is a specialized framework for training Diffusion Language Models at scale, built on top of Megatron-LM and NVIDIA Transformer-Engine.
  • The architecture combines multiple parallelism strategies (TP, PP, CP, EP, FSDP) to efficiently distribute training across thousands of GPUs.
  • Key implementation files include modeling_dlm.py for model architecture, arguments.py for training configuration, and text_generation_server.py for inference.
  • The framework supports end-to-end workflows from data preprocessing through checkpoint conversion to production deployment.
  • MegaDLMs achieves superior performance through FP8/BF16 precision support, flash-attention, and activation checkpointing.

Frequently Asked Questions

What is the difference between MegaDLMs and standard Megatron-LM?

MegaDLMs extends Megatron-LM by adding specialized support for Diffusion Language Models through custom architecture classes like DLMModel and DLMAttention in tools/weights_conversion/hf_configs/gptneox_1.7b_dlm/modeling_dlm.py. While it maintains compatibility with Megatron's parallelism infrastructure, it optimizes specifically for the non-autoregressive training patterns characteristic of diffusion models.

Can MegaDLMs train autoregressive models as well?

Yes, according to the source code, MegaDLMs supports both diffusion language models and conventional autoregressive (AR) language models. The framework's flexibility allows researchers to switch between model types while maintaining the same high-performance training infrastructure and parallelism configurations.

What hardware requirements does MegaDLMs have?

MegaDLMs is designed for NVIDIA GPUs and leverages Transformer-Engine optimizations. The framework supports training configurations from single nodes to thousands of GPUs using various parallelism strategies. Specific hardware requirements depend on model size, but the framework requires CUDA-capable GPUs with support for FP8 or BF16 precision for optimal performance.

How does checkpoint conversion work between HuggingFace and Megatron formats?

The conversion process uses tools/weights_conversion/hf_to_megatron_te.py to transform HuggingFace checkpoints into Megatron-compatible formats. After conversion, tools/weights_conversion/utils/verify_correctness_dlm.py validates numerical precision and architecture alignment. This enables seamless migration of pre-trained models into the MegaDLMs training ecosystem.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →