The Three Stages of the DLM Training Pipeline in MegaDLMs: A Complete Guide

The MegaDLMs repository implements a complete three-stage training pipeline for Diffusion Language Models (DLMs) consisting of pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL), with dedicated shell scripts for each stage located in the examples/dlm_training/ directory.

The jinjieni/megadlms repository provides a comprehensive framework for training Diffusion Language Models from scratch through deployment. According to the project documentation, the DLM training pipeline follows a structured three-stage approach explicitly described as supporting models "from pre-training and SFT to RL" (see README.md line 26). This architecture ensures models develop general language capabilities, task-specific proficiency, and human-aligned behavior sequentially.

Stage 1: Pre-Training

The first stage establishes foundational language representations by training the model from scratch on massive tokenized corpora. This unsupervised learning phase teaches the diffusion model general linguistic patterns and world knowledge.

To launch pre-training, execute the provided driver script:

source envs/.env
bash examples/dlm_training/dlm_pretrain_1.7b.sh

The examples/dlm_training/dlm_pretrain_1.7b.sh script contains the full configuration for training a 1.7 billion parameter DLM, including data paths, model architecture specifications, and distributed training parameters. This script serves as the primary entry point for the initial training phase.

Stage 2: Supervised Fine-Tuning (SFT)

After pre-training, the model undergoes supervised fine-tuning on labeled instruction or task-specific datasets. This stage adapts the general-purpose model to follow instructions and perform specific downstream tasks with higher accuracy.

Run the SFT stage using the skeleton script provided:

source envs/.env

# Replace <SFT_CFG> with your SFT config file

bash examples/dlm_training/dlm_sft.sh <SFT_CFG>

The examples/dlm_training/dlm_sft.sh script follows the same argument conventions as the pre-training script but accepts a configuration file specific to supervised fine-tuning. This design maintains consistency across pipeline stages while allowing stage-specific hyperparameter adjustments.

Stage 3: Reinforcement Learning (RL)

The final stage applies reinforcement learning (typically RLHF - Reinforcement Learning from Human Feedback) to align the model with human preferences and optimize for reward signals beyond simple loss minimization.

Execute the RL stage with:

source envs/.env

# Replace <RL_CFG> with your RL config file

bash examples/dlm_training/dlm_rl.sh <RL_CFG>

The examples/dlm_training/dlm_rl.sh script wraps the same Megatron-LM trainer used in previous stages but configures it for RL-specific arguments. This ensures hardware configurations (tensor parallelism, pipeline parallelism, data paths) remain consistent across all three stages of the DLM training pipeline.

Common Configuration and Shared Arguments

All three training scripts share a common interface for distributed training and model configuration. This architectural consistency allows seamless transitions between pipeline stages without reconfiguring hardware setups.

Key shared arguments include:

  • Model size specifications (hidden dimensions, number of layers, attention heads)
  • Data paths (training data, validation data, tokenizer paths)
  • Parallelism configurations (tensor parallelism size, pipeline parallelism size, data parallelism)
  • Optimization parameters (learning rate schedules, batch sizes, gradient clipping)

When moving from pre-training to SFT or RL, you typically only need to update the configuration file path and stage-specific hyperparameters while maintaining the same distributed training topology.

Verifying Model Weights Between Stages

The repository includes utilities to validate model integrity across the three-stage pipeline. After completing any training stage, you can verify weight correctness using:

python tools/weights_conversion/utils/verify_correctness_dlm.py --checkpoint-path <PATH>

The tools/weights_conversion/utils/verify_correctness_dlm.py script validates that model weights are correctly saved and can be loaded for subsequent training stages. This verification step is crucial when transitioning between pre-training, SFT, and RL to ensure weight compatibility and prevent corruption during the multi-stage DLM training pipeline.

Summary

  • MegaDLMs implements a three-stage training pipeline consisting of pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL) for Diffusion Language Models.
  • Each stage has dedicated entry scripts located in examples/dlm_training/: dlm_pretrain_1.7b.sh for pre-training, dlm_sft.sh for SFT, and dlm_rl.sh for RL.
  • All stages share common configuration patterns for distributed training, allowing consistent hardware setups across the pipeline while using stage-specific configuration files.
  • Weight verification utilities in tools/weights_conversion/utils/verify_correctness_dlm.py ensure model integrity between stages.

Frequently Asked Questions

What is the purpose of each stage in the MegaDLMs training pipeline?

Pre-training establishes general language understanding by training the diffusion model from scratch on large corpora. Supervised Fine-Tuning (SFT) adapts the pre-trained model to specific tasks and instruction-following using labeled datasets. Reinforcement Learning (RL) aligns the model with human preferences or specific reward functions, typically using RLHF techniques to optimize beyond standard loss metrics.

How do I transition a model from pre-training to SFT in MegaDLMs?

After completing pre-training using examples/dlm_training/dlm_pretrain_1.7b.sh, prepare an SFT-specific configuration file and launch examples/dlm_training/dlm_sft.sh <SFT_CFG>. The scripts share the same distributed training arguments, so you can maintain your tensor and pipeline parallelism settings while only updating the data paths and hyperparameters specific to supervised fine-tuning. Use tools/weights_conversion/utils/verify_correctness_dlm.py to validate checkpoint integrity between stages.

Can I skip the RL stage and use only pre-training and SFT?

Yes, the three-stage pipeline is modular. You can deploy a model after completing pre-training and SFT without proceeding to reinforcement learning. However, skipping RL means the model will not benefit from preference alignment or reward-based optimization that RLHF provides. The examples/dlm_training/dlm_rl.sh script remains available if you decide to add RL alignment later using the same checkpoint infrastructure.

What hardware configurations are supported across all three training stages?

All three scripts (dlm_pretrain_1.7b.sh, dlm_sft.sh, and dlm_rl.sh) utilize the same Megatron-LM backend and support configurable tensor parallelism, pipeline parallelism, and data parallelism. You define these parameters consistently across stages, typically through environment variables in envs/.env or command-line arguments. This design ensures that a model trained on a specific cluster topology during pre-training can continue training on the same hardware configuration during SFT and RL without requiring architectural changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →