What Types of Language Models Can Be Trained with MegaDLMs? A Complete Guide
MegaDLMs supports four distinct language model families: Diffusion Language Models (DLMs), Autoregressive Language Models (AR-LMs), Multimodal Vision-Language Models, and specialized encoder-decoder architectures like BERT and T5—all trainable with dense or Mixture-of-Experts (MoE) configurations.
MegaDLMs is a GPU-optimized training framework built atop Megatron-LM that enables large-scale pretraining and fine-tuning across diverse architectural paradigms. Whether you are developing diffusion-based generative models or classic causal transformers, the repository provides unified infrastructure for tensor, pipeline, and data parallelism with FP8/FP16/BF16 precision support.
Supported Language Model Architectures
MegaDLMs organizes model implementations into modular families, each residing in dedicated subdirectories under megatron/core/models/. The framework supports training from scratch, fine-tuning existing checkpoints, or distillation via utilities in tools/.
Diffusion Language Models (DLMs)
Diffusion Language Models implement a transformer-based denoising process through DiffLMTransformerBlock in megatron/core/models/difflm/gpt_model.py. Unlike autoregressive models that predict tokens sequentially, DLMs learn to reverse a noising process over the entire sequence.
The core implementation uses DiffLMTransformerBlock which exposes the same interface (forward, set_input_tensor) as standard transformer blocks but implements diffusion-specific forward passes. You can configure dense or Mixture-of-Experts (MoE) variants by adjusting args.num_experts and using the get_difflm_decoder_block_spec specification.
Entry point: pretrain_difflm.py
Core architecture: megatron/core/models/difflm/gpt_model.py
Autoregressive Language Models (AR-LMs)
Autoregressive Language Models represent the classic "next-token" prediction paradigm implemented via standard causal transformers. The TransformerBlock class in megatron/core/models/gpt/gpt_model.py provides the backbone for GPT-style architectures.
These models support both dense configurations and MoE scaling through args.num_experts combined with specifications like get_gpt_layer_with_transformer_engine_spec or get_gpt_layer_local_spec. The implementation maintains full compatibility with Megatron-LM's optimization kernels via NVIDIA Transformer Engine.
Core architecture: megatron/core/models/gpt/gpt_model.py
Layer specifications: megatron/core/models/gpt/gpt_layer_specs.py
Multimodal Vision-Language Models
Multimodal architectures combine language backbones (either AR-LM or DLM) with vision encoders such as CLIP-ViT. The integration layer resides in megatron/core/models/multimodal/llava_model.py, which wires vision front-ends to dense or MoE language cores.
This enables training of LLaVA-style models where visual inputs are projected into the language model's embedding space, allowing joint pretraining on interleaved image-text data.
Implementation: megatron/core/models/multimodal/llava_model.py
Specialized Encoder-Decoder Models
MegaDLMs retains support for task-specific architectures including BERT and T5 through adapters in megatron/core/models/bert/bert_model.py and megatron/core/models/T5/t5_model.py. These currently support dense configurations only and reuse the same Megatron-LM distributed training infrastructure for masked language modeling and sequence-to-sequence tasks.
BERT implementation: megatron/core/models/bert/bert_model.py
T5 implementation: megatron/core/models/T5/t5_model.py
Mixture-of-Experts (MoE) Scaling for All Architectures
Both DLMs and AR-LMs support seamless Mixture-of-Experts scaling without code changes. When args.num_experts is set to a value greater than zero, the framework automatically injects MoE routing layers through the specification functions:
- DLM MoE:
get_difflm_decoder_block_specinmegatron/core/models/difflm/gpt_layer_specs.py - AR-LM MoE:
get_gpt_layer_with_transformer_engine_specinmegatron/core/models/gpt/gpt_layer_specs.py
The moe-grouped-gemm flag enables optimized grouped GEMM kernels for expert computation, allowing efficient scaling to hundreds of billions of parameters.
Unified Training Infrastructure
MegaDLMs handles diverse model types through a unified abstraction layer that minimizes code duplication across training scripts.
Pluggable Transformer Blocks
All model families implement self.model_type = ModelType.encoder_or_decoder, enabling the same training loops in pretrain_difflm.py or generic pretrain.py to drive forward and backward passes. The training script switches between DiffLMTransformerBlock and TransformerBlock by passing different specifications to the model provider, while both blocks maintain identical interfaces.
Shared Data Pipeline and Parallelism
All model types consume tokenized .jsonl data processed via tools/preprocess_data.py. They share BlendedMegatronDatasetBuilder and GPTDataset utilities, ensuring tokenizer-agnostic pipelines. The framework automatically manages tensor parallelism, pipeline parallelism, and data parallelism regardless of whether you are training a DLM, AR-LM, or multimodal variant.
Training Examples
The following commands demonstrate how to train different types of language models with MegaDLMs.
Train a Dense Diffusion Language Model
python pretrain_difflm.py \
--model-type difflm \
--num-layers 32 \
--hidden-size 4096 \
--num-heads 32 \
--seq-length 2048 \
--batch-size 8 \
--data-path /path/to/tokenized_data \
--output-dir /tmp/dlm_checkpoints \
--max-steps 250000
Train a Dense Autoregressive GPT Model
python pretrain_difflm.py \
--model-type gpt \
--num-layers 24 \
--hidden-size 3072 \
--num-heads 24 \
--seq-length 1024 \
--batch-size 16 \
--data-path /path/to/tokenized_data \
--output-dir /tmp/gpt_checkpoints \
--max-steps 200000
Enable Mixture-of-Experts (MoE)
Add --num-experts 8 to either command above. For example, training an MoE DLM:
python pretrain_difflm.py \
--model-type difflm \
--num-experts 8 \
--moe-grouped-gemm true \
--num-layers 32 \
--hidden-size 4096 \
--num-heads 32 \
--data-path /path/to/tokenized_data \
--output-dir /tmp/moe_dlm_checkpoints
Minimal Model Provider Implementation
Training scripts use a provider function to instantiate models based on arguments:
from megatron.training import get_args
from megatron.core.models.difflm.gpt_model import GPTModel as DiffLM
from megatron.core.models.gpt.gpt_model import GPTModel as ARModel
from megatron.core.models.difflm.gpt_layer_specs import get_difflm_decoder_block_spec
from megatron.core.models.gpt.gpt_layer_specs import get_gpt_layer_local_spec
def model_provider(pre_process=True, post_process=True):
args = get_args()
spec = (get_difflm_decoder_block_spec if args.model_type == "difflm"
else get_gpt_layer_local_spec)
transformer_layer_spec = spec(args)
if args.model_type == "difflm":
return DiffLM(
config=args,
transformer_layer_spec=transformer_layer_spec,
vocab_size=args.padded_vocab_size,
max_sequence_length=args.max_position_embeddings,
pre_process=pre_process,
post_process=post_process
)
else:
return ARModel(
config=args,
transformer_layer_spec=transformer_layer_spec,
vocab_size=args.padded_vocab_size,
max_sequence_length=args.max_position_embeddings,
pre_process=pre_process,
post_process=post_process
)
Summary
- MegaDLMs supports four primary model families: Diffusion Language Models (DLMs), Autoregressive Language Models (AR-LMs), Multimodal Vision-Language Models, and specialized architectures (BERT/T5).
- Unified infrastructure: All models leverage
ModelType.encoder_or_decoder, pluggable transformer blocks, and shared Megatron-LM parallelism primitives. - Flexible scaling: Both DLMs and AR-LMs support dense and Mixture-of-Experts (MoE) configurations via
args.num_expertsand specification functions likeget_difflm_decoder_block_spec. - Common entry point: The
pretrain_difflm.pyscript handles both diffusion and autoregressive training, switching models via the--model-typeargument. - Multimodal ready: Vision-language integration is available through
megatron/core/models/multimodal/llava_model.py, combining vision encoders with either DLM or AR-LM backbones.
Frequently Asked Questions
Can I train both diffusion and autoregressive models using the same script?
Yes. The pretrain_difflm.py entry point handles both model families. You specify the architecture via --model-type difflm for Diffusion Language Models or --model-type gpt for Autoregressive Language Models. The script uses a unified model provider that switches between DiffLMTransformerBlock and standard TransformerBlock based on this flag, while maintaining identical data loading and distributed training configurations.
Does MegaDLMs support Mixture-of-Experts for all model types?
MoE is fully supported for DLMs and AR-LMs through the specification system in gpt_layer_specs.py files. Setting --num-experts greater than zero automatically injects routing layers into the transformer blocks. However, specialized variants like BERT and T5 currently support dense configurations only, as indicated in their respective model implementations under megatron/core/models/bert/ and megatron/core/models/T5/.
What multimodal capabilities does MegaDLMs offer?
MegaDLMs supports Vision-Language models through the LLaVA architecture implemented in megatron/core/models/multimodal/llava_model.py. This combines vision encoders (such as CLIP-ViT) with either diffusion or autoregressive language backbones. The framework handles vision-to-language projection layers and enables joint training on interleaved image-text datasets using the same parallelization strategies as pure language models.
How does MegaDLMs handle data preprocessing for different model types?
All model types use the same tokenizer-agnostic pipeline. The tools/preprocess_data.py script processes raw text or multimodal data into tokenized .jsonl format. During training, BlendedMegatronDatasetBuilder and GPTDataset classes handle data loading uniformly across DLM, AR-LM, and multimodal configurations, ensuring consistent batching and sequence packing regardless of the underlying architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →