# What Types of Language Models Can Be Trained with MegaDLMs? A Complete Guide

> Discover the diverse language models you can train with MegaDLMs including DLMs AR-LMs multimodal models and BERT T5 architectures Learn more about dense and MoE configurations with this complete guide

- Repository: [Jinjie Ni/megadlms](https://github.com/jinjieni/megadlms)
- Tags: deep-dive
- Published: 2026-03-04

---

**MegaDLMs supports four distinct language model families: Diffusion Language Models (DLMs), Autoregressive Language Models (AR-LMs), Multimodal Vision-Language Models, and specialized encoder-decoder architectures like BERT and T5—all trainable with dense or Mixture-of-Experts (MoE) configurations.**

MegaDLMs is a GPU-optimized training framework built atop Megatron-LM that enables large-scale pretraining and fine-tuning across diverse architectural paradigms. Whether you are developing diffusion-based generative models or classic causal transformers, the repository provides unified infrastructure for tensor, pipeline, and data parallelism with FP8/FP16/BF16 precision support.

## Supported Language Model Architectures

MegaDLMs organizes model implementations into modular families, each residing in dedicated subdirectories under `megatron/core/models/`. The framework supports training from scratch, fine-tuning existing checkpoints, or distillation via utilities in `tools/`.

### Diffusion Language Models (DLMs)

**Diffusion Language Models** implement a transformer-based denoising process through `DiffLMTransformerBlock` in [`megatron/core/models/difflm/gpt_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/difflm/gpt_model.py). Unlike autoregressive models that predict tokens sequentially, DLMs learn to reverse a noising process over the entire sequence.

The core implementation uses `DiffLMTransformerBlock` which exposes the same interface (`forward`, `set_input_tensor`) as standard transformer blocks but implements diffusion-specific forward passes. You can configure dense or **Mixture-of-Experts (MoE)** variants by adjusting `args.num_experts` and using the `get_difflm_decoder_block_spec` specification.

Entry point: [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py)  
Core architecture: [`megatron/core/models/difflm/gpt_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/difflm/gpt_model.py)

### Autoregressive Language Models (AR-LMs)

**Autoregressive Language Models** represent the classic "next-token" prediction paradigm implemented via standard causal transformers. The `TransformerBlock` class in [`megatron/core/models/gpt/gpt_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/gpt/gpt_model.py) provides the backbone for GPT-style architectures.

These models support both dense configurations and **MoE scaling** through `args.num_experts` combined with specifications like `get_gpt_layer_with_transformer_engine_spec` or `get_gpt_layer_local_spec`. The implementation maintains full compatibility with Megatron-LM's optimization kernels via NVIDIA Transformer Engine.

Core architecture: [`megatron/core/models/gpt/gpt_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/gpt/gpt_model.py)  
Layer specifications: [`megatron/core/models/gpt/gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/gpt/gpt_layer_specs.py)

### Multimodal Vision-Language Models

**Multimodal architectures** combine language backbones (either AR-LM or DLM) with vision encoders such as CLIP-ViT. The integration layer resides in [`megatron/core/models/multimodal/llava_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/multimodal/llava_model.py), which wires vision front-ends to dense or MoE language cores.

This enables training of LLaVA-style models where visual inputs are projected into the language model's embedding space, allowing joint pretraining on interleaved image-text data.

Implementation: [`megatron/core/models/multimodal/llava_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/multimodal/llava_model.py)

### Specialized Encoder-Decoder Models

MegaDLMs retains support for **task-specific architectures** including BERT and T5 through adapters in [`megatron/core/models/bert/bert_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/bert/bert_model.py) and [`megatron/core/models/T5/t5_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/T5/t5_model.py). These currently support dense configurations only and reuse the same Megatron-LM distributed training infrastructure for masked language modeling and sequence-to-sequence tasks.

BERT implementation: [`megatron/core/models/bert/bert_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/bert/bert_model.py)  
T5 implementation: [`megatron/core/models/T5/t5_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/T5/t5_model.py)

## Mixture-of-Experts (MoE) Scaling for All Architectures

Both DLMs and AR-LMs support seamless **Mixture-of-Experts** scaling without code changes. When `args.num_experts` is set to a value greater than zero, the framework automatically injects MoE routing layers through the specification functions:

- **DLM MoE**: `get_difflm_decoder_block_spec` in [`megatron/core/models/difflm/gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/difflm/gpt_layer_specs.py)
- **AR-LM MoE**: `get_gpt_layer_with_transformer_engine_spec` in [`megatron/core/models/gpt/gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/gpt/gpt_layer_specs.py)

The `moe-grouped-gemm` flag enables optimized grouped GEMM kernels for expert computation, allowing efficient scaling to hundreds of billions of parameters.

## Unified Training Infrastructure

MegaDLMs handles diverse model types through a unified abstraction layer that minimizes code duplication across training scripts.

### Pluggable Transformer Blocks

All model families implement `self.model_type = ModelType.encoder_or_decoder`, enabling the same training loops in [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py) or generic [`pretrain.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain.py) to drive forward and backward passes. The training script switches between `DiffLMTransformerBlock` and `TransformerBlock` by passing different specifications to the model provider, while both blocks maintain identical interfaces.

### Shared Data Pipeline and Parallelism

All model types consume tokenized `.jsonl` data processed via [`tools/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/preprocess_data.py). They share `BlendedMegatronDatasetBuilder` and `GPTDataset` utilities, ensuring tokenizer-agnostic pipelines. The framework automatically manages tensor parallelism, pipeline parallelism, and data parallelism regardless of whether you are training a DLM, AR-LM, or multimodal variant.

## Training Examples

The following commands demonstrate how to train different types of language models with MegaDLMs.

### Train a Dense Diffusion Language Model

```bash
python pretrain_difflm.py \
  --model-type difflm \
  --num-layers 32 \
  --hidden-size 4096 \
  --num-heads 32 \
  --seq-length 2048 \
  --batch-size 8 \
  --data-path /path/to/tokenized_data \
  --output-dir /tmp/dlm_checkpoints \
  --max-steps 250000

```

### Train a Dense Autoregressive GPT Model

```bash
python pretrain_difflm.py \
  --model-type gpt \
  --num-layers 24 \
  --hidden-size 3072 \
  --num-heads 24 \
  --seq-length 1024 \
  --batch-size 16 \
  --data-path /path/to/tokenized_data \
  --output-dir /tmp/gpt_checkpoints \
  --max-steps 200000

```

### Enable Mixture-of-Experts (MoE)

Add `--num-experts 8` to either command above. For example, training an MoE DLM:

```bash
python pretrain_difflm.py \
  --model-type difflm \
  --num-experts 8 \
  --moe-grouped-gemm true \
  --num-layers 32 \
  --hidden-size 4096 \
  --num-heads 32 \
  --data-path /path/to/tokenized_data \
  --output-dir /tmp/moe_dlm_checkpoints

```

### Minimal Model Provider Implementation

Training scripts use a provider function to instantiate models based on arguments:

```python
from megatron.training import get_args
from megatron.core.models.difflm.gpt_model import GPTModel as DiffLM
from megatron.core.models.gpt.gpt_model import GPTModel as ARModel
from megatron.core.models.difflm.gpt_layer_specs import get_difflm_decoder_block_spec
from megatron.core.models.gpt.gpt_layer_specs import get_gpt_layer_local_spec

def model_provider(pre_process=True, post_process=True):
    args = get_args()
    spec = (get_difflm_decoder_block_spec if args.model_type == "difflm"
            else get_gpt_layer_local_spec)
    transformer_layer_spec = spec(args)
    
    if args.model_type == "difflm":
        return DiffLM(
            config=args, 
            transformer_layer_spec=transformer_layer_spec,
            vocab_size=args.padded_vocab_size,
            max_sequence_length=args.max_position_embeddings,
            pre_process=pre_process, 
            post_process=post_process
        )
    else:
        return ARModel(
            config=args, 
            transformer_layer_spec=transformer_layer_spec,
            vocab_size=args.padded_vocab_size,
            max_sequence_length=args.max_position_embeddings,
            pre_process=pre_process, 
            post_process=post_process
        )

```

## Summary

- **MegaDLMs supports four primary model families**: Diffusion Language Models (DLMs), Autoregressive Language Models (AR-LMs), Multimodal Vision-Language Models, and specialized architectures (BERT/T5).
- **Unified infrastructure**: All models leverage `ModelType.encoder_or_decoder`, pluggable transformer blocks, and shared Megatron-LM parallelism primitives.
- **Flexible scaling**: Both DLMs and AR-LMs support dense and Mixture-of-Experts (MoE) configurations via `args.num_experts` and specification functions like `get_difflm_decoder_block_spec`.
- **Common entry point**: The [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py) script handles both diffusion and autoregressive training, switching models via the `--model-type` argument.
- **Multimodal ready**: Vision-language integration is available through [`megatron/core/models/multimodal/llava_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/multimodal/llava_model.py), combining vision encoders with either DLM or AR-LM backbones.

## Frequently Asked Questions

### Can I train both diffusion and autoregressive models using the same script?

Yes. The [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py) entry point handles both model families. You specify the architecture via `--model-type difflm` for Diffusion Language Models or `--model-type gpt` for Autoregressive Language Models. The script uses a unified model provider that switches between `DiffLMTransformerBlock` and standard `TransformerBlock` based on this flag, while maintaining identical data loading and distributed training configurations.

### Does MegaDLMs support Mixture-of-Experts for all model types?

MoE is fully supported for DLMs and AR-LMs through the specification system in [`gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/gpt_layer_specs.py) files. Setting `--num-experts` greater than zero automatically injects routing layers into the transformer blocks. However, specialized variants like BERT and T5 currently support dense configurations only, as indicated in their respective model implementations under `megatron/core/models/bert/` and `megatron/core/models/T5/`.

### What multimodal capabilities does MegaDLMs offer?

MegaDLMs supports Vision-Language models through the LLaVA architecture implemented in [`megatron/core/models/multimodal/llava_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/multimodal/llava_model.py). This combines vision encoders (such as CLIP-ViT) with either diffusion or autoregressive language backbones. The framework handles vision-to-language projection layers and enables joint training on interleaved image-text datasets using the same parallelization strategies as pure language models.

### How does MegaDLMs handle data preprocessing for different model types?

All model types use the same tokenizer-agnostic pipeline. The [`tools/preprocess_data.py`](https://github.com/jinjieni/megadlms/blob/main/tools/preprocess_data.py) script processes raw text or multimodal data into tokenized `.jsonl` format. During training, `BlendedMegatronDatasetBuilder` and `GPTDataset` classes handle data loading uniformly across DLM, AR-LM, and multimodal configurations, ensuring consistent batching and sequence packing regardless of the underlying architecture.