# DLM vs Autoregressive LM Training in MegaDLMs: Key Architectural Differences

> Explore DLM vs Autoregressive LM training within MegaDLMs. Understand the key architectural differences in bidirectional masking versus causal conditioning and how MegaDLMs switch modes for optimal performance.

- Repository: [Jinjie Ni/megadlms](https://github.com/jinjieni/megadlms)
- Tags: deep-dive
- Published: 2026-03-04

---

**MegaDLMs toggle between Diffusion Language Model (DLM) and Autoregressive Language Model (AR-LM) training via the `--model-running-mode` argument, switching between bidirectional masking with token corruption in `difflm_forward` and strict causal left-to-right conditioning in `vanilla_forward`.**

The MegaDLMs framework unifies two distinct training paradigms—diffusion-based and autoregressive language modeling—within a single codebase. Understanding the differences between DLM and Autoregressive LM training in MegaDLMs is essential for configuring your experiments, as each mode fundamentally alters the attention mechanisms, input processing, and loss computation strategies.

## Training Mode Selection and Forward Paths

The training regime is determined by the **`--model-running-mode`** argument defined in [[`custom_args/difflm.py`](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py)](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py). This flag selects between two distinct forward implementations in [[`megatron/core/models/difflm/gpt_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/difflm/gpt_model.py)](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/difflm/gpt_model.py):

- **`"difflm-noshift"`** (default): Activates the diffusion forward path via `GPTModel.difflm_forward`
- **`"vanilla"`**: Activates the standard autoregressive path via `GPTModel.vanilla_forward`

At the entry point of `GPTModel.forward` (lines 7‑8), the mode is stored in `args.model_running_mode_curr`, which gates all subsequent architectural decisions.

## Attention Masking and Conditioning

The attention mechanisms differ radically between the two modes. **DLM training** can utilize bidirectional attention masks, including the special `no_mask` configuration where the mask is temporarily removed entirely (`attention_mask = None`) during the diffusion step, or block-causal variants. This allows the model to attend to both past and future tokens when reconstructing noised positions.

In contrast, **AR-LM training** enforces strict left-to-right conditioning through the **`causal_bottom_right`** mask type. The autoregressive mode guarantees that each position only attends to preceding tokens, maintaining the unidirectional dependency structure required for standard language modeling.

## Input Corruption and Diffusion Process

The most visible difference occurs during input processing. In DLM mode, the `difflm_forward_process` method (lines 3‑25 of [`gpt_model.py`](https://github.com/jinjieni/megadlms/blob/main/gpt_model.py)) randomly **masks a subset of tokens** using `self.args.mask_token`, generating three key tensors:

- `noisy_batch`: The corrupted input sequence
- `masked_indices` (`difflm_mask`): Boolean indicators of which positions were masked
- `p_mask`: The per-token masking probability

Autoregressive training receives the **original token sequence without corruption**, as the model must predict the next token given all previous tokens in the standard left-to-right manner.

## Loss Computation and Weighting Strategies

Loss calculation diverges significantly between the two regimes. For **DLM training**, the cross-entropy loss is **re-weighted by the diffusion mask and masking probability** (lines 28‑30):

```python
loss = loss * difflm_mask / p_mask

```

This scaling focuses gradients exclusively on the noised tokens, effectively ignoring positions that were not corrupted during the forward process.

**AR-LM training** computes standard language-model cross-entropy across **all positions** without additional scaling or masking, treating every token as a prediction target under the causal constraint.

## Variable-Length Training and Packing Adjustments

DLM mode supports **optional random truncation** for variable-length training via the `--difflm-varilen-prob` argument. When `use_varilen_data` is enabled, the framework adjusts sequence packing through specialized methods:

- `update_packing_info_random_shrink`: Reduces packed sequence lengths to reflect random truncation
- `update_packing_info_shift_by_one`: Adjusts packing boundaries for shifted sequences

These packing modifications occur during `difflm_forward_process` (lines 39‑60) and are unique to diffusion training. Autoregressive mode maintains fixed sequence lengths determined by `--seq-length` and uses standard `GPTDataset` configurations without packing adjustments.

## Training Loop Implementation

The training entry point [[`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py)](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py) handles mode-specific logic in the training step:

- For **DLM mode**, the code extracts `difflm_mask` from the model output (lines 41‑44). When `model_running_mode_curr == "difflm-noshift"` and the attention mask type is `'no_mask'`, the loss mask is **cropped** to match potentially truncated sequences (lines 45‑46).
- For **AR-LM mode**, the training loop asserts that `difflm_mask` is `None` (lines 47‑48), confirming that no diffusion-specific tensors are produced.

## Practical Code Examples

### Launching a DLM Pre-training Run

Activate diffusion training with bidirectional attention and optional variable-length truncation:

```bash
torchrun --nproc_per_node=8 pretrain_difflm.py \
    --model-running-mode difflm-noshift \
    --attention-mask-type no_mask \
    --difflm-varilen-prob 0.01 \
    --seq-length 2048 \
    ...other arguments...

```

### Launching an Autoregressive LM Run

Force strict causal conditioning without token corruption:

```bash
torchrun --nproc_per_node=8 pretrain_difflm.py \
    --model-running-mode vanilla \
    --attention-mask-type causal_bottom_right \
    --seq-length 2048 \
    ...other arguments...

```

### Inspecting the Forward Path Programmatically

Verify which mode is active and inspect diffusion-specific outputs:

```python
from megatron.core.models.difflm.gpt_model import GPTModel

model = GPTModel(...)  # instantiated via model_provider()

print(model.args.model_running_mode)  # "difflm-noshift" or "vanilla"

# Forward pass (simplified)

output = model(
    input_ids, position_ids, attention_mask,
    labels=labels, packed_seq_params=None
)

if model.args.model_running_mode == "difflm-noshift":
    loss, logits, difflm_mask = output
    print("Diffusion mask shape:", difflm_mask.shape)
else:
    loss, logits = output
    print("No diffusion mask (AR mode).")

```

### Manual Token Corruption for Debugging

Access the diffusion corruption process directly:

```python
noisy_ids, mask, p_mask = model.difflm_forward_process(input_ids)
print("Masked token proportion per batch:", p_mask.mean().item())

```

## Summary

- **Mode Selection**: Use `--model-running-mode` with `"difflm-noshift"` for diffusion training or `"vanilla"` for autoregressive training, as defined in [`custom_args/difflm.py`](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py).
- **Attention**: DLM supports bidirectional (`no_mask`) or block-causal attention; AR-LM enforces strict `causal_bottom_right` masking.
- **Input Processing**: DLM applies random token masking via `difflm_forward_process`; AR-LM uses uncorrupted sequences.
- **Loss Weighting**: DLM scales loss by `difflm_mask / p_mask` to focus on noised tokens; AR-LM computes standard cross-entropy over all positions.
- **Variable Length**: DLM supports random truncation and packing adjustments via `--difflm-varilen-prob`; AR-LM uses fixed sequence lengths.
- **Training Loop**: [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py) handles `difflm_mask` extraction for DLM and asserts its absence for AR-LM.

## Frequently Asked Questions

### How do I switch between DLM and Autoregressive training in MegaDLMs?

Set the `--model-running-mode` argument when launching [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py). Use `"difflm-noshift"` to enable diffusion language model training with token corruption and bidirectional attention, or `"vanilla"` to enable standard autoregressive left-to-right training. This flag controls whether `GPTModel` invokes `difflm_forward` or `vanilla_forward` in [`megatron/core/models/difflm/gpt_model.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/difflm/gpt_model.py).

### What attention mask should I use for each training mode?

For DLM training, you can use `no_mask` (fully bidirectional) or block-causal masks, allowing the model to attend to future tokens when reconstructing noised positions. For Autoregressive LM training, you must use `causal_bottom_right` to enforce strict left-to-right dependencies. The mask type is specified via `--attention-mask-type` and is validated in the training loop.

### Why is the loss weighted differently in DLM compared to AR-LM?

DLM training weights the loss by `difflm_mask / p_mask` to concentrate gradients on the randomly masked tokens only, effectively ignoring uncorrupted positions. This reflects the diffusion objective of reconstructing noised inputs. AR-LM training computes standard cross-entropy across all positions because every token serves as a prediction target under the causal constraint.

### Can I use variable-length sequences with Autoregressive LM training?

No. Variable-length training with random truncation (`use_varilen_data`) and packing adjustments (`update_packing_info_random_shrink`) is implemented only for DLM mode via the `--difflm-varilen-prob` argument. Autoregressive training in MegaDLMs currently requires fixed sequence lengths as specified by `--seq-length`.