DLM vs Autoregressive LM Training in MegaDLMs: Key Architectural Differences

MegaDLMs toggle between Diffusion Language Model (DLM) and Autoregressive Language Model (AR-LM) training via the --model-running-mode argument, switching between bidirectional masking with token corruption in difflm_forward and strict causal left-to-right conditioning in vanilla_forward.

The MegaDLMs framework unifies two distinct training paradigms—diffusion-based and autoregressive language modeling—within a single codebase. Understanding the differences between DLM and Autoregressive LM training in MegaDLMs is essential for configuring your experiments, as each mode fundamentally alters the attention mechanisms, input processing, and loss computation strategies.

Training Mode Selection and Forward Paths

The training regime is determined by the --model-running-mode argument defined in [custom_args/difflm.py](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py). This flag selects between two distinct forward implementations in [megatron/core/models/difflm/gpt_model.py](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/difflm/gpt_model.py):

  • "difflm-noshift" (default): Activates the diffusion forward path via GPTModel.difflm_forward
  • "vanilla": Activates the standard autoregressive path via GPTModel.vanilla_forward

At the entry point of GPTModel.forward (lines 7‑8), the mode is stored in args.model_running_mode_curr, which gates all subsequent architectural decisions.

Attention Masking and Conditioning

The attention mechanisms differ radically between the two modes. DLM training can utilize bidirectional attention masks, including the special no_mask configuration where the mask is temporarily removed entirely (attention_mask = None) during the diffusion step, or block-causal variants. This allows the model to attend to both past and future tokens when reconstructing noised positions.

In contrast, AR-LM training enforces strict left-to-right conditioning through the causal_bottom_right mask type. The autoregressive mode guarantees that each position only attends to preceding tokens, maintaining the unidirectional dependency structure required for standard language modeling.

Input Corruption and Diffusion Process

The most visible difference occurs during input processing. In DLM mode, the difflm_forward_process method (lines 3‑25 of gpt_model.py) randomly masks a subset of tokens using self.args.mask_token, generating three key tensors:

  • noisy_batch: The corrupted input sequence
  • masked_indices (difflm_mask): Boolean indicators of which positions were masked
  • p_mask: The per-token masking probability

Autoregressive training receives the original token sequence without corruption, as the model must predict the next token given all previous tokens in the standard left-to-right manner.

Loss Computation and Weighting Strategies

Loss calculation diverges significantly between the two regimes. For DLM training, the cross-entropy loss is re-weighted by the diffusion mask and masking probability (lines 28‑30):

loss = loss * difflm_mask / p_mask

This scaling focuses gradients exclusively on the noised tokens, effectively ignoring positions that were not corrupted during the forward process.

AR-LM training computes standard language-model cross-entropy across all positions without additional scaling or masking, treating every token as a prediction target under the causal constraint.

Variable-Length Training and Packing Adjustments

DLM mode supports optional random truncation for variable-length training via the --difflm-varilen-prob argument. When use_varilen_data is enabled, the framework adjusts sequence packing through specialized methods:

  • update_packing_info_random_shrink: Reduces packed sequence lengths to reflect random truncation
  • update_packing_info_shift_by_one: Adjusts packing boundaries for shifted sequences

These packing modifications occur during difflm_forward_process (lines 39‑60) and are unique to diffusion training. Autoregressive mode maintains fixed sequence lengths determined by --seq-length and uses standard GPTDataset configurations without packing adjustments.

Training Loop Implementation

The training entry point [pretrain_difflm.py](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py) handles mode-specific logic in the training step:

  • For DLM mode, the code extracts difflm_mask from the model output (lines 41‑44). When model_running_mode_curr == "difflm-noshift" and the attention mask type is 'no_mask', the loss mask is cropped to match potentially truncated sequences (lines 45‑46).
  • For AR-LM mode, the training loop asserts that difflm_mask is None (lines 47‑48), confirming that no diffusion-specific tensors are produced.

Practical Code Examples

Launching a DLM Pre-training Run

Activate diffusion training with bidirectional attention and optional variable-length truncation:

torchrun --nproc_per_node=8 pretrain_difflm.py \
    --model-running-mode difflm-noshift \
    --attention-mask-type no_mask \
    --difflm-varilen-prob 0.01 \
    --seq-length 2048 \
    ...other arguments...

Launching an Autoregressive LM Run

Force strict causal conditioning without token corruption:

torchrun --nproc_per_node=8 pretrain_difflm.py \
    --model-running-mode vanilla \
    --attention-mask-type causal_bottom_right \
    --seq-length 2048 \
    ...other arguments...

Inspecting the Forward Path Programmatically

Verify which mode is active and inspect diffusion-specific outputs:

from megatron.core.models.difflm.gpt_model import GPTModel

model = GPTModel(...)  # instantiated via model_provider()

print(model.args.model_running_mode)  # "difflm-noshift" or "vanilla"

# Forward pass (simplified)

output = model(
    input_ids, position_ids, attention_mask,
    labels=labels, packed_seq_params=None
)

if model.args.model_running_mode == "difflm-noshift":
    loss, logits, difflm_mask = output
    print("Diffusion mask shape:", difflm_mask.shape)
else:
    loss, logits = output
    print("No diffusion mask (AR mode).")

Manual Token Corruption for Debugging

Access the diffusion corruption process directly:

noisy_ids, mask, p_mask = model.difflm_forward_process(input_ids)
print("Masked token proportion per batch:", p_mask.mean().item())

Summary

  • Mode Selection: Use --model-running-mode with "difflm-noshift" for diffusion training or "vanilla" for autoregressive training, as defined in custom_args/difflm.py.
  • Attention: DLM supports bidirectional (no_mask) or block-causal attention; AR-LM enforces strict causal_bottom_right masking.
  • Input Processing: DLM applies random token masking via difflm_forward_process; AR-LM uses uncorrupted sequences.
  • Loss Weighting: DLM scales loss by difflm_mask / p_mask to focus on noised tokens; AR-LM computes standard cross-entropy over all positions.
  • Variable Length: DLM supports random truncation and packing adjustments via --difflm-varilen-prob; AR-LM uses fixed sequence lengths.
  • Training Loop: pretrain_difflm.py handles difflm_mask extraction for DLM and asserts its absence for AR-LM.

Frequently Asked Questions

How do I switch between DLM and Autoregressive training in MegaDLMs?

Set the --model-running-mode argument when launching pretrain_difflm.py. Use "difflm-noshift" to enable diffusion language model training with token corruption and bidirectional attention, or "vanilla" to enable standard autoregressive left-to-right training. This flag controls whether GPTModel invokes difflm_forward or vanilla_forward in megatron/core/models/difflm/gpt_model.py.

What attention mask should I use for each training mode?

For DLM training, you can use no_mask (fully bidirectional) or block-causal masks, allowing the model to attend to future tokens when reconstructing noised positions. For Autoregressive LM training, you must use causal_bottom_right to enforce strict left-to-right dependencies. The mask type is specified via --attention-mask-type and is validated in the training loop.

Why is the loss weighted differently in DLM compared to AR-LM?

DLM training weights the loss by difflm_mask / p_mask to concentrate gradients on the randomly masked tokens only, effectively ignoring uncorrupted positions. This reflects the diffusion objective of reconstructing noised inputs. AR-LM training computes standard cross-entropy across all positions because every token serves as a prediction target under the causal constraint.

Can I use variable-length sequences with Autoregressive LM training?

No. Variable-length training with random truncation (use_varilen_data) and packing adjustments (update_packing_info_random_shrink) is implemented only for DLM mode via the --difflm-varilen-prob argument. Autoregressive training in MegaDLMs currently requires fixed sequence lengths as specified by --seq-length.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →