DLM vs Autoregressive LM Training in MegaDLMs: Key Architectural Differences
MegaDLMs toggle between Diffusion Language Model (DLM) and Autoregressive Language Model (AR-LM) training via the --model-running-mode argument, switching between bidirectional masking with token corruption in difflm_forward and strict causal left-to-right conditioning in vanilla_forward.
The MegaDLMs framework unifies two distinct training paradigms—diffusion-based and autoregressive language modeling—within a single codebase. Understanding the differences between DLM and Autoregressive LM training in MegaDLMs is essential for configuring your experiments, as each mode fundamentally alters the attention mechanisms, input processing, and loss computation strategies.
Training Mode Selection and Forward Paths
The training regime is determined by the --model-running-mode argument defined in [custom_args/difflm.py](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py). This flag selects between two distinct forward implementations in [megatron/core/models/difflm/gpt_model.py](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/difflm/gpt_model.py):
"difflm-noshift"(default): Activates the diffusion forward path viaGPTModel.difflm_forward"vanilla": Activates the standard autoregressive path viaGPTModel.vanilla_forward
At the entry point of GPTModel.forward (lines 7‑8), the mode is stored in args.model_running_mode_curr, which gates all subsequent architectural decisions.
Attention Masking and Conditioning
The attention mechanisms differ radically between the two modes. DLM training can utilize bidirectional attention masks, including the special no_mask configuration where the mask is temporarily removed entirely (attention_mask = None) during the diffusion step, or block-causal variants. This allows the model to attend to both past and future tokens when reconstructing noised positions.
In contrast, AR-LM training enforces strict left-to-right conditioning through the causal_bottom_right mask type. The autoregressive mode guarantees that each position only attends to preceding tokens, maintaining the unidirectional dependency structure required for standard language modeling.
Input Corruption and Diffusion Process
The most visible difference occurs during input processing. In DLM mode, the difflm_forward_process method (lines 3‑25 of gpt_model.py) randomly masks a subset of tokens using self.args.mask_token, generating three key tensors:
noisy_batch: The corrupted input sequencemasked_indices(difflm_mask): Boolean indicators of which positions were maskedp_mask: The per-token masking probability
Autoregressive training receives the original token sequence without corruption, as the model must predict the next token given all previous tokens in the standard left-to-right manner.
Loss Computation and Weighting Strategies
Loss calculation diverges significantly between the two regimes. For DLM training, the cross-entropy loss is re-weighted by the diffusion mask and masking probability (lines 28‑30):
loss = loss * difflm_mask / p_mask
This scaling focuses gradients exclusively on the noised tokens, effectively ignoring positions that were not corrupted during the forward process.
AR-LM training computes standard language-model cross-entropy across all positions without additional scaling or masking, treating every token as a prediction target under the causal constraint.
Variable-Length Training and Packing Adjustments
DLM mode supports optional random truncation for variable-length training via the --difflm-varilen-prob argument. When use_varilen_data is enabled, the framework adjusts sequence packing through specialized methods:
update_packing_info_random_shrink: Reduces packed sequence lengths to reflect random truncationupdate_packing_info_shift_by_one: Adjusts packing boundaries for shifted sequences
These packing modifications occur during difflm_forward_process (lines 39‑60) and are unique to diffusion training. Autoregressive mode maintains fixed sequence lengths determined by --seq-length and uses standard GPTDataset configurations without packing adjustments.
Training Loop Implementation
The training entry point [pretrain_difflm.py](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py) handles mode-specific logic in the training step:
- For DLM mode, the code extracts
difflm_maskfrom the model output (lines 41‑44). Whenmodel_running_mode_curr == "difflm-noshift"and the attention mask type is'no_mask', the loss mask is cropped to match potentially truncated sequences (lines 45‑46). - For AR-LM mode, the training loop asserts that
difflm_maskisNone(lines 47‑48), confirming that no diffusion-specific tensors are produced.
Practical Code Examples
Launching a DLM Pre-training Run
Activate diffusion training with bidirectional attention and optional variable-length truncation:
torchrun --nproc_per_node=8 pretrain_difflm.py \
--model-running-mode difflm-noshift \
--attention-mask-type no_mask \
--difflm-varilen-prob 0.01 \
--seq-length 2048 \
...other arguments...
Launching an Autoregressive LM Run
Force strict causal conditioning without token corruption:
torchrun --nproc_per_node=8 pretrain_difflm.py \
--model-running-mode vanilla \
--attention-mask-type causal_bottom_right \
--seq-length 2048 \
...other arguments...
Inspecting the Forward Path Programmatically
Verify which mode is active and inspect diffusion-specific outputs:
from megatron.core.models.difflm.gpt_model import GPTModel
model = GPTModel(...) # instantiated via model_provider()
print(model.args.model_running_mode) # "difflm-noshift" or "vanilla"
# Forward pass (simplified)
output = model(
input_ids, position_ids, attention_mask,
labels=labels, packed_seq_params=None
)
if model.args.model_running_mode == "difflm-noshift":
loss, logits, difflm_mask = output
print("Diffusion mask shape:", difflm_mask.shape)
else:
loss, logits = output
print("No diffusion mask (AR mode).")
Manual Token Corruption for Debugging
Access the diffusion corruption process directly:
noisy_ids, mask, p_mask = model.difflm_forward_process(input_ids)
print("Masked token proportion per batch:", p_mask.mean().item())
Summary
- Mode Selection: Use
--model-running-modewith"difflm-noshift"for diffusion training or"vanilla"for autoregressive training, as defined incustom_args/difflm.py. - Attention: DLM supports bidirectional (
no_mask) or block-causal attention; AR-LM enforces strictcausal_bottom_rightmasking. - Input Processing: DLM applies random token masking via
difflm_forward_process; AR-LM uses uncorrupted sequences. - Loss Weighting: DLM scales loss by
difflm_mask / p_maskto focus on noised tokens; AR-LM computes standard cross-entropy over all positions. - Variable Length: DLM supports random truncation and packing adjustments via
--difflm-varilen-prob; AR-LM uses fixed sequence lengths. - Training Loop:
pretrain_difflm.pyhandlesdifflm_maskextraction for DLM and asserts its absence for AR-LM.
Frequently Asked Questions
How do I switch between DLM and Autoregressive training in MegaDLMs?
Set the --model-running-mode argument when launching pretrain_difflm.py. Use "difflm-noshift" to enable diffusion language model training with token corruption and bidirectional attention, or "vanilla" to enable standard autoregressive left-to-right training. This flag controls whether GPTModel invokes difflm_forward or vanilla_forward in megatron/core/models/difflm/gpt_model.py.
What attention mask should I use for each training mode?
For DLM training, you can use no_mask (fully bidirectional) or block-causal masks, allowing the model to attend to future tokens when reconstructing noised positions. For Autoregressive LM training, you must use causal_bottom_right to enforce strict left-to-right dependencies. The mask type is specified via --attention-mask-type and is validated in the training loop.
Why is the loss weighted differently in DLM compared to AR-LM?
DLM training weights the loss by difflm_mask / p_mask to concentrate gradients on the randomly masked tokens only, effectively ignoring uncorrupted positions. This reflects the diffusion objective of reconstructing noised inputs. AR-LM training computes standard cross-entropy across all positions because every token serves as a prediction target under the causal constraint.
Can I use variable-length sequences with Autoregressive LM training?
No. Variable-length training with random truncation (use_varilen_data) and packing adjustments (update_packing_info_random_shrink) is implemented only for DLM mode via the --difflm-varilen-prob argument. Autoregressive training in MegaDLMs currently requires fixed sequence lengths as specified by --seq-length.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →