How to Configure Training Arguments for DLM Models in MegaDLMs
Configure DLM training arguments in MegaDLMs by extending Megatron-LM's base parser with the extra_args_provider function from custom_args/difflm.py, then access unified settings via get_args() in your training scripts.
MegaDLMs builds on the Megatron-LM framework to train Diffusion Language Models (DLM). While standard Megatron arguments control model parallelism and optimization, DLM-specific behaviors—such as masking strategies, attention configurations, and validation modes—require specialized configuration. This guide explains how to set up both base and DLM-specific training arguments using command-line interfaces and YAML files.
Understanding the MegaDLMs Argument Architecture
MegaDLMs uses a two-layer argument system that separates generic distributed training parameters from diffusion-specific model behaviors.
Base Megatron-LM Arguments
The foundation resides in megatron/training/arguments.py, which defines standard flags for model size, data paths, tensor parallelism, pipeline parallelism, and optimizer settings. These arguments apply to any Megatron-based training run and handle hardware configuration, checkpointing, and basic hyperparameters.
DLM-Specific Extensions
DLM-specific arguments are injected through custom_args/difflm.py via the extra_args_provider function. This function returns an argparse.ArgumentParser containing flags such as --model-running-mode, --mask-token, and --attention-mask-type. When pretrain_difflm.py invokes megatron.training.pretrain, it passes this provider to merge DLM flags into the base namespace.
# pretrain_difflm.py (excerpt)
from custom_args.difflm import extra_args_provider
pretrain(
...
extra_args_provider=extra_args_provider, # Injects DLM-specific flags
)
Configuring DLM Training Arguments via Command Line
You can specify all training parameters directly when launching your distributed training job. The parser automatically recognizes both Megatron base flags and DLM extensions.
# Example launch (single-node, 8 GPUs)
torchrun --standalone --nproc_per_node=8 \
pretrain_difflm.py \
--model-type difflm \
--model-running-mode difflm-noshift \
--base-model Qwen_2 \
--mask-token 151656 \
--attention-mask-type causal_bottom_right \
--core-attn-implementation TEDotProductAttention \
--difflm-varilen-prob 0.01 \
--use-difflm-validation \
--micro-batch-size 4 \
--global-batch-size 64 \
--seq-length 4096 \
--data-path /path/to/data \
--save /output/checkpoints
Key DLM flags in this example include:
--model-running-mode: Controls the diffusion execution strategy (e.g.,difflm-noshiftfor standard diffusion without position shifting).--mask-token: Specifies the token ID used for masking during diffusion training.--attention-mask-type: Defines the attention pattern (e.g.,causal_bottom_rightfor causal masking).--core-attn-implementation: Selects the attention backend (e.g.,TEDotProductAttentionfor Transformer Engine).--difflm-varilen-prob: Sets the probability for variable-length diffusion sampling.--use-difflm-validation: Enables DLM-specific validation logic during training.
Using YAML Configuration Files for DLM Training
For complex experiments with many hyperparameters, MegaDLMs supports YAML-based configuration via the --yaml_cfg argument. This approach improves reproducibility and version control.
Create a configuration file dlm_config.yaml:
model_type: difflm
model_running_mode: difflm-noshift
base_model: Qwen_2
mask_token: 151656
attention_mask_type: causal_bottom_right
core_attn_implementation: TEDotProductAttention
difflm_varilen_prob: 0.01
use_difflm_validation: true
micro_batch_size: 4
global_batch_size: 64
seq_length: 4096
data_path: /path/to/data
save: /output/checkpoints
Launch training with the YAML configuration:
torchrun --standalone --nproc_per_node=8 \
pretrain_difflm.py \
--yaml_cfg dlm_config.yaml
The yaml_arguments.py module processes the YAML file and merges its contents with the argument namespace. When yaml_arguments.core_transformer_config_from_yaml builds the TransformerConfig, DLM flags remain accessible through the same args namespace used for command-line arguments.
Accessing and Validating Arguments in Your Code
Once parsed, all configuration values are accessible via get_args() from megatron.training. This function returns a SimpleNamespace containing both base Megatron and DLM-specific attributes.
from megatron.training import get_args
def dump_args():
args = get_args()
for k, v in vars(args).items():
print(f'{k}: {v}')
# Called from model_provider or forward_step after initialization
During model construction in pretrain_difflm.py, the model_provider function uses these arguments to instantiate the DLM architecture:
if args.base_model == "vanilla":
model = GPTModel(
config=config,
transformer_layer_spec=transformer_layer_spec,
vocab_size=args.padded_vocab_size,
max_sequence_length=args.max_position_embeddings,
# DLM-specific options propagate through transformer_layer_spec
)
else:
raise NotImplementedError("Models other than vanilla GPT models are not supported yet.")
The transformer_layer_spec (defined in megatron/core/models/difflm/gpt_layer_specs.py) receives DLM flags such as attn_mask_type and core_attn_implementation to configure the attention mechanism appropriately.
During the forward pass, the training loop checks args.model_running_mode_curr to determine how to post-process outputs and slice the loss_mask based on diffusion-specific masking logic.
Key DLM Configuration Parameters Explained
Understanding these critical flags helps optimize your Diffusion Language Model training:
-
--model-running-mode– Selects the diffusion execution strategy. Options likedifflm-noshiftenable standard diffusion training without position shifting, affecting how the model processes token positions during the diffusion steps. -
--mask-token– The specific token ID used for masking during the diffusion process. This token replaces original tokens during the noising phase of diffusion training. -
--attention-mask-type– Controls the attention pattern.causal_bottom_rightprovides causal masking suitable for autoregressive diffusion, while other patterns may support different diffusion variants. -
--core-attn-implementation– Specifies the attention backend implementation.TEDotProductAttentionleverages Transformer Engine for optimized attention computation on modern GPUs. -
--difflm-varilen-prob– Probability for variable-length diffusion sampling, enabling training with variable sequence lengths to improve generalization. -
--use-difflm-validation– Boolean flag enabling DLM-specific validation logic that evaluates diffusion-specific metrics during training checkpoints.
Summary
-
MegaDLMs extends Megatron-LM's argument parser through
custom_args/difflm.pyto support DLM-specific training configurations. -
Use
extra_args_providerwhen callingmegatron.training.pretraininpretrain_difflm.pyto inject diffusion-specific flags into the base argument namespace. -
Access all configuration values via
get_args(), which returns aSimpleNamespacecontaining both standard Megatron and DLM-specific settings likemodel_running_modeandmask_token. -
Configure training via command-line arguments or YAML files using
--yaml_cfg, withyaml_arguments.pyhandling the merge logic while preserving DLM flag accessibility. -
Key DLM parameters control diffusion execution mode, masking tokens, attention implementations, and validation strategies, all consumed by
gpt_layer_specs.pyand the training loop to construct the appropriate model architecture and forward pass behavior.
Frequently Asked Questions
What is the difference between base Megatron arguments and DLM-specific arguments?
Base Megatron arguments, defined in megatron/training/arguments.py, handle generic distributed training concerns such as tensor parallelism, pipeline parallelism, optimizer settings, and data paths. DLM-specific arguments, defined in custom_args/difflm.py, control diffusion-specific behaviors including the masking token ID, attention mask types for diffusion, variable length sampling probabilities, and model running modes that determine how the diffusion process executes during training.
How do I add custom arguments for my DLM experiment?
To add custom arguments, extend the extra_args_provider function in custom_args/difflm.py by adding new arguments to the parser object using group.add_argument(). For example, add a flag like --my-custom-flag with appropriate type and default values. After modifying the provider, the new argument will automatically appear in the unified namespace returned by get_args() and can be accessed throughout your training scripts, model providers, and forward steps.
Can I mix command-line arguments with YAML configuration?
Yes, MegaDLMs supports mixing both methods. When you specify --yaml_cfg path/to/config.yaml alongside command-line flags, the YAML loader in megatron/training/yaml_arguments.py first processes the YAML file, then command-line arguments override any duplicate keys in the configuration. This allows you to maintain base configurations in YAML while quickly overriding specific settings like batch size or learning rate via command-line flags for individual experimental runs.
Where are the DLM arguments actually used during training?
DLM arguments are consumed at multiple stages of the training pipeline. During model construction in pretrain_difflm.py, the model_provider function passes DLM flags like attention_mask_type and core_attn_implementation to gpt_layer_specs.py to build the appropriate transformer layer specifications. During the forward pass, the training loop in megatron/training/training.py checks args.model_running_mode_curr to determine how to process the loss mask and handle diffusion-specific post-processing of model outputs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →