# How to Configure Training Arguments for DLM Models in MegaDLMs

> Configure DLM training arguments in MegaDLMs by extending Megatron-LM's base parser with extra_args_provider. Access unified settings via get_args() in your training scripts.

- Repository: [Jinjie Ni/megadlms](https://github.com/jinjieni/megadlms)
- Tags: how-to-guide
- Published: 2026-03-04

---

**Configure DLM training arguments in MegaDLMs by extending Megatron-LM's base parser with the `extra_args_provider` function from [`custom_args/difflm.py`](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py), then access unified settings via `get_args()` in your training scripts.**

MegaDLMs builds on the Megatron-LM framework to train Diffusion Language Models (DLM). While standard Megatron arguments control model parallelism and optimization, DLM-specific behaviors—such as masking strategies, attention configurations, and validation modes—require specialized configuration. This guide explains how to set up both base and DLM-specific training arguments using command-line interfaces and YAML files.

## Understanding the MegaDLMs Argument Architecture

MegaDLMs uses a two-layer argument system that separates generic distributed training parameters from diffusion-specific model behaviors.

### Base Megatron-LM Arguments

The foundation resides in [`megatron/training/arguments.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/arguments.py), which defines standard flags for model size, data paths, tensor parallelism, pipeline parallelism, and optimizer settings. These arguments apply to any Megatron-based training run and handle hardware configuration, checkpointing, and basic hyperparameters.

### DLM-Specific Extensions

DLM-specific arguments are injected through [`custom_args/difflm.py`](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py) via the `extra_args_provider` function. This function returns an `argparse.ArgumentParser` containing flags such as `--model-running-mode`, `--mask-token`, and `--attention-mask-type`. When [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py) invokes `megatron.training.pretrain`, it passes this provider to merge DLM flags into the base namespace.

```python

# pretrain_difflm.py (excerpt)

from custom_args.difflm import extra_args_provider

pretrain(
    ...
    extra_args_provider=extra_args_provider,  # Injects DLM-specific flags

)

```

## Configuring DLM Training Arguments via Command Line

You can specify all training parameters directly when launching your distributed training job. The parser automatically recognizes both Megatron base flags and DLM extensions.

```bash

# Example launch (single-node, 8 GPUs)

torchrun --standalone --nproc_per_node=8 \
    pretrain_difflm.py \
    --model-type difflm \
    --model-running-mode difflm-noshift \
    --base-model Qwen_2 \
    --mask-token 151656 \
    --attention-mask-type causal_bottom_right \
    --core-attn-implementation TEDotProductAttention \
    --difflm-varilen-prob 0.01 \
    --use-difflm-validation \
    --micro-batch-size 4 \
    --global-batch-size 64 \
    --seq-length 4096 \
    --data-path /path/to/data \
    --save /output/checkpoints

```

Key DLM flags in this example include:

- **`--model-running-mode`**: Controls the diffusion execution strategy (e.g., `difflm-noshift` for standard diffusion without position shifting).
- **`--mask-token`**: Specifies the token ID used for masking during diffusion training.
- **`--attention-mask-type`**: Defines the attention pattern (e.g., `causal_bottom_right` for causal masking).
- **`--core-attn-implementation`**: Selects the attention backend (e.g., `TEDotProductAttention` for Transformer Engine).
- **`--difflm-varilen-prob`**: Sets the probability for variable-length diffusion sampling.
- **`--use-difflm-validation`**: Enables DLM-specific validation logic during training.

## Using YAML Configuration Files for DLM Training

For complex experiments with many hyperparameters, MegaDLMs supports YAML-based configuration via the `--yaml_cfg` argument. This approach improves reproducibility and version control.

Create a configuration file [`dlm_config.yaml`](https://github.com/jinjieni/megadlms/blob/main/dlm_config.yaml):

```yaml
model_type: difflm
model_running_mode: difflm-noshift
base_model: Qwen_2
mask_token: 151656
attention_mask_type: causal_bottom_right
core_attn_implementation: TEDotProductAttention
difflm_varilen_prob: 0.01
use_difflm_validation: true
micro_batch_size: 4
global_batch_size: 64
seq_length: 4096
data_path: /path/to/data
save: /output/checkpoints

```

Launch training with the YAML configuration:

```bash
torchrun --standalone --nproc_per_node=8 \
    pretrain_difflm.py \
    --yaml_cfg dlm_config.yaml

```

The [`yaml_arguments.py`](https://github.com/jinjieni/megadlms/blob/main/yaml_arguments.py) module processes the YAML file and merges its contents with the argument namespace. When `yaml_arguments.core_transformer_config_from_yaml` builds the `TransformerConfig`, DLM flags remain accessible through the same `args` namespace used for command-line arguments.

## Accessing and Validating Arguments in Your Code

Once parsed, all configuration values are accessible via `get_args()` from `megatron.training`. This function returns a `SimpleNamespace` containing both base Megatron and DLM-specific attributes.

```python
from megatron.training import get_args

def dump_args():
    args = get_args()
    for k, v in vars(args).items():
        print(f'{k}: {v}')

# Called from model_provider or forward_step after initialization

```

During model construction in [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py), the `model_provider` function uses these arguments to instantiate the DLM architecture:

```python
if args.base_model == "vanilla":
    model = GPTModel(
        config=config,
        transformer_layer_spec=transformer_layer_spec,
        vocab_size=args.padded_vocab_size,
        max_sequence_length=args.max_position_embeddings,
        # DLM-specific options propagate through transformer_layer_spec

    )
else:
    raise NotImplementedError("Models other than vanilla GPT models are not supported yet.")

```

The `transformer_layer_spec` (defined in [`megatron/core/models/difflm/gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/difflm/gpt_layer_specs.py)) receives DLM flags such as `attn_mask_type` and `core_attn_implementation` to configure the attention mechanism appropriately.

During the forward pass, the training loop checks `args.model_running_mode_curr` to determine how to post-process outputs and slice the `loss_mask` based on diffusion-specific masking logic.

## Key DLM Configuration Parameters Explained

Understanding these critical flags helps optimize your Diffusion Language Model training:

- **`--model-running-mode`** – Selects the diffusion execution strategy. Options like `difflm-noshift` enable standard diffusion training without position shifting, affecting how the model processes token positions during the diffusion steps.

- **`--mask-token`** – The specific token ID used for masking during the diffusion process. This token replaces original tokens during the noising phase of diffusion training.

- **`--attention-mask-type`** – Controls the attention pattern. `causal_bottom_right` provides causal masking suitable for autoregressive diffusion, while other patterns may support different diffusion variants.

- **`--core-attn-implementation`** – Specifies the attention backend implementation. `TEDotProductAttention` leverages Transformer Engine for optimized attention computation on modern GPUs.

- **`--difflm-varilen-prob`** – Probability for variable-length diffusion sampling, enabling training with variable sequence lengths to improve generalization.

- **`--use-difflm-validation`** – Boolean flag enabling DLM-specific validation logic that evaluates diffusion-specific metrics during training checkpoints.

## Summary

- MegaDLMs extends Megatron-LM's argument parser through [`custom_args/difflm.py`](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py) to support DLM-specific training configurations.

- Use `extra_args_provider` when calling `megatron.training.pretrain` in [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py) to inject diffusion-specific flags into the base argument namespace.

- Access all configuration values via `get_args()`, which returns a `SimpleNamespace` containing both standard Megatron and DLM-specific settings like `model_running_mode` and `mask_token`.

- Configure training via command-line arguments or YAML files using `--yaml_cfg`, with [`yaml_arguments.py`](https://github.com/jinjieni/megadlms/blob/main/yaml_arguments.py) handling the merge logic while preserving DLM flag accessibility.

- Key DLM parameters control diffusion execution mode, masking tokens, attention implementations, and validation strategies, all consumed by [`gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/gpt_layer_specs.py) and the training loop to construct the appropriate model architecture and forward pass behavior.

## Frequently Asked Questions

### What is the difference between base Megatron arguments and DLM-specific arguments?

Base Megatron arguments, defined in [`megatron/training/arguments.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/arguments.py), handle generic distributed training concerns such as tensor parallelism, pipeline parallelism, optimizer settings, and data paths. DLM-specific arguments, defined in [`custom_args/difflm.py`](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py), control diffusion-specific behaviors including the masking token ID, attention mask types for diffusion, variable length sampling probabilities, and model running modes that determine how the diffusion process executes during training.

### How do I add custom arguments for my DLM experiment?

To add custom arguments, extend the `extra_args_provider` function in [`custom_args/difflm.py`](https://github.com/jinjieni/megadlms/blob/main/custom_args/difflm.py) by adding new arguments to the parser object using `group.add_argument()`. For example, add a flag like `--my-custom-flag` with appropriate type and default values. After modifying the provider, the new argument will automatically appear in the unified namespace returned by `get_args()` and can be accessed throughout your training scripts, model providers, and forward steps.

### Can I mix command-line arguments with YAML configuration?

Yes, MegaDLMs supports mixing both methods. When you specify `--yaml_cfg path/to/config.yaml` alongside command-line flags, the YAML loader in [`megatron/training/yaml_arguments.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/yaml_arguments.py) first processes the YAML file, then command-line arguments override any duplicate keys in the configuration. This allows you to maintain base configurations in YAML while quickly overriding specific settings like batch size or learning rate via command-line flags for individual experimental runs.

### Where are the DLM arguments actually used during training?

DLM arguments are consumed at multiple stages of the training pipeline. During model construction in [`pretrain_difflm.py`](https://github.com/jinjieni/megadlms/blob/main/pretrain_difflm.py), the `model_provider` function passes DLM flags like `attention_mask_type` and `core_attn_implementation` to [`gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/gpt_layer_specs.py) to build the appropriate transformer layer specifications. During the forward pass, the training loop in [`megatron/training/training.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/training/training.py) checks `args.model_running_mode_curr` to determine how to process the loss mask and handle diffusion-specific post-processing of model outputs.