Parallelism Strategies Supported by MegaDLMs: The Complete Technical Guide
MegaDLMs leverages Megatron-Core to provide six orthogonal parallelism dimensions—Data Parallelism (DP), Tensor Parallelism (TP), Pipeline Parallelism (PP), Context Parallelism (CP), Expert Parallelism (EP), and Sequence Parallelism (SP)—that can be combined to scale diffusion language models across thousands of GPUs.
The MegaDLMs repository (jinjieni/megadlms) builds upon NVIDIA's Megatron-Core infrastructure to enable distributed training of diffusion language models at massive scale. Understanding the parallelism strategies supported by MegaDLMs is essential for optimizing memory usage, communication overhead, and throughput across multi-GPU and multi-node configurations.
Data Parallelism (DP)
Data Parallelism replicates the complete model on every GPU while partitioning the training data across workers. MegaDLMs supports both standard Distributed Data Parallel (DDP) and Fully-Sharded Data Parallel (FSDP) implementations.
In megatron/training/arguments.py, the framework defines flags to control DP behavior:
--data-parallel-sharding-strategy no_shardenables standard DDP where each GPU maintains a full copy of parameters and gradients.--use-custom-fsdpor--use-torch-fsdpactivates FSDP, which shards model parameters across data-parallel ranks to reduce per-GPU memory consumption.
The runtime automatically selects the appropriate PyTorch DDP or FSDP APIs based on these arguments, allowing seamless scaling from single-node to multi-node clusters without code modifications.
Tensor Parallelism (TP) and Sequence Parallelism (SP)
Tensor Parallelism splits individual layer weight matrices and computations across multiple GPUs, reducing the memory footprint per device and enabling larger hidden dimensions than would fit on a single accelerator. This is implemented in megatron/core/tensor_parallel/ (specifically layers.py and utils.py), which provides scatter/gather operations and all-reduce kernels.
When TP is enabled with --tensor-model-parallel-size > 1, MegaDLMs recommends simultaneously enabling Sequence Parallelism (SP) via the --sequence-parallel flag. SP overlaps communication with computation by distributing sequence chunks across TP ranks, maximizing hardware utilization. The combination of TP and SP is particularly effective for transformer-based diffusion models with large hidden sizes.
Pipeline Parallelism (PP)
Pipeline Parallelism distributes distinct model stages (contiguous blocks of layers) across GPUs, creating a processing pipeline where different micro-batches flow through different stages simultaneously. This approach reduces per-GPU memory by sharding parameters vertically through the network depth.
Core logic resides in megatron/core/pipeline_parallel/, including p2p_communication.py which handles point-to-point transfers between stages. MegaDLMs supports virtual pipeline stages via --virtual-pipeline-model-parallel-size, which further subdivides physical stages to improve load balancing and bubble reduction in the training pipeline.
Context Parallelism (CP)
Context Parallelism addresses the challenge of training on extremely long input sequences by distributing sequence chunks across GPUs. This strategy enables context lengths that would otherwise exceed the memory capacity of any single device.
Configuration occurs through:
--context-parallel-sizeto set the degree of sequence splitting--cp-comm-type p2pto select the communication backend--hierarchical-context-parallel-sizes 2 4to enable hierarchical CP for complex cluster topologies
The implementation in megatron/core/context_parallel/ provides sequence chunking utilities and optimized collective communication patterns to minimize latency when exchanging attention contexts across ranks.
Expert Parallelism (EP) for Mixture-of-Experts
Expert Parallelism is specific to Mixture-of-Experts (MoE) architectures, partitioning the set of experts across GPUs so each accelerator hosts only a subset of the expert parameters. This is critical for scaling MoE models with billions to trillions of total parameters.
Key flags defined in megatron/training/arguments.py and validated in megatron/core/model_parallel_config.py include:
--expert-model-parallel-size 4to set the EP degree--num-experts 8to specify experts per layer--moe-grouped-gemmto optimize expert computation
According to the MoE implementation in megatron/core/transformer/moe/, Expert Parallelism requires Sequence Parallelism when combined with Tensor Parallelism (EP + TP without SP is an invalid configuration that the framework explicitly rejects during argument validation).
Combining Parallelism Strategies
MegaDLMs supports arbitrary combinations of the above dimensions to create 3D parallelism or higher-order configurations. For example, you can simultaneously enable TP (intra-layer), PP (inter-layer), CP (sequence), and EP (expert) parallelism to maximize hardware utilization on large clusters.
The framework validates incompatible combinations in megatron/core/model_parallel_config.py and the training argument parser. For instance, attempting to use Expert Parallelism with Tensor Parallelism without enabling Sequence Parallelism will trigger a configuration error before training begins.
Configuration and Implementation Files
The following source files define and implement the parallelism strategies supported by MegaDLMs:
megatron/training/arguments.py— Defines all command-line flags (e.g.,--tensor-model-parallel-size,--pipeline-model-parallel-size,--context-parallel-size,--expert-model-parallel-size) and performs initial argument validation.megatron/core/model_parallel_config.py— Contains theModelParallelConfigclass that stores runtime settings for DP, TP, SP, PP, CP, and EP, enforcing compatibility rules between strategies.megatron/core/tensor_parallel/— Implements tensor-parallel operators including column/row parallel linear layers and sequence-parallel utilities.megatron/core/pipeline_parallel/— Handles pipeline stage scheduling, micro-batch partitioning, and virtual pipeline logic.megatron/core/context_parallel/— Provides context-parallel kernels for distributed attention over long sequences.megatron/core/transformer/moe/— Contains expert-parallel utilities and token dispatching mechanisms for MoE models.
Practical Configuration Examples
Below are minimal command-line configurations demonstrating how to enable each parallelism strategy when launching training with pretrain_difflm.py:
# Standard Data Parallel (DDP)
torchrun --nproc_per_node=8 pretrain_difflm.py \
--data-parallel-sharding-strategy no_shard
# Tensor Parallelism (4-way) + Sequence Parallelism
torchrun --nproc_per_node=8 pretrain_difflm.py \
--tensor-model-parallel-size 4 \
--sequence-parallel
# Pipeline Parallelism (8 stages) with Virtual Pipeline (4 stages)
torchrun --nproc_per_node=8 pretrain_difflm.py \
--pipeline-model-parallel-size 8 \
--virtual-pipeline-model-parallel-size 4
# Context Parallelism for long sequences
torchrun --nproc_per_node=8 pretrain_difflm.py \
--context-parallel-size 2 \
--cp-comm-type p2p \
--hierarchical-context-parallel-sizes 2 4
# Expert Parallelism (MoE) with 4-way EP and 8 experts
torchrun --nproc_per_node=8 pretrain_difflm.py \
--expert-model-parallel-size 4 \
--num-experts 8 \
--moe-grouped-gemm \
--sequence-parallel
# Full 3D Parallelism (TP + PP + EP + CP)
torchrun --nproc_per_node=8 pretrain_difflm.py \
--tensor-model-parallel-size 4 \
--pipeline-model-parallel-size 4 \
--expert-model-parallel-size 2 \
--context-parallel-size 2 \
--sequence-parallel
Summary
- MegaDLMs provides six orthogonal parallelism dimensions: Data Parallelism (DP), Tensor Parallelism (TP), Sequence Parallelism (SP), Pipeline Parallelism (PP), Context Parallelism (CP), and Expert Parallelism (EP).
- Configuration flags are centralized in
megatron/training/arguments.pyand validated inmegatron/core/model_parallel_config.py. - Tensor Parallelism requires Sequence Parallelism when combined with Expert Parallelism to ensure valid communication patterns.
- Context Parallelism enables training on sequences longer than single-GPU memory limits through hierarchical communication strategies.
- All strategies can be combined arbitrarily to match specific hardware topologies and model architectures, from single-node multi-GPU setups to thousand-GPU clusters.
Frequently Asked Questions
What is the difference between Tensor Parallelism and Pipeline Parallelism in MegaDLMs?
Tensor Parallelism splits individual layers horizontally across GPUs, dividing weight matrices and matrix multiplications to reduce per-GPU memory for large hidden dimensions. Pipeline Parallelism splits the model vertically by distributing contiguous blocks of layers (stages) across GPUs, allowing different micro-batches to be processed simultaneously at different depths of the network. TP is implemented in megatron/core/tensor_parallel/, while PP logic resides in megatron/core/pipeline_parallel/.
Can I use Expert Parallelism without Tensor Parallelism?
Yes, Expert Parallelism (EP) can be used independently or combined with other strategies. However, if you combine EP with Tensor Parallelism (TP), you must enable Sequence Parallelism (SP) via the --sequence-parallel flag. The framework validates this requirement in megatron/core/model_parallel_config.py and will raise a configuration error if you attempt EP+TP without SP.
Which parallelism strategy should I use to train on very long sequences?
For sequences that exceed the memory capacity of a single GPU, enable Context Parallelism (CP) using --context-parallel-size and optionally --hierarchical-context-parallel-sizes. CP divides the input sequence across GPUs and is specifically designed to support extremely long context lengths in diffusion language models without sacrificing batch size or hidden dimensions.
How do I choose the right combination of parallelism strategies for my hardware?
Select strategies based on your model size and cluster topology: use TP (2-8 GPUs) to fit large hidden dimensions in memory; add PP when layers exceed single-GPU capacity; add CP for long sequences; add EP for MoE models; and use DP (or FSDP) to scale to multiple nodes. The product of all parallelism degrees (TP × PP × CP × EP × DP) must equal your total GPU count, and MegaDLMs validates these constraints in megatron/training/arguments.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →