How MegaDLMs Implements Mixture of Experts (MoE) Architectures

MegaDLMs supports Mixture of Experts by replacing standard MLP blocks with configurable MoE layers that use expert parallelism, Top-K or Expert-Choice routing, and optimized token dispatchers like All-to-All communication.

MegaDLMs extends the Megatron-LM framework with native Mixture of Experts (MoE) support, allowing transformer models to scale to trillions of parameters by conditionally activating only a subset of experts per token. This implementation enables seamless integration of MoE architectures into existing GPT-style models through configuration-driven activation and modular layer design.

Configuration-Driven MoE Activation

MoE support in MegaDLMs is controlled entirely through TransformerConfig in megatron/core/transformer/transformer_config.py. Setting the num_moe_experts parameter to a non-null value triggers the global replacement of standard MLP layers with MoE variants.

Key configuration fields include:

  • num_moe_experts: Total number of experts in the MoE layer (must be divisible by expert parallel size)
  • moe_token_dispatcher_type: Communication backend—"allgather", "alltoall", or "alltoall_seq"
  • moe_shared_expert_intermediate_size: Enables an optional shared expert alongside routed experts
  • moe_shared_expert_overlap: Allows overlapping shared expert computation with token dispatching
  • moe_layer_recompute: Enables activation checkpointing for MoE layers to reduce memory usage

When num_moe_experts is configured, the model specification utilities in megatron/core/models/gpt/gpt_layer_specs.py automatically substitute MoESubmodules for the standard MLPSubmodules during layer construction.

Core MoE Layer Architecture

The MoE implementation centers on three classes in megatron/core/transformer/moe/moe_layer.py:

BaseMoELayer

BaseMoELayer (lines 30-61) handles the foundational setup for expert parallelism. It partitions the global expert pool across the expert parallel group, calculating self.local_expert_indices so each GPU knows which experts it owns. It also initializes shared expert parameters when moe_shared_expert_intermediate_size is specified.

MoELayer

MoELayer implements the standard Top-K routing strategy. It instantiates a TopKRouter and a token dispatcher based on the configuration. The forward pass executes a precise three-step pipeline:

  1. Permutation: token_dispatcher.token_permutation() routes tokens to their assigned experts
  2. Expert Computation: Local experts process the dispatched tokens
  3. Unpermutation: token_dispatcher.token_unpermutation() reconstructs the sequence from expert outputs

MoELayerExpertChoice

MoELayerExpertChoice (lines 165+) uses ExpertChoiceRouter instead of Top-K, implementing the alternative routing strategy where experts select tokens rather than tokens selecting experts. This variant is particularly effective for load balancing in certain training scenarios.

Routing Mechanisms: Top-K vs Expert Choice

MegaDLMs provides two distinct routing implementations in megatron/core/transformer/moe/router.py:

TopKRouter computes dense gate logits through a learned linear layer, applies softmax, and selects the top-k experts for each token. It returns probability weights and a routing map used by the dispatcher to shuffle tokens.

ExpertChoiceRouter inverts the selection process: each expert chooses its preferred tokens based on the routing scores. This approach naturally balances load across experts and is used exclusively by MoELayerExpertChoice.

Both routers respect the num_moe_experts configuration and handle the mathematical gating logic that determines which experts process which tokens during the forward pass.

Token Dispatching and Communication

The token_dispatcher.py module implements three communication strategies for moving tokens between GPUs during the permutation phase:

MoEAllGatherTokenDispatcher uses an all-gather operation followed by a scatter, suitable for configurations where tensor_model_parallel_size == 1. This is the simplest but least bandwidth-efficient approach.

MoEAlltoAllTokenDispatcher performs an all-to-all exchange of token buckets, maximizing bandwidth efficiency when scaling to large expert counts across many GPUs.

MoEAlltoAllSEQTokenDispatcher provides sequence-aware all-to-all communication that overlaps with computation when moe_shared_expert_overlap is enabled, hiding communication latency behind the shared expert forward pass.

Each dispatcher implements token_permutation() to distribute tokens to experts and token_unpermutation() to reconstruct the output sequence, handling the complex tensor reshaping required for expert parallelism.

Integration with Transformer Specifications

MoE integration occurs at the model specification level in megatron/core/models/gpt/gpt_layer_specs.py. When config.num_moe_experts is set, the specification logic replaces the standard MLPSubmodules with MoESubmodules:

if config.num_moe_experts:
    submodules.mlp = MoESubmodules(
        experts=MLPSubmodules(...),
        shared_experts=... if config.moe_shared_expert_intermediate_size else None,
    )

The build_module utility in megatron/core/transformer/spec_utils.py then constructs the actual expert modules from these specifications, creating local expert instances based on the expert parallel distribution calculated by BaseMoELayer.

Activation Checkpointing and Constraints

MegaDLMs provides memory optimization for MoE training through the moe_layer_recompute configuration flag. When enabled, both MoELayer and MoELayerExpertChoice wrap their forward computations with tensor_parallel.checkpoint, trading computation for memory by recomputing the expert forward passes during backward propagation.

Current implementation constraints include:

  • Sequence Parallelism Compatibility: MoE can only be combined with sequence parallelism when tensor_model_parallel_size > 1, enforced by validation checks in the layer initialization.
  • No Token Dropping: Unlike some MoE implementations, MegaDLMs does not support token dropping—every token must be processed by its selected expert(s), ensuring deterministic computation but potentially limiting load balancing optimizations.

Summary

  • Configuration-driven activation: Set num_moe_experts in TransformerConfig to globally replace MLP layers with MoE architectures.
  • Modular routing: Choose between TopKRouter for standard token-to-expert selection or ExpertChoiceRouter for expert-to-token selection via MoELayerExpertChoice.
  • Efficient communication: Three dispatcher implementations (AllGather, AlltoAll, AlltoAllSEQ) optimize token movement across expert parallel groups.
  • Seamless integration: gpt_layer_specs.py automatically substitutes MoESubmodules for MLPSubmodules when MoE is enabled, requiring no manual layer surgery.
  • Memory optimization: Enable moe_layer_recompute for activation checkpointing, with constraints on sequence parallelism compatibility and no token dropping support.

Frequently Asked Questions

What is the difference between Top-K and Expert-Choice routing in MegaDLMs?

Top-K routing uses the TopKRouter class to compute routing probabilities via a learned gate, allowing each token to select its top-k preferred experts. Expert-Choice routing inverts this relationship using ExpertChoiceRouter, where each expert selects the tokens it wants to process based on routing scores. Expert-Choice naturally balances load across experts and is implemented in the MoELayerExpertChoice class, while standard MoELayer uses Top-K.

How does expert parallelism work in MegaDLMs MoE layers?

Expert parallelism distributes the total pool of experts across multiple GPUs according to the expert parallel world size. The BaseMoELayer class calculates local_expert_indices for each GPU, determining which subset of experts reside locally. During the forward pass, the token dispatcher routes tokens to the appropriate GPUs based on routing decisions, and only local experts process their assigned tokens. This sharding enables scaling to hundreds of experts without replicating every expert on every GPU.

Can I use sequence parallelism with MoE in MegaDLMs?

Sequence parallelism can be used with MoE, but only when tensor_model_parallel_size > 1. The implementation includes validation checks that enforce this constraint during layer initialization. If you enable sequence parallelism with tensor model parallelism equal to 1, the MoE layer will raise a configuration error. This limitation ensures correct gradient synchronization across the expert parallel groups when sequence parallelism is active.

What token dispatcher should I choose for my MoE model?

Choose based on your parallelism configuration and performance requirements. Use MoEAllGatherTokenDispatcher only when tensor_model_parallel_size == 1 and you prefer simplicity over bandwidth efficiency. For most large-scale training scenarios with many experts, MoEAlltoAllTokenDispatcher provides optimal bandwidth utilization via all-to-all communication. If you are using shared experts and want to hide communication latency, select MoEAlltoAllSEQTokenDispatcher, which overlaps the all-to-all exchange with shared expert computation when moe_shared_expert_overlap is enabled.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →