# How MegaDLMs Implements Mixture of Experts (MoE) Architectures

> MegaDLMs implements Mixture of Experts (MoE) with expert parallelism, configurable routing, and optimized token dispatchers. Learn how MegaDLMs supports MoE architectures.

- Repository: [Jinjie Ni/megadlms](https://github.com/jinjieni/megadlms)
- Tags: deep-dive
- Published: 2026-03-04

---

**MegaDLMs supports Mixture of Experts by replacing standard MLP blocks with configurable MoE layers that use expert parallelism, Top-K or Expert-Choice routing, and optimized token dispatchers like All-to-All communication.**

MegaDLMs extends the Megatron-LM framework with native Mixture of Experts (MoE) support, allowing transformer models to scale to trillions of parameters by conditionally activating only a subset of experts per token. This implementation enables seamless integration of MoE architectures into existing GPT-style models through configuration-driven activation and modular layer design.

## Configuration-Driven MoE Activation

MoE support in MegaDLMs is controlled entirely through `TransformerConfig` in [`megatron/core/transformer/transformer_config.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/transformer/transformer_config.py). Setting the `num_moe_experts` parameter to a non-null value triggers the global replacement of standard MLP layers with MoE variants.

Key configuration fields include:

- **`num_moe_experts`**: Total number of experts in the MoE layer (must be divisible by expert parallel size)
- **`moe_token_dispatcher_type`**: Communication backend—`"allgather"`, `"alltoall"`, or `"alltoall_seq"`
- **`moe_shared_expert_intermediate_size`**: Enables an optional shared expert alongside routed experts
- **`moe_shared_expert_overlap`**: Allows overlapping shared expert computation with token dispatching
- **`moe_layer_recompute`**: Enables activation checkpointing for MoE layers to reduce memory usage

When `num_moe_experts` is configured, the model specification utilities in [`megatron/core/models/gpt/gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/gpt/gpt_layer_specs.py) automatically substitute `MoESubmodules` for the standard `MLPSubmodules` during layer construction.

## Core MoE Layer Architecture

The MoE implementation centers on three classes in [`megatron/core/transformer/moe/moe_layer.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/transformer/moe/moe_layer.py):

### BaseMoELayer

`BaseMoELayer` (lines 30-61) handles the foundational setup for **expert parallelism**. It partitions the global expert pool across the expert parallel group, calculating `self.local_expert_indices` so each GPU knows which experts it owns. It also initializes shared expert parameters when `moe_shared_expert_intermediate_size` is specified.

### MoELayer

`MoELayer` implements the standard Top-K routing strategy. It instantiates a `TopKRouter` and a token dispatcher based on the configuration. The forward pass executes a precise three-step pipeline:

1. **Permutation**: `token_dispatcher.token_permutation()` routes tokens to their assigned experts
2. **Expert Computation**: Local experts process the dispatched tokens
3. **Unpermutation**: `token_dispatcher.token_unpermutation()` reconstructs the sequence from expert outputs

### MoELayerExpertChoice

`MoELayerExpertChoice` (lines 165+) uses `ExpertChoiceRouter` instead of Top-K, implementing the alternative routing strategy where experts select tokens rather than tokens selecting experts. This variant is particularly effective for load balancing in certain training scenarios.

## Routing Mechanisms: Top-K vs Expert Choice

MegaDLMs provides two distinct routing implementations in [`megatron/core/transformer/moe/router.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/transformer/moe/router.py):

**TopKRouter** computes dense gate logits through a learned linear layer, applies softmax, and selects the top-k experts for each token. It returns probability weights and a routing map used by the dispatcher to shuffle tokens.

**ExpertChoiceRouter** inverts the selection process: each expert chooses its preferred tokens based on the routing scores. This approach naturally balances load across experts and is used exclusively by `MoELayerExpertChoice`.

Both routers respect the `num_moe_experts` configuration and handle the mathematical gating logic that determines which experts process which tokens during the forward pass.

## Token Dispatching and Communication

The [`token_dispatcher.py`](https://github.com/jinjieni/megadlms/blob/main/token_dispatcher.py) module implements three communication strategies for moving tokens between GPUs during the permutation phase:

**MoEAllGatherTokenDispatcher** uses an all-gather operation followed by a scatter, suitable for configurations where `tensor_model_parallel_size == 1`. This is the simplest but least bandwidth-efficient approach.

**MoEAlltoAllTokenDispatcher** performs an all-to-all exchange of token buckets, maximizing bandwidth efficiency when scaling to large expert counts across many GPUs.

**MoEAlltoAllSEQTokenDispatcher** provides sequence-aware all-to-all communication that overlaps with computation when `moe_shared_expert_overlap` is enabled, hiding communication latency behind the shared expert forward pass.

Each dispatcher implements `token_permutation()` to distribute tokens to experts and `token_unpermutation()` to reconstruct the output sequence, handling the complex tensor reshaping required for expert parallelism.

## Integration with Transformer Specifications

MoE integration occurs at the model specification level in [`megatron/core/models/gpt/gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/models/gpt/gpt_layer_specs.py). When `config.num_moe_experts` is set, the specification logic replaces the standard `MLPSubmodules` with `MoESubmodules`:

```python
if config.num_moe_experts:
    submodules.mlp = MoESubmodules(
        experts=MLPSubmodules(...),
        shared_experts=... if config.moe_shared_expert_intermediate_size else None,
    )

```

The `build_module` utility in [`megatron/core/transformer/spec_utils.py`](https://github.com/jinjieni/megadlms/blob/main/megatron/core/transformer/spec_utils.py) then constructs the actual expert modules from these specifications, creating local expert instances based on the expert parallel distribution calculated by `BaseMoELayer`.

## Activation Checkpointing and Constraints

MegaDLMs provides memory optimization for MoE training through the `moe_layer_recompute` configuration flag. When enabled, both `MoELayer` and `MoELayerExpertChoice` wrap their forward computations with `tensor_parallel.checkpoint`, trading computation for memory by recomputing the expert forward passes during backward propagation.

Current implementation constraints include:

- **Sequence Parallelism Compatibility**: MoE can only be combined with sequence parallelism when `tensor_model_parallel_size > 1`, enforced by validation checks in the layer initialization.
- **No Token Dropping**: Unlike some MoE implementations, MegaDLMs does not support token dropping—every token must be processed by its selected expert(s), ensuring deterministic computation but potentially limiting load balancing optimizations.

## Summary

- **Configuration-driven activation**: Set `num_moe_experts` in `TransformerConfig` to globally replace MLP layers with MoE architectures.
- **Modular routing**: Choose between `TopKRouter` for standard token-to-expert selection or `ExpertChoiceRouter` for expert-to-token selection via `MoELayerExpertChoice`.
- **Efficient communication**: Three dispatcher implementations (`AllGather`, `AlltoAll`, `AlltoAllSEQ`) optimize token movement across expert parallel groups.
- **Seamless integration**: [`gpt_layer_specs.py`](https://github.com/jinjieni/megadlms/blob/main/gpt_layer_specs.py) automatically substitutes `MoESubmodules` for `MLPSubmodules` when MoE is enabled, requiring no manual layer surgery.
- **Memory optimization**: Enable `moe_layer_recompute` for activation checkpointing, with constraints on sequence parallelism compatibility and no token dropping support.

## Frequently Asked Questions

### What is the difference between Top-K and Expert-Choice routing in MegaDLMs?

**Top-K routing** uses the `TopKRouter` class to compute routing probabilities via a learned gate, allowing each token to select its top-k preferred experts. **Expert-Choice routing** inverts this relationship using `ExpertChoiceRouter`, where each expert selects the tokens it wants to process based on routing scores. Expert-Choice naturally balances load across experts and is implemented in the `MoELayerExpertChoice` class, while standard `MoELayer` uses Top-K.

### How does expert parallelism work in MegaDLMs MoE layers?

Expert parallelism distributes the total pool of experts across multiple GPUs according to the expert parallel world size. The `BaseMoELayer` class calculates `local_expert_indices` for each GPU, determining which subset of experts reside locally. During the forward pass, the token dispatcher routes tokens to the appropriate GPUs based on routing decisions, and only local experts process their assigned tokens. This sharding enables scaling to hundreds of experts without replicating every expert on every GPU.

### Can I use sequence parallelism with MoE in MegaDLMs?

Sequence parallelism can be used with MoE, but only when `tensor_model_parallel_size > 1`. The implementation includes validation checks that enforce this constraint during layer initialization. If you enable sequence parallelism with tensor model parallelism equal to 1, the MoE layer will raise a configuration error. This limitation ensures correct gradient synchronization across the expert parallel groups when sequence parallelism is active.

### What token dispatcher should I choose for my MoE model?

Choose based on your parallelism configuration and performance requirements. Use **`MoEAllGatherTokenDispatcher`** only when `tensor_model_parallel_size == 1` and you prefer simplicity over bandwidth efficiency. For most large-scale training scenarios with many experts, **`MoEAlltoAllTokenDispatcher`** provides optimal bandwidth utilization via all-to-all communication. If you are using shared experts and want to hide communication latency, select **`MoEAlltoAllSEQTokenDispatcher`**, which overlaps the all-to-all exchange with shared expert computation when `moe_shared_expert_overlap` is enabled.