# How MiniMind Implements Mixture of Experts (MoE): Architecture and Code Deep Dive

> Discover how MiniMind implements Mixture of Experts MoE. Explore its architecture, code, gating network, expert routing, and load balancing loss for efficient AI models.

- Repository: [jingyaogong/minimind](https://github.com/jingyaogong/minimind)
- Tags: deep-dive
- Published: 2026-03-24

---

**TLDR:** MiniMind implements Mixture of Experts through a conditional `MOEFeedForward` module activated by the `use_moe` config flag, consisting of a gating network (`MoEGate`) that routes tokens to top-k experts via Softmax scoring, separate routed and shared expert lists, and an auxiliary loss for load balancing aggregated across transformer layers.

The `jingyaogong/minimind` repository provides a lightweight, educational implementation of Large Language Model architectures, featuring a complete **Mixture of Experts (MoE)** system that scales model capacity without linearly increasing computational cost. Unlike dense feed-forward networks, MiniMind’s MoE design routes each token to specialized expert sub-networks, activated through a learned gating mechanism configured via `MiniMindConfig`.

## Configuration and Activation

MoE behavior in MiniMind is controlled entirely through the `MiniMindConfig` class defined in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) (lines 29‑40). Setting `use_moe=True` switches the architecture from dense to sparse activation.

Key **MoE-specific hyper-parameters** include:

- `n_routed_experts`: Total number of expert networks available for routing
- `n_shared_experts`: Number of experts applied to every token unconditionally
- `num_experts_per_tok`: How many routed experts process each token (top-k selection)
- `scoring_func`: Gating normalization method (typically `'softmax'`)
- `aux_loss_alpha`: Weighting factor for the load-balancing auxiliary loss

When `use_moe` is enabled, the transformer stack automatically replaces standard feed-forward layers with the sparse MoE variant.

## The Gating Mechanism (MoEGate)

The `MoEGate` class in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) (lines 32‑86) serves as the routing controller. It projects input hidden states to a **gating dimension** equal to `hidden_size`, then produces a probability distribution over all routed experts.

During the forward pass, `MoEGate`:

1. Computes logits via a lightweight linear layer
2. Applies Softmax to generate routing probabilities
3. Selects the `topk` experts per token based on `num_experts_per_tok`
4. Optionally normalizes the top-k probabilities
5. Computes an **auxiliary loss** that penalizes uneven expert utilization, encouraging balanced load across the expert pool

## Expert Feed-Forward Architecture (MOEFeedForward)

The `MOEFeedForward` class (lines 88‑148 in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)) encapsulates the actual expert computation. It maintains:

- A `ModuleList` of routed experts (`n_routed_experts` instances of standard `FeedForward` blocks)
- An optional list of shared experts (`n_shared_experts`) applied to every token
- The `MoEGate` instance that manages routing decisions

### Training Mode Implementation

During training (`self.training == True`), the module **repeats** each input token `num_experts_per_tok` times. This allows parallel processing where the repeated token copies are dispatched to their respective selected experts, and outputs are gathered and weighted by the top-k probabilities.

### Inference Optimization

During inference (`self.training == False`), MiniMind switches to the `moe_infer` method (lines 128‑148), which groups tokens by their assigned expert indices before processing. This eliminates redundant tensor copies and memory overhead, making autoregressive generation efficient despite the sparse architecture.

## Block-Level Integration

Integration occurs at the transformer block level in `MiniMindBlock` (line 63, [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)):

```python
self.mlp = FeedForward(config) if not config.use_moe else MOEFeedForward(config)

```

This conditional instantiation ensures that when `use_moe=True`, every transformer layer automatically employs the sparse `MOEFeedForward` while the attention mechanism, RMSNorm layers, and residual connections remain unchanged. The rest of the architecture ( Rotary Position Embeddings, RMSNorm, etc.) requires no modification to support MoE.

## Load Balancing via Auxiliary Loss

To prevent **expert collapse** (where a few experts dominate all traffic), MiniMind computes a layer-wise auxiliary loss inside `MoEGate`. After the full forward pass, `MiniMindModel` aggregates these values (line 23, [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)):

```python
aux_loss = sum(layer.mlp.aux_loss for layer in self.layers if hasattr(layer.mlp, 'aux_loss'))

```

This summed loss is returned alongside the standard language modeling loss, allowing the optimizer to penalize routing imbalance explicitly during backpropagation.

## End-to-End Execution Flow

The complete Mixture of Experts execution proceeds as follows:

1. **Initialization**: Configure `MiniMindConfig(use_moe=True, num_experts_per_tok=2, n_routed_experts=4, ...)`
2. **Per-Layer Forward Pass**:
   - Hidden states enter `MOEFeedForward`
   - `MoEGate` computes top-k expert indices and routing weights
   - Tokens are either duplicated (training) or grouped by expert (inference)
   - Selected routed experts process their assigned token subsets
   - Outputs are weighted, summed, and combined with shared expert contributions
   - Auxiliary loss is accumulated on the layer
3. **Model Output**: `MiniMindModel` returns `(hidden_states, presents, aux_loss)`; `MiniMindForCausalLM` combines the auxiliary loss with the cross-entropy loss for training.

## Practical Code Examples

### Enabling MoE in MiniMind Configuration

```python
from model.model_minimind import MiniMindForCausalLM, MiniMindConfig
import torch

# Initialize model with 4 routed experts, 2 experts per token, and 1 shared expert

config = MiniMindConfig(
    use_moe=True,
    n_routed_experts=4,
    n_shared_experts=1,
    num_experts_per_tok=2,
    scoring_func='softmax',
    aux_loss_alpha=0.01,
    seq_aux=True,
)

model = MiniMindForCausalLM(config)
input_ids = torch.randint(0, config.vocab_size, (2, 8))  # Batch size 2, sequence 8

# Forward pass returns logits and auxiliary loss

output = model(input_ids=input_ids, use_cache=False)
print(f"Logits shape: {output.logits.shape}")
print(f"Auxiliary loss: {output.aux_loss.item()}")

```

### Inspecting Gating Decisions During Training

```python
model.train()  # Enable training mode (aux loss computed)

outputs = model(input_ids=input_ids, output_hidden_states=True)

# Access the first layer's gate

gate = model.model.layers[0].mlp.gate
hidden = outputs.hidden_states[0]  # First layer hidden states

topk_idx, topk_weight, aux_loss = gate(hidden)
print(f"Expert assignments: {topk_idx.shape}")  # (batch*seq, num_experts_per_tok)

print(f"Routing weights: {topk_weight.shape}")

```

### Optimized Inference Mode

```python
model.eval()  # Switches to moe_infer path

with torch.no_grad():
    outputs = model(input_ids=input_ids)
    

# Auxiliary loss is zero during inference as load balancing is not required

logits = outputs.logits

```

## Summary

- **Conditional Activation**: MoE is enabled via `use_moe=True` in `MiniMindConfig`, switching `MiniMindBlock` to use `MOEFeedForward` instead of dense layers.
- **Sparse Routing**: The `MoEGate` network in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) performs top-k expert selection and computes load-balancing auxiliary loss.
- **Expert Types**: MiniMind supports both **routed experts** (token-specific) and **shared experts** (applied to all tokens) within the same layer.
- **Dual Mode Execution**: Training repeats tokens for parallel expert processing, while inference uses the `moe_infer` method to group tokens by expert for memory efficiency.
- **Load Balancing**: Layer-wise auxiliary losses are aggregated in `MiniMindModel` to prevent expert collapse and ensure uniform utilization.

## Frequently Asked Questions

### How do I activate Mixture of Experts in MiniMind?

Set `use_moe=True` in the `MiniMindConfig` initialization along with expert counts (`n_routed_experts`, `n_shared_experts`) and the top-k value (`num_experts_per_tok`). According to the source code in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) at line 63, `MiniMindBlock` automatically instantiates `MOEFeedForward` when this flag is enabled.

### What is the difference between routed and shared experts in MiniMind's MoE?

**Routed experts** are selected dynamically by the `MoEGate` network based on hidden state content, while **shared experts** process every token unconditionally. This architecture allows shared experts to handle common linguistic patterns while routed experts specialize in specific knowledge domains.

### How does MiniMind optimize MoE routing during inference versus training?

During training, `MOEFeedForward` duplicates each token `num_experts_per_tok` times to facilitate parallel expert processing. During inference, the `moe_infer` method (lines 128‑148 in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)) groups tokens by their assigned expert indices first, eliminating redundant memory copies and improving generation speed.

### What is the purpose of the auxiliary loss in MiniMind's MoE implementation?

The auxiliary loss penalizes imbalanced usage of routed experts by encouraging the `MoEGate` to distribute tokens evenly across the expert pool. Computed per layer and aggregated in `MiniMindModel` (line 23), this loss prevents expert collapse and ensures the full capacity of the sparse network is utilized during training.