# How MiniMind's MoE Routing Mechanism Works: From Gate Scores to Expert Dispatch

> Explore MiniMind's MoE routing mechanism. Learn how gate scores select top-k experts per token and dispatch representations with auxiliary loss for uniform utilization.

- Repository: [jingyaogong/minimind](https://github.com/jingyaogong/minimind)
- Tags: deep-dive
- Published: 2026-03-24

---

**MiniMind's Mixture-of-Experts routing mechanism uses a learned gating network to compute softmax probabilities over expert feed-forward networks, selects the top-k experts per token, and dispatches representations to selected experts with an auxiliary load-balancing loss to ensure uniform utilization.**

The jingyaogong/minimind repository implements a sparse Mixture-of-Experts (MoE) transformer architecture that routes tokens dynamically through specialized feed-forward networks. This **MiniMind MoE routing mechanism** balances computational efficiency with model capacity by activating only a subset of experts per token. The implementation spans configuration management, gating logic, and efficient dispatch strategies for both training and inference modes.

## Architecture Components

The MoE system consists of three interconnected components defined in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py).

### Configuration in MiniMindConfig

The `MiniMindConfig` class (lines 28-38) defines all hyperparameters governing the routing behavior. Key fields include `use_moe` to enable the architecture, `num_experts_per_tok` (top-k value), `n_routed_experts` (total expert pool size), `scoring_func` (typically 'softmax'), and `norm_topk_prob` for optional probability normalization. The configuration also specifies `aux_loss_alpha` to control the load-balancing loss weight during training.

### The MoEGate Module

The `MoEGate` class (lines 32-86) implements the routing decision logic. Its `forward` method accepts hidden states of shape *(B, S, H)*, projects them to expert logits via `F.linear(hidden_states, self.weight, None)` (line 55), and computes selection probabilities. When `scoring_func` is set to `'softmax'`, the logits transform into probability distributions using `scores = logits.softmax(dim=-1)` (line 56).

### MOEFeedForward Dispatch

The `MOEFeedForward` class (lines 88-127) manages expert execution. It maintains a list of `FeedForward` expert modules and implements two dispatch strategies: a training path with explicit token duplication and an optimized inference path via `moe_infer`. The `MiniMindBlock` integrates this component when `config.use_moe` is `True`, substituting the standard MLP with `MOEFeedForward(config)` (line 63).

## The Routing Pipeline Step-by-Step

MiniMind's routing follows a deterministic sequence from score computation to expert dispatch.

### 1. Computing Expert Scores with Learned Projection

The gating mechanism begins by flattening input hidden states from *(B, S, H)* to *(B·S, H)*. The `MoEGate` module multiplies these states by a learnable weight matrix to produce logits for every expert in the pool. This linear projection happens at line 55: `logits = F.linear(hidden_states, self.weight, None)`, generating a score vector of dimension `n_routed_experts` for each token position.

### 2. Top-k Expert Selection via torch.topk

After applying softmax to obtain probability distributions (line 56: `scores = logits.softmax(dim=-1)`), the router selects the most appropriate experts using `torch.topk`. The operation at line 60—`topk_weight, topk_idx = torch.topk(scores, k=self.top_k, dim=-1, sorted=False)`—returns the `num_experts_per_tok` highest probabilities and their corresponding expert indices for every token in the batch.

### 3. Optional Probability Normalization

When `norm_topk_prob=True`, the selected top-k probabilities undergo normalization to sum to 1.0. This rescaling occurs in lines 62-65 of the gate implementation, ensuring that the contribution weights for the selected experts form a proper convex combination before being passed to the dispatch logic.

### 4. Auxiliary Load-Balancing Loss Calculation

To prevent expert collapse—where a few experts dominate all tokens—the mechanism computes an auxiliary loss encouraging uniform utilization. Lines 66-83 implement two variants: sequence-level balancing (`seq_aux=True`) and token-level balancing (`seq_aux=False`). The loss aggregates across layers in `MiniMindModel` (line 23: `aux_loss = sum([...])`) and adds to the primary language modeling objective during backpropagation.

## Training vs. Inference Dispatch Strategies

MiniMind employs distinct routing implementations optimized for each computational phase.

### Training Path: Explicit Token Duplication

During training, `MOEFeedForward.forward` (lines 13-18) duplicates each token `num_experts_per_tok` times using `x = x.repeat_interleave(self.config.num_experts_per_tok, dim=0)`. Tokens route to their assigned experts via boolean masking (`flat_topk_idx == i`), where each expert processes its subset and outputs are weighted by `topk_weight` and summed. This explicit routing ensures gradient flow to all selected experts but incurs memory overhead from token replication.

### Inference Path: Sorted Batch Processing

For inference, the `moe_infer` method (lines 29-48) eliminates duplication by sorting tokens according to their assigned expert indices. The implementation uses `flat_expert_indices.argsort()` to group tokens by expert, processes them in contiguous batches via `expert(x[exp_token_idx])`, and scatters results back using `scatter_add_`. This batch-by-expert approach maximizes GPU utilization and minimizes memory overhead during generation.

## Implementation Example

The following code demonstrates routing behavior in a MiniMind model with MoE enabled:

```python
import torch
from model.model_minimind import MiniMindConfig, MiniMindForCausalLM

# Configure MoE with 4 experts, selecting top-2 per token

cfg = MiniMindConfig(
    use_moe=True,
    hidden_size=512,
    num_hidden_layers=2,
    num_experts_per_tok=2,
    n_routed_experts=4,
    n_shared_experts=0,
    aux_loss_alpha=0.01,
)

model = MiniMindForCausalLM(cfg).eval()
input_ids = torch.randint(0, cfg.vocab_size, (1, 4))

# Capture routing decisions

with torch.no_grad():
    outputs = model(input_ids, logits_to_keep=0)

# Inspect first layer gate outputs

gate = model.model.layers[0].mlp.gate
hidden = model.model.embed_tokens(input_ids).view(-1, cfg.hidden_size)
topk_idx, topk_weight, aux_loss = gate(hidden)

print("Expert indices per token:", topk_idx)  # Shape: (B*S, 2)

print("Routing weights:", topk_weight)       # Shape: (B*S, 2)

print("Load-balancing loss:", aux_loss.item())

```

The output displays which experts process each token, their normalized contribution weights, and the auxiliary loss value that optimizes expert utilization during training.

## Summary

- **MiniMind's MoE routing mechanism** relies on a learned `MoEGate` that projects hidden states to expert logits and applies softmax to generate routing probabilities.
- The system selects `num_experts_per_tok` experts per token using `torch.topk`, with optional probability normalization controlled by `norm_topk_prob`.
- An auxiliary load-balancing loss prevents expert collapse, calculated either at the sequence or token level depending on `seq_aux`.
- Training uses explicit token duplication for gradient flow, while inference employs sorted batching via `moe_infer` for efficiency.
- Configuration occurs through `MiniMindConfig` in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py), with seamless integration into the transformer blocks via `MOEFeedForward`.

## Frequently Asked Questions

### What determines how many experts process each token in MiniMind?

The `num_experts_per_tok` parameter in `MiniMindConfig` controls the top-k selection count. During the forward pass in `MoEGate`, `torch.topk` selects exactly this many experts from the total pool defined by `n_routed_experts`, typically defaulting to 2-4 experts per token while maintaining a larger total pool of 8-64 experts.

### How does MiniMind prevent expert collapse during MoE training?

The implementation computes an auxiliary load-balancing loss (lines 66-83 of [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)) that penalizes uneven expert utilization. This loss aggregates across all layers and adds to the main training objective with weight `aux_loss_alpha`, forcing the gate to distribute tokens uniformly across the expert pool rather than routing everything to a single high-performing expert.

### Why does MiniMind use different routing paths for training and inference?

Training requires `MOEFeedForward.forward` to duplicate tokens explicitly (line 13: `repeat_interleave`) to maintain differentiable gradients through all selected experts. Inference uses `moe_infer` (lines 29-48) which sorts tokens by expert assignment and processes them in batches without duplication, reducing memory overhead and improving throughput during autoregressive generation.

### Where is the MoE routing logic configured in the MiniMind codebase?

All routing configuration resides in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py). The `MiniMindConfig` class (lines 28-38) defines MoE hyperparameters, `MoEGate` (lines 32-86) implements the selection logic, and `MOEFeedForward` (lines 88-127) handles expert dispatch. Training scripts in `trainer/` expose these parameters via command-line arguments like `--use_moe` for end-to-end configuration.