How MiniMind's MoE Routing Mechanism Works: From Gate Scores to Expert Dispatch

MiniMind's Mixture-of-Experts routing mechanism uses a learned gating network to compute softmax probabilities over expert feed-forward networks, selects the top-k experts per token, and dispatches representations to selected experts with an auxiliary load-balancing loss to ensure uniform utilization.

The jingyaogong/minimind repository implements a sparse Mixture-of-Experts (MoE) transformer architecture that routes tokens dynamically through specialized feed-forward networks. This MiniMind MoE routing mechanism balances computational efficiency with model capacity by activating only a subset of experts per token. The implementation spans configuration management, gating logic, and efficient dispatch strategies for both training and inference modes.

Architecture Components

The MoE system consists of three interconnected components defined in model/model_minimind.py.

Configuration in MiniMindConfig

The MiniMindConfig class (lines 28-38) defines all hyperparameters governing the routing behavior. Key fields include use_moe to enable the architecture, num_experts_per_tok (top-k value), n_routed_experts (total expert pool size), scoring_func (typically 'softmax'), and norm_topk_prob for optional probability normalization. The configuration also specifies aux_loss_alpha to control the load-balancing loss weight during training.

The MoEGate Module

The MoEGate class (lines 32-86) implements the routing decision logic. Its forward method accepts hidden states of shape (B, S, H), projects them to expert logits via F.linear(hidden_states, self.weight, None) (line 55), and computes selection probabilities. When scoring_func is set to 'softmax', the logits transform into probability distributions using scores = logits.softmax(dim=-1) (line 56).

MOEFeedForward Dispatch

The MOEFeedForward class (lines 88-127) manages expert execution. It maintains a list of FeedForward expert modules and implements two dispatch strategies: a training path with explicit token duplication and an optimized inference path via moe_infer. The MiniMindBlock integrates this component when config.use_moe is True, substituting the standard MLP with MOEFeedForward(config) (line 63).

The Routing Pipeline Step-by-Step

MiniMind's routing follows a deterministic sequence from score computation to expert dispatch.

1. Computing Expert Scores with Learned Projection

The gating mechanism begins by flattening input hidden states from (B, S, H) to (B·S, H). The MoEGate module multiplies these states by a learnable weight matrix to produce logits for every expert in the pool. This linear projection happens at line 55: logits = F.linear(hidden_states, self.weight, None), generating a score vector of dimension n_routed_experts for each token position.

2. Top-k Expert Selection via torch.topk

After applying softmax to obtain probability distributions (line 56: scores = logits.softmax(dim=-1)), the router selects the most appropriate experts using torch.topk. The operation at line 60—topk_weight, topk_idx = torch.topk(scores, k=self.top_k, dim=-1, sorted=False)—returns the num_experts_per_tok highest probabilities and their corresponding expert indices for every token in the batch.

3. Optional Probability Normalization

When norm_topk_prob=True, the selected top-k probabilities undergo normalization to sum to 1.0. This rescaling occurs in lines 62-65 of the gate implementation, ensuring that the contribution weights for the selected experts form a proper convex combination before being passed to the dispatch logic.

4. Auxiliary Load-Balancing Loss Calculation

To prevent expert collapse—where a few experts dominate all tokens—the mechanism computes an auxiliary loss encouraging uniform utilization. Lines 66-83 implement two variants: sequence-level balancing (seq_aux=True) and token-level balancing (seq_aux=False). The loss aggregates across layers in MiniMindModel (line 23: aux_loss = sum([...])) and adds to the primary language modeling objective during backpropagation.

Training vs. Inference Dispatch Strategies

MiniMind employs distinct routing implementations optimized for each computational phase.

Training Path: Explicit Token Duplication

During training, MOEFeedForward.forward (lines 13-18) duplicates each token num_experts_per_tok times using x = x.repeat_interleave(self.config.num_experts_per_tok, dim=0). Tokens route to their assigned experts via boolean masking (flat_topk_idx == i), where each expert processes its subset and outputs are weighted by topk_weight and summed. This explicit routing ensures gradient flow to all selected experts but incurs memory overhead from token replication.

Inference Path: Sorted Batch Processing

For inference, the moe_infer method (lines 29-48) eliminates duplication by sorting tokens according to their assigned expert indices. The implementation uses flat_expert_indices.argsort() to group tokens by expert, processes them in contiguous batches via expert(x[exp_token_idx]), and scatters results back using scatter_add_. This batch-by-expert approach maximizes GPU utilization and minimizes memory overhead during generation.

Implementation Example

The following code demonstrates routing behavior in a MiniMind model with MoE enabled:

import torch
from model.model_minimind import MiniMindConfig, MiniMindForCausalLM

# Configure MoE with 4 experts, selecting top-2 per token

cfg = MiniMindConfig(
    use_moe=True,
    hidden_size=512,
    num_hidden_layers=2,
    num_experts_per_tok=2,
    n_routed_experts=4,
    n_shared_experts=0,
    aux_loss_alpha=0.01,
)

model = MiniMindForCausalLM(cfg).eval()
input_ids = torch.randint(0, cfg.vocab_size, (1, 4))

# Capture routing decisions

with torch.no_grad():
    outputs = model(input_ids, logits_to_keep=0)

# Inspect first layer gate outputs

gate = model.model.layers[0].mlp.gate
hidden = model.model.embed_tokens(input_ids).view(-1, cfg.hidden_size)
topk_idx, topk_weight, aux_loss = gate(hidden)

print("Expert indices per token:", topk_idx)  # Shape: (B*S, 2)

print("Routing weights:", topk_weight)       # Shape: (B*S, 2)

print("Load-balancing loss:", aux_loss.item())

The output displays which experts process each token, their normalized contribution weights, and the auxiliary loss value that optimizes expert utilization during training.

Summary

  • MiniMind's MoE routing mechanism relies on a learned MoEGate that projects hidden states to expert logits and applies softmax to generate routing probabilities.
  • The system selects num_experts_per_tok experts per token using torch.topk, with optional probability normalization controlled by norm_topk_prob.
  • An auxiliary load-balancing loss prevents expert collapse, calculated either at the sequence or token level depending on seq_aux.
  • Training uses explicit token duplication for gradient flow, while inference employs sorted batching via moe_infer for efficiency.
  • Configuration occurs through MiniMindConfig in model/model_minimind.py, with seamless integration into the transformer blocks via MOEFeedForward.

Frequently Asked Questions

What determines how many experts process each token in MiniMind?

The num_experts_per_tok parameter in MiniMindConfig controls the top-k selection count. During the forward pass in MoEGate, torch.topk selects exactly this many experts from the total pool defined by n_routed_experts, typically defaulting to 2-4 experts per token while maintaining a larger total pool of 8-64 experts.

How does MiniMind prevent expert collapse during MoE training?

The implementation computes an auxiliary load-balancing loss (lines 66-83 of model/model_minimind.py) that penalizes uneven expert utilization. This loss aggregates across all layers and adds to the main training objective with weight aux_loss_alpha, forcing the gate to distribute tokens uniformly across the expert pool rather than routing everything to a single high-performing expert.

Why does MiniMind use different routing paths for training and inference?

Training requires MOEFeedForward.forward to duplicate tokens explicitly (line 13: repeat_interleave) to maintain differentiable gradients through all selected experts. Inference uses moe_infer (lines 29-48) which sorts tokens by expert assignment and processes them in batches without duplication, reducing memory overhead and improving throughput during autoregressive generation.

Where is the MoE routing logic configured in the MiniMind codebase?

All routing configuration resides in model/model_minimind.py. The MiniMindConfig class (lines 28-38) defines MoE hyperparameters, MoEGate (lines 32-86) implements the selection logic, and MOEFeedForward (lines 88-127) handles expert dispatch. Training scripts in trainer/ expose these parameters via command-line arguments like --use_moe for end-to-end configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →