How MiniMind Implements Mixture of Experts (MoE): Architecture and Code Deep Dive

TLDR: MiniMind implements Mixture of Experts through a conditional MOEFeedForward module activated by the use_moe config flag, consisting of a gating network (MoEGate) that routes tokens to top-k experts via Softmax scoring, separate routed and shared expert lists, and an auxiliary loss for load balancing aggregated across transformer layers.

The jingyaogong/minimind repository provides a lightweight, educational implementation of Large Language Model architectures, featuring a complete Mixture of Experts (MoE) system that scales model capacity without linearly increasing computational cost. Unlike dense feed-forward networks, MiniMind’s MoE design routes each token to specialized expert sub-networks, activated through a learned gating mechanism configured via MiniMindConfig.

Configuration and Activation

MoE behavior in MiniMind is controlled entirely through the MiniMindConfig class defined in model/model_minimind.py (lines 29‑40). Setting use_moe=True switches the architecture from dense to sparse activation.

Key MoE-specific hyper-parameters include:

  • n_routed_experts: Total number of expert networks available for routing
  • n_shared_experts: Number of experts applied to every token unconditionally
  • num_experts_per_tok: How many routed experts process each token (top-k selection)
  • scoring_func: Gating normalization method (typically 'softmax')
  • aux_loss_alpha: Weighting factor for the load-balancing auxiliary loss

When use_moe is enabled, the transformer stack automatically replaces standard feed-forward layers with the sparse MoE variant.

The Gating Mechanism (MoEGate)

The MoEGate class in model/model_minimind.py (lines 32‑86) serves as the routing controller. It projects input hidden states to a gating dimension equal to hidden_size, then produces a probability distribution over all routed experts.

During the forward pass, MoEGate:

  1. Computes logits via a lightweight linear layer
  2. Applies Softmax to generate routing probabilities
  3. Selects the topk experts per token based on num_experts_per_tok
  4. Optionally normalizes the top-k probabilities
  5. Computes an auxiliary loss that penalizes uneven expert utilization, encouraging balanced load across the expert pool

Expert Feed-Forward Architecture (MOEFeedForward)

The MOEFeedForward class (lines 88‑148 in model/model_minimind.py) encapsulates the actual expert computation. It maintains:

  • A ModuleList of routed experts (n_routed_experts instances of standard FeedForward blocks)
  • An optional list of shared experts (n_shared_experts) applied to every token
  • The MoEGate instance that manages routing decisions

Training Mode Implementation

During training (self.training == True), the module repeats each input token num_experts_per_tok times. This allows parallel processing where the repeated token copies are dispatched to their respective selected experts, and outputs are gathered and weighted by the top-k probabilities.

Inference Optimization

During inference (self.training == False), MiniMind switches to the moe_infer method (lines 128‑148), which groups tokens by their assigned expert indices before processing. This eliminates redundant tensor copies and memory overhead, making autoregressive generation efficient despite the sparse architecture.

Block-Level Integration

Integration occurs at the transformer block level in MiniMindBlock (line 63, model/model_minimind.py):

self.mlp = FeedForward(config) if not config.use_moe else MOEFeedForward(config)

This conditional instantiation ensures that when use_moe=True, every transformer layer automatically employs the sparse MOEFeedForward while the attention mechanism, RMSNorm layers, and residual connections remain unchanged. The rest of the architecture ( Rotary Position Embeddings, RMSNorm, etc.) requires no modification to support MoE.

Load Balancing via Auxiliary Loss

To prevent expert collapse (where a few experts dominate all traffic), MiniMind computes a layer-wise auxiliary loss inside MoEGate. After the full forward pass, MiniMindModel aggregates these values (line 23, model/model_minimind.py):

aux_loss = sum(layer.mlp.aux_loss for layer in self.layers if hasattr(layer.mlp, 'aux_loss'))

This summed loss is returned alongside the standard language modeling loss, allowing the optimizer to penalize routing imbalance explicitly during backpropagation.

End-to-End Execution Flow

The complete Mixture of Experts execution proceeds as follows:

  1. Initialization: Configure MiniMindConfig(use_moe=True, num_experts_per_tok=2, n_routed_experts=4, ...)
  2. Per-Layer Forward Pass:
    • Hidden states enter MOEFeedForward
    • MoEGate computes top-k expert indices and routing weights
    • Tokens are either duplicated (training) or grouped by expert (inference)
    • Selected routed experts process their assigned token subsets
    • Outputs are weighted, summed, and combined with shared expert contributions
    • Auxiliary loss is accumulated on the layer
  3. Model Output: MiniMindModel returns (hidden_states, presents, aux_loss); MiniMindForCausalLM combines the auxiliary loss with the cross-entropy loss for training.

Practical Code Examples

Enabling MoE in MiniMind Configuration

from model.model_minimind import MiniMindForCausalLM, MiniMindConfig
import torch

# Initialize model with 4 routed experts, 2 experts per token, and 1 shared expert

config = MiniMindConfig(
    use_moe=True,
    n_routed_experts=4,
    n_shared_experts=1,
    num_experts_per_tok=2,
    scoring_func='softmax',
    aux_loss_alpha=0.01,
    seq_aux=True,
)

model = MiniMindForCausalLM(config)
input_ids = torch.randint(0, config.vocab_size, (2, 8))  # Batch size 2, sequence 8

# Forward pass returns logits and auxiliary loss

output = model(input_ids=input_ids, use_cache=False)
print(f"Logits shape: {output.logits.shape}")
print(f"Auxiliary loss: {output.aux_loss.item()}")

Inspecting Gating Decisions During Training

model.train()  # Enable training mode (aux loss computed)

outputs = model(input_ids=input_ids, output_hidden_states=True)

# Access the first layer's gate

gate = model.model.layers[0].mlp.gate
hidden = outputs.hidden_states[0]  # First layer hidden states

topk_idx, topk_weight, aux_loss = gate(hidden)
print(f"Expert assignments: {topk_idx.shape}")  # (batch*seq, num_experts_per_tok)

print(f"Routing weights: {topk_weight.shape}")

Optimized Inference Mode

model.eval()  # Switches to moe_infer path

with torch.no_grad():
    outputs = model(input_ids=input_ids)
    

# Auxiliary loss is zero during inference as load balancing is not required

logits = outputs.logits

Summary

  • Conditional Activation: MoE is enabled via use_moe=True in MiniMindConfig, switching MiniMindBlock to use MOEFeedForward instead of dense layers.
  • Sparse Routing: The MoEGate network in model/model_minimind.py performs top-k expert selection and computes load-balancing auxiliary loss.
  • Expert Types: MiniMind supports both routed experts (token-specific) and shared experts (applied to all tokens) within the same layer.
  • Dual Mode Execution: Training repeats tokens for parallel expert processing, while inference uses the moe_infer method to group tokens by expert for memory efficiency.
  • Load Balancing: Layer-wise auxiliary losses are aggregated in MiniMindModel to prevent expert collapse and ensure uniform utilization.

Frequently Asked Questions

How do I activate Mixture of Experts in MiniMind?

Set use_moe=True in the MiniMindConfig initialization along with expert counts (n_routed_experts, n_shared_experts) and the top-k value (num_experts_per_tok). According to the source code in model/model_minimind.py at line 63, MiniMindBlock automatically instantiates MOEFeedForward when this flag is enabled.

What is the difference between routed and shared experts in MiniMind's MoE?

Routed experts are selected dynamically by the MoEGate network based on hidden state content, while shared experts process every token unconditionally. This architecture allows shared experts to handle common linguistic patterns while routed experts specialize in specific knowledge domains.

How does MiniMind optimize MoE routing during inference versus training?

During training, MOEFeedForward duplicates each token num_experts_per_tok times to facilitate parallel expert processing. During inference, the moe_infer method (lines 128‑148 in model/model_minimind.py) groups tokens by their assigned expert indices first, eliminating redundant memory copies and improving generation speed.

What is the purpose of the auxiliary loss in MiniMind's MoE implementation?

The auxiliary loss penalizes imbalanced usage of routed experts by encouraging the MoEGate to distribute tokens evenly across the expert pool. Computed per layer and aggregated in MiniMindModel (line 23), this loss prevents expert collapse and ensures the full capacity of the sparse network is utilized during training.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →