MiniMind MoE Architecture Configurable Parameters: Complete Technical Guide
MiniMind’s Mixture-of-Experts (MoE) architecture provides eight configurable parameters—use_moe, num_experts_per_tok, n_routed_experts, n_shared_experts, scoring_func, aux_loss_alpha, seq_aux, and norm_topk_prob—all centralized in the MiniMindConfig class to control expert routing density, auxiliary loss weighting, and computation flow.
MiniMind is a lightweight large language model implementation designed for efficient training and inference on limited hardware. Its optional MoE architecture activates sparse expert networks per token rather than using dense feed-forward layers, significantly reducing computational cost while maintaining model capacity. Understanding these configurable parameters enables precise control over the trade-off between inference speed and predictive performance.
Core MoE Configuration Parameters
The MiniMindConfig class defined in model/model_minimind.py (lines 32-40) exposes the following MoE-specific fields:
| Parameter | Type | Default | Description |
|---|---|---|---|
use_moe |
bool |
False |
Master switch enabling the MoE architecture. When False, the model uses standard dense FeedForward layers. |
num_experts_per_tok |
int |
2 |
Top-k value determining how many routed experts process each token. |
n_routed_experts |
int |
4 |
Total number of routed experts available for selection by the gating network. |
n_shared_experts |
int |
1 |
Number of shared experts applied to every token unconditionally, added after MoE routing. |
scoring_func |
str |
'softmax' |
Gating network scoring function (currently supports 'softmax' only). |
aux_loss_alpha |
float |
0.01 |
Weight coefficient for the auxiliary load-balancing loss. |
seq_aux |
bool |
True |
When True, computes auxiliary loss at the sequence level (averaged across tokens); otherwise per-token. |
norm_topk_prob |
bool |
True |
Normalizes top-k gating probabilities to sum to 1 before weighting expert outputs. |
These values initialize the MOEFeedForward and MoEGate modules when use_moe=True.
How Parameters Control MoE Execution
Gating and Expert Routing
The MoEGate class (implemented in model/model_minimind.py, lines 32-86) consumes the configuration to implement top-k routing. It uses num_experts_per_tok to select the highest-scoring experts from the pool of n_routed_experts, applies the scoring_func to generate probabilities, and respects norm_topk_prob to ensure the selected expert weights sum to unity.
The aux_loss_alpha and seq_aux parameters determine how the load-balancing loss is calculated. This auxiliary loss encourages uniform utilization across all routed experts, preventing collapse to a single expert. When seq_aux=True, the loss aggregates across the entire sequence rather than individual tokens, providing more stable gradients during training.
Feed-Forward Layer Selection
Within MiniMindBlock (lines 63-64 of model/model_minimind.py), the boolean use_moe parameter acts as a conditional switch. When enabled, the block instantiates MOEFeedForward instead of the standard FeedForward, inserting the sparse expert computation into the transformer layer. This architectural decision occurs at model construction time and remains fixed throughout training and inference.
Training Utilities and Checkpointing
The training infrastructure references these parameters for experiment tracking. In trainer/trainer_utils.py (lines 65-68), utility functions access num_experts_per_tok and n_routed_experts when constructing checkpoint filenames, ensuring MoE-specific configurations are preserved in model artifacts. Additionally, the command-line interface in scripts/serve_openai_api.py (lines 171-176) exposes the --use_moe flag, allowing runtime toggling of the architecture via the config object.
Practical Configuration Examples
Instantiating a Custom MoE Model
Configure a 145M parameter MoE variant with increased expert diversity:
from model.model_minimind import MiniMindConfig, MiniMindModel
cfg = MiniMindConfig(
hidden_size=640,
num_hidden_layers=8,
use_moe=True,
num_experts_per_tok=2,
n_routed_experts=8,
n_shared_experts=2,
scoring_func='softmax',
aux_loss_alpha=0.02,
seq_aux=True,
norm_topk_prob=True,
)
model = MiniMindModel(cfg)
Monitoring Auxiliary Loss During Training
Access the load-balancing loss computed by the MoE layer for custom loss scaling:
outputs = model(input_ids)
# Access the first transformer block's MoE auxiliary loss
aux_loss = model.layers[0].mlp.aux_loss
total_loss = ce_loss + aux_loss
total_loss.backward()
Disabling MoE for Dense Training
Revert to the dense 26M parameter baseline by disabling the MoE flag:
cfg = MiniMindConfig(use_moe=False)
model = MiniMindModel(cfg)
Command-Line Configuration
Enable MoE when launching the OpenAI-compatible API server:
python -m scripts.serve_openai_api \
--hidden_size 640 \
--num_hidden_layers 8 \
--use_moe 1
Summary
- MiniMindConfig centralizes all MoE parameters in
model/model_minimind.py, providing a single source of truth for architecture configuration. use_moeacts as the master switch, determining whetherMiniMindBlockinstantiates sparseMOEFeedForwardor denseFeedForwardlayers.- Routing parameters (
num_experts_per_tok,n_routed_experts,n_shared_experts) control the number of active experts per token and the total expert pool size. - Training parameters (
aux_loss_alpha,seq_aux,norm_topk_prob) fine-tune the load-balancing behavior and loss computation strategy.
Frequently Asked Questions
What happens when use_moe is set to False?
When use_moe=False, the model ignores all other MoE-specific parameters and constructs standard dense transformer blocks using the FeedForward class. Each token processes through a single dense network rather than being routed to multiple expert networks, resulting in lower memory usage but reduced model capacity compared to the sparse variant.
How does num_experts_per_tok affect inference speed?
The num_experts_per_tok parameter (default 2) determines the sparsity of computation. Lower values reduce the number of expert networks activated per token, decreasing floating-point operations and improving latency. However, reducing this value below 2 may degrade model quality as each token receives input from fewer specialized experts.
What is the purpose of n_shared_experts versus n_routed_experts?
The n_shared_experts (default 1) represents experts applied to every token unconditionally, providing a stable baseline representation, while n_routed_experts (default 4) represents the pool from which the gating network dynamically selects experts per token. This hybrid approach combines the stability of universal computation with the efficiency of sparse, context-dependent processing.
Where is the auxiliary loss calculated in the codebase?
The auxiliary loss computation resides in the MoEGate class within model/model_minimind.py. This class references aux_loss_alpha to scale the load-balancing penalty and checks seq_aux to determine whether to average the loss across the sequence or compute it per token. The resulting loss value is stored in the MOEFeedForward instance and accessed during the backward pass through model.layers[i].mlp.aux_loss.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →