MiniMind MoE Architecture Configurable Parameters: Complete Technical Guide

MiniMind’s Mixture-of-Experts (MoE) architecture provides eight configurable parameters—use_moe, num_experts_per_tok, n_routed_experts, n_shared_experts, scoring_func, aux_loss_alpha, seq_aux, and norm_topk_prob—all centralized in the MiniMindConfig class to control expert routing density, auxiliary loss weighting, and computation flow.

MiniMind is a lightweight large language model implementation designed for efficient training and inference on limited hardware. Its optional MoE architecture activates sparse expert networks per token rather than using dense feed-forward layers, significantly reducing computational cost while maintaining model capacity. Understanding these configurable parameters enables precise control over the trade-off between inference speed and predictive performance.

Core MoE Configuration Parameters

The MiniMindConfig class defined in model/model_minimind.py (lines 32-40) exposes the following MoE-specific fields:

Parameter Type Default Description
use_moe bool False Master switch enabling the MoE architecture. When False, the model uses standard dense FeedForward layers.
num_experts_per_tok int 2 Top-k value determining how many routed experts process each token.
n_routed_experts int 4 Total number of routed experts available for selection by the gating network.
n_shared_experts int 1 Number of shared experts applied to every token unconditionally, added after MoE routing.
scoring_func str 'softmax' Gating network scoring function (currently supports 'softmax' only).
aux_loss_alpha float 0.01 Weight coefficient for the auxiliary load-balancing loss.
seq_aux bool True When True, computes auxiliary loss at the sequence level (averaged across tokens); otherwise per-token.
norm_topk_prob bool True Normalizes top-k gating probabilities to sum to 1 before weighting expert outputs.

These values initialize the MOEFeedForward and MoEGate modules when use_moe=True.

How Parameters Control MoE Execution

Gating and Expert Routing

The MoEGate class (implemented in model/model_minimind.py, lines 32-86) consumes the configuration to implement top-k routing. It uses num_experts_per_tok to select the highest-scoring experts from the pool of n_routed_experts, applies the scoring_func to generate probabilities, and respects norm_topk_prob to ensure the selected expert weights sum to unity.

The aux_loss_alpha and seq_aux parameters determine how the load-balancing loss is calculated. This auxiliary loss encourages uniform utilization across all routed experts, preventing collapse to a single expert. When seq_aux=True, the loss aggregates across the entire sequence rather than individual tokens, providing more stable gradients during training.

Feed-Forward Layer Selection

Within MiniMindBlock (lines 63-64 of model/model_minimind.py), the boolean use_moe parameter acts as a conditional switch. When enabled, the block instantiates MOEFeedForward instead of the standard FeedForward, inserting the sparse expert computation into the transformer layer. This architectural decision occurs at model construction time and remains fixed throughout training and inference.

Training Utilities and Checkpointing

The training infrastructure references these parameters for experiment tracking. In trainer/trainer_utils.py (lines 65-68), utility functions access num_experts_per_tok and n_routed_experts when constructing checkpoint filenames, ensuring MoE-specific configurations are preserved in model artifacts. Additionally, the command-line interface in scripts/serve_openai_api.py (lines 171-176) exposes the --use_moe flag, allowing runtime toggling of the architecture via the config object.

Practical Configuration Examples

Instantiating a Custom MoE Model

Configure a 145M parameter MoE variant with increased expert diversity:

from model.model_minimind import MiniMindConfig, MiniMindModel

cfg = MiniMindConfig(
    hidden_size=640,
    num_hidden_layers=8,
    use_moe=True,
    num_experts_per_tok=2,
    n_routed_experts=8,
    n_shared_experts=2,
    scoring_func='softmax',
    aux_loss_alpha=0.02,
    seq_aux=True,
    norm_topk_prob=True,
)

model = MiniMindModel(cfg)

Monitoring Auxiliary Loss During Training

Access the load-balancing loss computed by the MoE layer for custom loss scaling:

outputs = model(input_ids)

# Access the first transformer block's MoE auxiliary loss

aux_loss = model.layers[0].mlp.aux_loss
total_loss = ce_loss + aux_loss
total_loss.backward()

Disabling MoE for Dense Training

Revert to the dense 26M parameter baseline by disabling the MoE flag:

cfg = MiniMindConfig(use_moe=False)
model = MiniMindModel(cfg)

Command-Line Configuration

Enable MoE when launching the OpenAI-compatible API server:

python -m scripts.serve_openai_api \
    --hidden_size 640 \
    --num_hidden_layers 8 \
    --use_moe 1

Summary

  • MiniMindConfig centralizes all MoE parameters in model/model_minimind.py, providing a single source of truth for architecture configuration.
  • use_moe acts as the master switch, determining whether MiniMindBlock instantiates sparse MOEFeedForward or dense FeedForward layers.
  • Routing parameters (num_experts_per_tok, n_routed_experts, n_shared_experts) control the number of active experts per token and the total expert pool size.
  • Training parameters (aux_loss_alpha, seq_aux, norm_topk_prob) fine-tune the load-balancing behavior and loss computation strategy.

Frequently Asked Questions

What happens when use_moe is set to False?

When use_moe=False, the model ignores all other MoE-specific parameters and constructs standard dense transformer blocks using the FeedForward class. Each token processes through a single dense network rather than being routed to multiple expert networks, resulting in lower memory usage but reduced model capacity compared to the sparse variant.

How does num_experts_per_tok affect inference speed?

The num_experts_per_tok parameter (default 2) determines the sparsity of computation. Lower values reduce the number of expert networks activated per token, decreasing floating-point operations and improving latency. However, reducing this value below 2 may degrade model quality as each token receives input from fewer specialized experts.

What is the purpose of n_shared_experts versus n_routed_experts?

The n_shared_experts (default 1) represents experts applied to every token unconditionally, providing a stable baseline representation, while n_routed_experts (default 4) represents the pool from which the gating network dynamically selects experts per token. This hybrid approach combines the stability of universal computation with the efficiency of sparse, context-dependent processing.

Where is the auxiliary loss calculated in the codebase?

The auxiliary loss computation resides in the MoEGate class within model/model_minimind.py. This class references aux_loss_alpha to scale the load-balancing penalty and checks seq_aux to determine whether to average the loss across the sequence or compute it per token. The resulting loss value is stored in the MOEFeedForward instance and accessed during the backward pass through model.layers[i].mlp.aux_loss.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →