# MiniMind MoE Architecture Configurable Parameters: Complete Technical Guide

> Explore MiniMind's MoE architecture and its eight configurable parameters in this technical guide. Learn to control expert routing, auxiliary loss, and computation flow with MiniMindConfig.

- Repository: [jingyaogong/minimind](https://github.com/jingyaogong/minimind)
- Tags: technical-guide
- Published: 2026-03-24

---

**MiniMind’s Mixture-of-Experts (MoE) architecture provides eight configurable parameters—`use_moe`, `num_experts_per_tok`, `n_routed_experts`, `n_shared_experts`, `scoring_func`, `aux_loss_alpha`, `seq_aux`, and `norm_topk_prob`—all centralized in the `MiniMindConfig` class to control expert routing density, auxiliary loss weighting, and computation flow.**

MiniMind is a lightweight large language model implementation designed for efficient training and inference on limited hardware. Its optional **MoE architecture** activates sparse expert networks per token rather than using dense feed-forward layers, significantly reducing computational cost while maintaining model capacity. Understanding these **configurable parameters** enables precise control over the trade-off between inference speed and predictive performance.

## Core MoE Configuration Parameters

The `MiniMindConfig` class defined in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) (lines 32-40) exposes the following MoE-specific fields:

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `use_moe` | `bool` | `False` | Master switch enabling the MoE architecture. When `False`, the model uses standard dense `FeedForward` layers. |
| `num_experts_per_tok` | `int` | `2` | Top-k value determining how many routed experts process each token. |
| `n_routed_experts` | `int` | `4` | Total number of routed experts available for selection by the gating network. |
| `n_shared_experts` | `int` | `1` | Number of shared experts applied to every token unconditionally, added after MoE routing. |
| `scoring_func` | `str` | `'softmax'` | Gating network scoring function (currently supports `'softmax'` only). |
| `aux_loss_alpha` | `float` | `0.01` | Weight coefficient for the auxiliary load-balancing loss. |
| `seq_aux` | `bool` | `True` | When `True`, computes auxiliary loss at the sequence level (averaged across tokens); otherwise per-token. |
| `norm_topk_prob` | `bool` | `True` | Normalizes top-k gating probabilities to sum to 1 before weighting expert outputs. |

These values initialize the `MOEFeedForward` and `MoEGate` modules when `use_moe=True`.

## How Parameters Control MoE Execution

### Gating and Expert Routing

The `MoEGate` class (implemented in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py), lines 32-86) consumes the configuration to implement **top-k routing**. It uses `num_experts_per_tok` to select the highest-scoring experts from the pool of `n_routed_experts`, applies the `scoring_func` to generate probabilities, and respects `norm_topk_prob` to ensure the selected expert weights sum to unity.

The `aux_loss_alpha` and `seq_aux` parameters determine how the **load-balancing loss** is calculated. This auxiliary loss encourages uniform utilization across all routed experts, preventing collapse to a single expert. When `seq_aux=True`, the loss aggregates across the entire sequence rather than individual tokens, providing more stable gradients during training.

### Feed-Forward Layer Selection

Within `MiniMindBlock` (lines 63-64 of [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)), the boolean `use_moe` parameter acts as a conditional switch. When enabled, the block instantiates `MOEFeedForward` instead of the standard `FeedForward`, inserting the sparse expert computation into the transformer layer. This architectural decision occurs at model construction time and remains fixed throughout training and inference.

### Training Utilities and Checkpointing

The training infrastructure references these parameters for experiment tracking. In [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py) (lines 65-68), utility functions access `num_experts_per_tok` and `n_routed_experts` when constructing checkpoint filenames, ensuring MoE-specific configurations are preserved in model artifacts. Additionally, the command-line interface in [`scripts/serve_openai_api.py`](https://github.com/jingyaogong/minimind/blob/main/scripts/serve_openai_api.py) (lines 171-176) exposes the `--use_moe` flag, allowing runtime toggling of the architecture via the config object.

## Practical Configuration Examples

### Instantiating a Custom MoE Model

Configure a 145M parameter MoE variant with increased expert diversity:

```python
from model.model_minimind import MiniMindConfig, MiniMindModel

cfg = MiniMindConfig(
    hidden_size=640,
    num_hidden_layers=8,
    use_moe=True,
    num_experts_per_tok=2,
    n_routed_experts=8,
    n_shared_experts=2,
    scoring_func='softmax',
    aux_loss_alpha=0.02,
    seq_aux=True,
    norm_topk_prob=True,
)

model = MiniMindModel(cfg)

```

### Monitoring Auxiliary Loss During Training

Access the load-balancing loss computed by the MoE layer for custom loss scaling:

```python
outputs = model(input_ids)

# Access the first transformer block's MoE auxiliary loss

aux_loss = model.layers[0].mlp.aux_loss
total_loss = ce_loss + aux_loss
total_loss.backward()

```

### Disabling MoE for Dense Training

Revert to the dense 26M parameter baseline by disabling the MoE flag:

```python
cfg = MiniMindConfig(use_moe=False)
model = MiniMindModel(cfg)

```

### Command-Line Configuration

Enable MoE when launching the OpenAI-compatible API server:

```bash
python -m scripts.serve_openai_api \
    --hidden_size 640 \
    --num_hidden_layers 8 \
    --use_moe 1

```

## Summary

- **MiniMindConfig** centralizes all MoE parameters in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py), providing a single source of truth for architecture configuration.
- **`use_moe`** acts as the master switch, determining whether `MiniMindBlock` instantiates sparse `MOEFeedForward` or dense `FeedForward` layers.
- **Routing parameters** (`num_experts_per_tok`, `n_routed_experts`, `n_shared_experts`) control the number of active experts per token and the total expert pool size.
- **Training parameters** (`aux_loss_alpha`, `seq_aux`, `norm_topk_prob`) fine-tune the load-balancing behavior and loss computation strategy.

## Frequently Asked Questions

### What happens when `use_moe` is set to `False`?

When `use_moe=False`, the model ignores all other MoE-specific parameters and constructs standard dense transformer blocks using the `FeedForward` class. Each token processes through a single dense network rather than being routed to multiple expert networks, resulting in lower memory usage but reduced model capacity compared to the sparse variant.

### How does `num_experts_per_tok` affect inference speed?

The `num_experts_per_tok` parameter (default 2) determines the sparsity of computation. Lower values reduce the number of expert networks activated per token, decreasing floating-point operations and improving latency. However, reducing this value below 2 may degrade model quality as each token receives input from fewer specialized experts.

### What is the purpose of `n_shared_experts` versus `n_routed_experts`?

The `n_shared_experts` (default 1) represents experts applied to every token unconditionally, providing a stable baseline representation, while `n_routed_experts` (default 4) represents the pool from which the gating network dynamically selects experts per token. This hybrid approach combines the stability of universal computation with the efficiency of sparse, context-dependent processing.

### Where is the auxiliary loss calculated in the codebase?

The auxiliary loss computation resides in the `MoEGate` class within [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py). This class references `aux_loss_alpha` to scale the load-balancing penalty and checks `seq_aux` to determine whether to average the loss across the sequence or compute it per token. The resulting loss value is stored in the `MOEFeedForward` instance and accessed during the backward pass through `model.layers[i].mlp.aux_loss`.