# What is the Transformer Architecture Used in MiniMind?

> Discover the Transformer architecture in MiniMind. Explore its compact decoder-only design featuring RMSNorm, SwiGLU, RoPE, and optional FlashAttention/MoE.

- Repository: [jingyaogong/minimind](https://github.com/jingyaogong/minimind)
- Tags: internals
- Published: 2026-03-24

---

**MiniMind implements a compact decoder-only Transformer with pre-normalization RMSNorm, SwiGLU feed-forward networks, Rotary Positional Embeddings (RoPE), and optional FlashAttention and Mixture-of-Experts (MoE) support.**

The jingyaogong/minimind repository provides a minimal yet complete implementation of a modern large language model built entirely in PyTorch. Understanding the Transformer architecture used in MiniMind reveals how this compact decoder-only model achieves competitive performance while remaining small enough to train on consumer hardware.

## Decoder-Only Transformer Design

MiniMind follows the architectural patterns established by GPT-3 and LLaMA, implementing a decoder-only autoregressive model. The core structure consists of a stack of `MiniMindBlock` layers assembled within `MiniMindModel`, processing input sequences without an encoder branch. According to the source code in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) (lines 76-84), the model inherits from `PreTrainedModel` and `GenerationMixin`, exposing the standard HuggingFace `.generate()` interface through the `MiniMindForCausalLM` wrapper class (lines 27-35).

## Core Architectural Components

### Pre-Normalization with RMSNorm

Instead of standard LayerNorm, MiniMind employs **RMSNorm** (Root Mean Square Layer Normalization) applied before each sub-layer in a pre-normalization configuration. The `RMSNorm` class in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) (lines 96-107) computes normalization without centering, which has been shown to improve training stability in smaller-scale models while reducing computational overhead.

### SwiGLU Feed-Forward Networks

The feed-forward network utilizes a **SwiGLU** architecture implemented through the `FeedForward` class (lines 16-27 in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)). This design employs three linear projections—`gate_proj`, `up_proj`, and `down_proj`—combined with **SiLU** (Sigmoid Linear Unit) activation. The gated mechanism provides superior expressiveness compared to standard ReLU or GELU variants commonly found in earlier Transformer implementations.

### Rotary Positional Embeddings with YaRN

MiniMind implements **RoPE** (Rotary Positional Embeddings) through the `precompute_freqs_cis` function (lines 9-28 in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)). This approach encodes relative positional information directly into the attention query and key representations via rotation matrices. The implementation includes optional **YaRN** (Yet Another RoPE extensioN) scaling parameters, enabling the model to extrapolate to context lengths significantly longer than those encountered during training.

### Multi-Head Attention with FlashAttention

The `Attention` class (lines 50-70 in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)) implements multi-head scaled dot-product attention with optional **FlashAttention** support when running on PyTorch 2.0 or newer. The module efficiently handles **KV-cache** management, storing past key and value tensors to eliminate redundant computation during autoregressive generation. The attention mechanism supports configurable numbers of key-value heads (`num_key_value_heads`) for grouped-query attention patterns.

## Optional Mixture-of-Experts Architecture

When configured with `use_moe=True`, MiniMind replaces the dense feed-forward network with a **Mixture-of-Experts** module. The `MOEFeedForward` class (lines 88-108 in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)) implements sparse activation where each token is routed to `num_experts_per_tok` experts selected from a pool of `n_routed_experts`, alongside `n_shared_experts` available to all tokens. This architecture includes an auxiliary load-balancing loss term to prevent expert collapse and ensure uniform utilization across the expert pool.

## Configuration and Model Classes

All architectural hyperparameters are centralized in `MiniMindConfig` (lines 8-41 in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)), including hidden dimensions, layer counts, attention head configurations, MoE settings, RoPE scaling factors, and FlashAttention flags. The `MiniMindForCausalLM` class wraps the base `MiniMindModel` to provide the complete causal language modeling interface, compatible with standard HuggingFace training and inference pipelines.

## Practical Implementation Examples

The following examples demonstrate how to instantiate and utilize the MiniMind Transformer architecture with different configurations.

```python

# Load a standard dense MiniMind checkpoint

from model.model_minimind import MiniMindForCausalLM, MiniMindConfig
import torch

cfg = MiniMindConfig(
    hidden_size=512,
    num_hidden_layers=8,
    num_attention_heads=8,
    num_key_value_heads=2,
    vocab_size=6400,
    use_moe=False,               # Dense FFN variant

    flash_attn=True,             # Enable FlashAttention if PyTorch >= 2.0

)

model = MiniMindForCausalLM(config=cfg)
model.load_state_dict(torch.load("out/full_sft_512.pth"))
model.eval()

# Generate text using the HuggingFace-compatible API

prompt = "请介绍一下 MiniMind 的模型结构。"

# Note: Use the repository's tokenizer in production (trainer/train_tokenizer.py)

input_ids = torch.zeros((1, 10), dtype=torch.long)  # Placeholder for actual tokenized input

output = model.generate(
    input_ids=input_ids,
    max_new_tokens=64,
    temperature=0.85,
    top_p=0.85,
)

```

```python

# Instantiate the MoE variant for increased capacity

cfg_moe = MiniMindConfig(
    hidden_size=512,
    num_hidden_layers=8,
    use_moe=True,
    num_experts_per_tok=2,
    n_routed_experts=4,
    n_shared_experts=1,
)

model_moe = MiniMindForCausalLM(config=cfg_moe)

# Loading and generation proceed identically to the dense variant

```

## Summary

- MiniMind implements a **decoder-only Transformer** architecture following modern LLM design patterns established by GPT-3 and LLaMA.
- **RMSNorm** with pre-normalization provides training stability for compact model sizes.
- **SwiGLU** activation functions in the feed-forward networks utilize three linear projections (`gate_proj`, `up_proj`, `down_proj`) with SiLU non-linearity.
- **RoPE** with optional **YaRN** scaling enables flexible context length handling and long-sequence extrapolation.
- **FlashAttention** and **KV-caching** optimize memory usage and computation speed during training and inference.
- Optional **Mixture-of-Experts** support allows scaling model capacity without proportionally increasing active parameters.

## Frequently Asked Questions

### What normalization technique does MiniMind use?

MiniMind employs **RMSNorm** (Root Mean Square Layer Normalization) rather than standard LayerNorm. The implementation in the `RMSNorm` class (lines 96-107 of [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)) applies normalization before each sub-layer using only the root-mean-square statistic, omitting the mean-centering step. This pre-normalization configuration enhances training stability specifically for smaller-scale language models.

### How does MiniMind handle positional encoding?

MiniMind uses **Rotary Positional Embeddings (RoPE)** implemented in the `precompute_freqs_cis` function (lines 9-28 of [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)). This method encodes relative positional information by rotating query and key vectors in the attention mechanism. The implementation supports **YaRN** scaling parameters to extend the effective context window beyond the training sequence length without additional fine-tuning.

### Can MiniMind utilize FlashAttention for accelerated training?

Yes, the `Attention` class (lines 50-70 of [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)) includes conditional **FlashAttention** support activated when `flash_attn=True` is set in the configuration and PyTorch 2.0 or newer is available. The module also implements **KV-cache** management to store previously computed key and value tensors, significantly reducing computational redundancy during autoregressive text generation.

### What is the Mixture-of-Experts configuration in MiniMind?

When enabled via `use_moe=True`, MiniMind activates the `MOEFeedForward` class (lines 88-108 of [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py)), which replaces the dense feed-forward network with a sparse expert routing mechanism. Each token is processed by `num_experts_per_tok` experts selected from `n_routed_experts` total experts, plus `n_shared_experts` available to all tokens. The implementation includes an auxiliary load-balancing loss to ensure uniform distribution of tokens across the expert pool and prevent routing collapse.