What is the Transformer Architecture Used in MiniMind?

MiniMind implements a compact decoder-only Transformer with pre-normalization RMSNorm, SwiGLU feed-forward networks, Rotary Positional Embeddings (RoPE), and optional FlashAttention and Mixture-of-Experts (MoE) support.

The jingyaogong/minimind repository provides a minimal yet complete implementation of a modern large language model built entirely in PyTorch. Understanding the Transformer architecture used in MiniMind reveals how this compact decoder-only model achieves competitive performance while remaining small enough to train on consumer hardware.

Decoder-Only Transformer Design

MiniMind follows the architectural patterns established by GPT-3 and LLaMA, implementing a decoder-only autoregressive model. The core structure consists of a stack of MiniMindBlock layers assembled within MiniMindModel, processing input sequences without an encoder branch. According to the source code in model/model_minimind.py (lines 76-84), the model inherits from PreTrainedModel and GenerationMixin, exposing the standard HuggingFace .generate() interface through the MiniMindForCausalLM wrapper class (lines 27-35).

Core Architectural Components

Pre-Normalization with RMSNorm

Instead of standard LayerNorm, MiniMind employs RMSNorm (Root Mean Square Layer Normalization) applied before each sub-layer in a pre-normalization configuration. The RMSNorm class in model/model_minimind.py (lines 96-107) computes normalization without centering, which has been shown to improve training stability in smaller-scale models while reducing computational overhead.

SwiGLU Feed-Forward Networks

The feed-forward network utilizes a SwiGLU architecture implemented through the FeedForward class (lines 16-27 in model/model_minimind.py). This design employs three linear projections—gate_proj, up_proj, and down_proj—combined with SiLU (Sigmoid Linear Unit) activation. The gated mechanism provides superior expressiveness compared to standard ReLU or GELU variants commonly found in earlier Transformer implementations.

Rotary Positional Embeddings with YaRN

MiniMind implements RoPE (Rotary Positional Embeddings) through the precompute_freqs_cis function (lines 9-28 in model/model_minimind.py). This approach encodes relative positional information directly into the attention query and key representations via rotation matrices. The implementation includes optional YaRN (Yet Another RoPE extensioN) scaling parameters, enabling the model to extrapolate to context lengths significantly longer than those encountered during training.

Multi-Head Attention with FlashAttention

The Attention class (lines 50-70 in model/model_minimind.py) implements multi-head scaled dot-product attention with optional FlashAttention support when running on PyTorch 2.0 or newer. The module efficiently handles KV-cache management, storing past key and value tensors to eliminate redundant computation during autoregressive generation. The attention mechanism supports configurable numbers of key-value heads (num_key_value_heads) for grouped-query attention patterns.

Optional Mixture-of-Experts Architecture

When configured with use_moe=True, MiniMind replaces the dense feed-forward network with a Mixture-of-Experts module. The MOEFeedForward class (lines 88-108 in model/model_minimind.py) implements sparse activation where each token is routed to num_experts_per_tok experts selected from a pool of n_routed_experts, alongside n_shared_experts available to all tokens. This architecture includes an auxiliary load-balancing loss term to prevent expert collapse and ensure uniform utilization across the expert pool.

Configuration and Model Classes

All architectural hyperparameters are centralized in MiniMindConfig (lines 8-41 in model/model_minimind.py), including hidden dimensions, layer counts, attention head configurations, MoE settings, RoPE scaling factors, and FlashAttention flags. The MiniMindForCausalLM class wraps the base MiniMindModel to provide the complete causal language modeling interface, compatible with standard HuggingFace training and inference pipelines.

Practical Implementation Examples

The following examples demonstrate how to instantiate and utilize the MiniMind Transformer architecture with different configurations.


# Load a standard dense MiniMind checkpoint

from model.model_minimind import MiniMindForCausalLM, MiniMindConfig
import torch

cfg = MiniMindConfig(
    hidden_size=512,
    num_hidden_layers=8,
    num_attention_heads=8,
    num_key_value_heads=2,
    vocab_size=6400,
    use_moe=False,               # Dense FFN variant

    flash_attn=True,             # Enable FlashAttention if PyTorch >= 2.0

)

model = MiniMindForCausalLM(config=cfg)
model.load_state_dict(torch.load("out/full_sft_512.pth"))
model.eval()

# Generate text using the HuggingFace-compatible API

prompt = "请介绍一下 MiniMind 的模型结构。"

# Note: Use the repository's tokenizer in production (trainer/train_tokenizer.py)

input_ids = torch.zeros((1, 10), dtype=torch.long)  # Placeholder for actual tokenized input

output = model.generate(
    input_ids=input_ids,
    max_new_tokens=64,
    temperature=0.85,
    top_p=0.85,
)

# Instantiate the MoE variant for increased capacity

cfg_moe = MiniMindConfig(
    hidden_size=512,
    num_hidden_layers=8,
    use_moe=True,
    num_experts_per_tok=2,
    n_routed_experts=4,
    n_shared_experts=1,
)

model_moe = MiniMindForCausalLM(config=cfg_moe)

# Loading and generation proceed identically to the dense variant

Summary

  • MiniMind implements a decoder-only Transformer architecture following modern LLM design patterns established by GPT-3 and LLaMA.
  • RMSNorm with pre-normalization provides training stability for compact model sizes.
  • SwiGLU activation functions in the feed-forward networks utilize three linear projections (gate_proj, up_proj, down_proj) with SiLU non-linearity.
  • RoPE with optional YaRN scaling enables flexible context length handling and long-sequence extrapolation.
  • FlashAttention and KV-caching optimize memory usage and computation speed during training and inference.
  • Optional Mixture-of-Experts support allows scaling model capacity without proportionally increasing active parameters.

Frequently Asked Questions

What normalization technique does MiniMind use?

MiniMind employs RMSNorm (Root Mean Square Layer Normalization) rather than standard LayerNorm. The implementation in the RMSNorm class (lines 96-107 of model/model_minimind.py) applies normalization before each sub-layer using only the root-mean-square statistic, omitting the mean-centering step. This pre-normalization configuration enhances training stability specifically for smaller-scale language models.

How does MiniMind handle positional encoding?

MiniMind uses Rotary Positional Embeddings (RoPE) implemented in the precompute_freqs_cis function (lines 9-28 of model/model_minimind.py). This method encodes relative positional information by rotating query and key vectors in the attention mechanism. The implementation supports YaRN scaling parameters to extend the effective context window beyond the training sequence length without additional fine-tuning.

Can MiniMind utilize FlashAttention for accelerated training?

Yes, the Attention class (lines 50-70 of model/model_minimind.py) includes conditional FlashAttention support activated when flash_attn=True is set in the configuration and PyTorch 2.0 or newer is available. The module also implements KV-cache management to store previously computed key and value tensors, significantly reducing computational redundancy during autoregressive text generation.

What is the Mixture-of-Experts configuration in MiniMind?

When enabled via use_moe=True, MiniMind activates the MOEFeedForward class (lines 88-108 of model/model_minimind.py), which replaces the dense feed-forward network with a sparse expert routing mechanism. Each token is processed by num_experts_per_tok experts selected from n_routed_experts total experts, plus n_shared_experts available to all tokens. The implementation includes an auxiliary load-balancing loss to ensure uniform distribution of tokens across the expert pool and prevent routing collapse.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →