# minimind | jingyaogong | Knowledge Base | Instagit

🚀🚀 「大模型」2小时完全从0训练26M的小参数GPT！🌏 Train a 26M-parameter GPT from scratch in just 2h!

GitHub Stars: 42.9k

Repository: https://github.com/jingyaogong/minimind

---

## Articles

### [MiniMind Pre-Training Resume Across Different GPU Configurations: Automatic Step Scaling Explained](/jingyaogong/minimind/minimind-pretraining-resume-gpu-change)

Learn how MiniMind resumes pre-training across different GPU configurations using automatic step scaling. Discover the lm_checkpoint utility and seamless distributed world size adjustments.

- Tags: deep-dive
- Published: 2026-03-24

### [Where Are MiniMind Model Checkpoints Saved? File Paths and Management Guide](/jingyaogong/minimind/minimind-model-checkpoint-saving-location)

Discover where MiniMind model checkpoints are saved. Learn about default file paths and management using the lm_checkpoint utility in this quick guide.

- Tags: how-to-guide
- Published: 2026-03-24

### [How to Resume Pre-Training from a Checkpoint in MiniMind](/jingyaogong/minimind/resume-pretraining-minimind-checkpoint)

Effortlessly resume pre-training from a checkpoint in MiniMind using --from_resume 1. Seamlessly continue training with full state restoration.

- Tags: how-to-guide
- Published: 2026-03-24

### [Where to Find the Recommended Pre-Training Data for MiniMind: Complete Download and Setup Guide](/jingyaogong/minimind/minimind-pretraining-data-location)

Find the recommended pre-training data for MiniMind. Download the pretrain_hq.jsonl file from ModelScope or Hugging Face with our complete guide.

- Tags: how-to-guide
- Published: 2026-03-24

### [How to Perform Multi-GPU Pre-Training with MiniMind: A Complete DDP Guide](/jingyaogong/minimind/multi-gpu-pretraining-minimind)

Learn multi-GPU pre-training with MiniMind using PyTorch DDP. Effortlessly shard data and synchronize checkpoints across devices with torchrun for efficient model training.

- Tags: how-to-guide
- Published: 2026-03-24

### [How to Start Pre‑Training a MiniMind Model: Complete Guide](/jingyaogong/minimind/start-pretraining-minimind-model)

Learn how to start pre-training a MiniMind model with our comprehensive guide. This tutorial details running the train_pretrain.py script and configuring hyperparameters for efficient distributed training.

- Tags: how-to-guide
- Published: 2026-03-24

### [How to Train a Custom Tokenizer for MiniMind: A Complete Guide](/jingyaogong/minimind/train-custom-tokenizer-minimind)

Learn how to train a custom tokenizer for MiniMind with this complete guide. Follow our step-by-step instructions to generate and load your own Byte-Level BPE tokenizer for enhanced model performance.

- Tags: how-to-guide
- Published: 2026-03-24

### [What Is the Vocabulary Size of MiniMind’s Custom Tokenizer?](/jingyaogong/minimind/minimind-tokenizer-vocabulary-size)

Discover MiniMind's custom tokenizer vocabulary size, 6400 tokens. Learn how this compact design ensures a lightweight model with essential language understanding.

- Tags: deep-dive
- Published: 2026-03-24

### [How to Configure MiniMind to Use YaRN for Longer Context Windows](/jingyaogong/minimind/configure-minimind-yarn-long-context)

Configure MiniMind to use YaRN for longer contexts by setting inference_rope_scaling to True or passing the --inference_rope_scaling flag. Extend context beyond 32768 tokens.

- Tags: how-to-guide
- Published: 2026-03-24

### [How MiniMind Achieves Long Context Extrapolation with YaRN](/jingyaogong/minimind/minimind-yarn-long-context-extrapolation)

Discover how MiniMind achieves long context extrapolation up to 32,000 tokens using the YaRN scaling algorithm and a learned linear ramp for RoPE frequency components.

- Tags: deep-dive
- Published: 2026-03-24

### [RoPE (Rotary Positional Embedding) in MiniMind: Implementation and Usage Guide](/jingyaogong/minimind/minimind-rope-positional-embedding)

Learn how MiniMind implements Rotary Positional Embedding RoPE for native sequence order encoding. Understand its usage and optional YaRN scaling for extended context lengths.

- Tags: implementation-and-usage-guide
- Published: 2026-03-24

### [MiniMind MoE Architecture Configurable Parameters: Complete Technical Guide](/jingyaogong/minimind/minimind-moe-configuration-parameters)

Explore MiniMind's MoE architecture and its eight configurable parameters in this technical guide. Learn to control expert routing, auxiliary loss, and computation flow with MiniMindConfig.

- Tags: technical-guide
- Published: 2026-03-24

### [How MiniMind's MoE Routing Mechanism Works: From Gate Scores to Expert Dispatch](/jingyaogong/minimind/minimind-moe-routing-mechanism)

Explore MiniMind's MoE routing mechanism. Learn how gate scores select top-k experts per token and dispatch representations with auxiliary loss for uniform utilization.

- Tags: deep-dive
- Published: 2026-03-24

### [How MiniMind Implements Grouped-Query Attention (GQA): A Deep Dive into the Source Code](/jingyaogong/minimind/minimind-gqa-implementation)

Explore how MiniMind implements Grouped-Query Attention GQA by adjusting key-value heads and replicating them. Dive into the source code for a technical deep dive.

- Tags: deep-dive
- Published: 2026-03-24

### [RMSNorm Normalization in MiniMind: Implementation and Architecture Guide](/jingyaogong/minimind/minimind-rmsnorm-normalization)

Discover RMSNorm in MiniMind. This guide explains how RMSNorm replaces LayerNorm, stabilizing training and cutting computation by removing mean-subtraction.

- Tags: architecture
- Published: 2026-03-24

### [What is the Transformer Architecture Used in MiniMind?](/jingyaogong/minimind/minimind-transformer-architecture)

Discover the Transformer architecture in MiniMind. Explore its compact decoder-only design featuring RMSNorm, SwiGLU, RoPE, and optional FlashAttention/MoE.

- Tags: internals
- Published: 2026-03-24

### [MiniMind Compatible Frameworks: Training, Fine-Tuning, and Inference Integration](/jingyaogong/minimind/minimind-framework-compatibility)

Discover MiniMind compatible frameworks for training fine-tuning and inference including HuggingFace DeepSpeed llama.cpp vllm and ollama Integrate seamlessly with your ML workflow.

- Tags: api-reference
- Published: 2026-03-24

### [How MiniMind Implements Mixture of Experts (MoE): Architecture and Code Deep Dive](/jingyaogong/minimind/minimind-moe-implementation-details)

Discover how MiniMind implements Mixture of Experts MoE. Explore its architecture, code, gating network, expert routing, and load balancing loss for efficient AI models.

- Tags: deep-dive
- Published: 2026-03-24

### [MiniMind Full LLM Training Pipeline: SFT, LoRA, DPO, and RLHF Explained](/jingyaogong/minimind/minimind-llm-training-pipeline-features)

Explore the MiniMind LLM training pipeline covering SFT, LoRA, DPO, and RLHF. Train powerful language models efficiently with this comprehensive toolkit.

- Tags: deep-dive
- Published: 2026-03-24

### [MiniMind Parameter Size Compared to GPT-3: From 26M to 145M Parameters Explained](/jingyaogong/minimind/minimind-parameter-size-vs-gpt3)

Explore how MiniMind models with 26M to 145M parameters compare to GPT-3's 175B parameters. Discover efficient chat capabilities in a significantly smaller model.

- Tags: comparison
- Published: 2026-03-24

### [MiniMind Training Cost and Time Requirements: Complete Hardware & Budget Guide](/jingyaogong/minimind/minimind-model-training-cost-time)

Discover MiniMind training cost and time. Learn about hardware needs and budget considerations for training a MiniMind model, from 3 Yuan and 2 hours to 160 Yuan and 122 hours.

- Tags: guide
- Published: 2026-03-24

### [How to Train Your Own Small Language Model from Scratch with MiniMind](/jingyaogong/minimind/how-to-train-small-language-model-minimind)

Train your own small language model from scratch with MiniMind. Follow sequential steps from tokenizer prep to RL alignment using the jingyaogong/minimind PyTorch components.

- Tags: how-to-guide
- Published: 2026-03-24

