minimind

🚀🚀 「大模型」2小时完全从0训练26M的小参数GPT!🌏 Train a 26M-parameter GPT from scratch in just 2h!

22 articles 42.9k View on GitHub ↗
22 articles
MiniMind Pre-Training Resume Across Different GPU Configurations: Automatic Step Scaling Explained

Learn how MiniMind resumes pre-training across different GPU configurations using automatic step scaling. Discover the lm_checkpoint utility and seamless distributed world size adjustments.

deep-dive
Mar 24, 2026
Where Are MiniMind Model Checkpoints Saved? File Paths and Management Guide

Discover where MiniMind model checkpoints are saved. Learn about default file paths and management using the lm_checkpoint utility in this quick guide.

how-to-guide
Mar 24, 2026
How to Resume Pre-Training from a Checkpoint in MiniMind

Effortlessly resume pre-training from a checkpoint in MiniMind using --from_resume 1. Seamlessly continue training with full state restoration.

how-to-guide
Mar 24, 2026
Where to Find the Recommended Pre-Training Data for MiniMind: Complete Download and Setup Guide

Find the recommended pre-training data for MiniMind. Download the pretrain_hq.jsonl file from ModelScope or Hugging Face with our complete guide.

how-to-guide
Mar 24, 2026
How to Perform Multi-GPU Pre-Training with MiniMind: A Complete DDP Guide

Learn multi-GPU pre-training with MiniMind using PyTorch DDP. Effortlessly shard data and synchronize checkpoints across devices with torchrun for efficient model training.

how-to-guide
Mar 24, 2026
How to Start Pre‑Training a MiniMind Model: Complete Guide

Learn how to start pre-training a MiniMind model with our comprehensive guide. This tutorial details running the train_pretrain.py script and configuring hyperparameters for efficient distributed training.

how-to-guide
Mar 24, 2026
How to Train a Custom Tokenizer for MiniMind: A Complete Guide

Learn how to train a custom tokenizer for MiniMind with this complete guide. Follow our step-by-step instructions to generate and load your own Byte-Level BPE tokenizer for enhanced model performance.

how-to-guide
Mar 24, 2026
What Is the Vocabulary Size of MiniMind’s Custom Tokenizer?

Discover MiniMind's custom tokenizer vocabulary size, 6400 tokens. Learn how this compact design ensures a lightweight model with essential language understanding.

deep-dive
Mar 24, 2026
How to Configure MiniMind to Use YaRN for Longer Context Windows

Configure MiniMind to use YaRN for longer contexts by setting inference_rope_scaling to True or passing the --inference_rope_scaling flag. Extend context beyond 32768 tokens.

how-to-guide
Mar 24, 2026
How MiniMind Achieves Long Context Extrapolation with YaRN

Discover how MiniMind achieves long context extrapolation up to 32,000 tokens using the YaRN scaling algorithm and a learned linear ramp for RoPE frequency components.

deep-dive
Mar 24, 2026
RoPE (Rotary Positional Embedding) in MiniMind: Implementation and Usage Guide

Learn how MiniMind implements Rotary Positional Embedding RoPE for native sequence order encoding. Understand its usage and optional YaRN scaling for extended context lengths.

implementation-and-usage-guide
Mar 24, 2026
MiniMind MoE Architecture Configurable Parameters: Complete Technical Guide

Explore MiniMind's MoE architecture and its eight configurable parameters in this technical guide. Learn to control expert routing, auxiliary loss, and computation flow with MiniMindConfig.

technical-guide
Mar 24, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →