minimind
🚀🚀 「大模型」2小时完全从0训练26M的小参数GPT!🌏 Train a 26M-parameter GPT from scratch in just 2h!
Learn how MiniMind resumes pre-training across different GPU configurations using automatic step scaling. Discover the lm_checkpoint utility and seamless distributed world size adjustments.
Where Are MiniMind Model Checkpoints Saved? File Paths and Management GuideDiscover where MiniMind model checkpoints are saved. Learn about default file paths and management using the lm_checkpoint utility in this quick guide.
How to Resume Pre-Training from a Checkpoint in MiniMindEffortlessly resume pre-training from a checkpoint in MiniMind using --from_resume 1. Seamlessly continue training with full state restoration.
Where to Find the Recommended Pre-Training Data for MiniMind: Complete Download and Setup GuideFind the recommended pre-training data for MiniMind. Download the pretrain_hq.jsonl file from ModelScope or Hugging Face with our complete guide.
How to Perform Multi-GPU Pre-Training with MiniMind: A Complete DDP GuideLearn multi-GPU pre-training with MiniMind using PyTorch DDP. Effortlessly shard data and synchronize checkpoints across devices with torchrun for efficient model training.
How to Start Pre‑Training a MiniMind Model: Complete GuideLearn how to start pre-training a MiniMind model with our comprehensive guide. This tutorial details running the train_pretrain.py script and configuring hyperparameters for efficient distributed training.
How to Train a Custom Tokenizer for MiniMind: A Complete GuideLearn how to train a custom tokenizer for MiniMind with this complete guide. Follow our step-by-step instructions to generate and load your own Byte-Level BPE tokenizer for enhanced model performance.
What Is the Vocabulary Size of MiniMind’s Custom Tokenizer?Discover MiniMind's custom tokenizer vocabulary size, 6400 tokens. Learn how this compact design ensures a lightweight model with essential language understanding.
How to Configure MiniMind to Use YaRN for Longer Context WindowsConfigure MiniMind to use YaRN for longer contexts by setting inference_rope_scaling to True or passing the --inference_rope_scaling flag. Extend context beyond 32768 tokens.
How MiniMind Achieves Long Context Extrapolation with YaRNDiscover how MiniMind achieves long context extrapolation up to 32,000 tokens using the YaRN scaling algorithm and a learned linear ramp for RoPE frequency components.
RoPE (Rotary Positional Embedding) in MiniMind: Implementation and Usage GuideLearn how MiniMind implements Rotary Positional Embedding RoPE for native sequence order encoding. Understand its usage and optional YaRN scaling for extended context lengths.
MiniMind MoE Architecture Configurable Parameters: Complete Technical GuideExplore MiniMind's MoE architecture and its eight configurable parameters in this technical guide. Learn to control expert routing, auxiliary loss, and computation flow with MiniMindConfig.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →