# MiniMind Training Cost and Time Requirements: Complete Hardware & Budget Guide

> Discover MiniMind training cost and time. Learn about hardware needs and budget considerations for training a MiniMind model, from 3 Yuan and 2 hours to 160 Yuan and 122 hours.

- Repository: [jingyaogong/minimind](https://github.com/jingyaogong/minimind)
- Tags: guide
- Published: 2026-03-24

---

**Training a MiniMind model from scratch costs between 3–160 Yuan (approximately $0.40–$23 USD) and requires 2–122 hours on a single NVIDIA RTX 3090 GPU, depending on the parameter count and training pipeline selected.**

The MiniMind repository by `jingyaogong/minimind` is an open-source implementation of lightweight large language models designed specifically for resource-constrained environments. According to the project's training documentation, these cost and time requirements for training a MiniMind model make it one of the most affordable options for researchers and developers looking to experiment with full LLM training pipelines without multi-million dollar budgets.

## Training Costs by Model Size

MiniMind provides detailed cost breakdowns for two primary model variants, assuming cloud GPU rental rates of **≈ 1.3 ¥/hour (≈ 0.18 USD/hour)** for a single RTX 3090.

### MiniMind2-Small (26M Parameters)

The smallest variant offers the fastest iteration cycle for experimentation:

- **Pre-training**: ≈ 1.1 hours costing ≈ 1.43 ¥
- **Lightweight SFT** (`sft_mini_512`): ≈ 1 hour costing ≈ 1.3 ¥
- **Full SFT** (`sft_512`): ≈ 6 hours costing ≈ 7.8 ¥
- **RLHF/DPO** (`dpo`): ≈ 1 hour costing ≈ 1.3 ¥

### MiniMind2 (104M Parameters)

The full-sized variant requires proportionally more compute but remains under $25:

- **Pre-training**: ≈ 3.9 hours costing ≈ 5.07 ¥
- **Lightweight SFT**: ≈ 3.3 hours costing ≈ 4.29 ¥
- **Full SFT**: ≈ 20 hours costing ≈ 26 ¥
- **RLHF/DPO**: ≈ 3 hours costing ≈ 3.9 ¥

## End-to-End Training Recipes and Total Investment

The repository defines three standard training recipes that combine different data pipelines and model sizes. These represent complete workflows from random initialization to instruction-tuned model.

### MiniMind-Zero: The 2-Hour Chatbot (≈ 2.73 ¥)

The fastest path to a functional conversational model uses minimal datasets:

- **Datasets**: `pretrain_hq.jsonl` + `sft_mini_512.jsonl`
- **Total GPU-hours**: ≈ 2.1 hours
- **Total cost**: ≈ 2.73 ¥ (≈ $0.38 USD)
- **Output**: 26M parameter model capable of basic chat

### Full MiniMind-Small Pipeline (≈ 49.6 ¥)

A complete production-quality training run for the 26M model:

- **Datasets**: `pretrain_hq.jsonl` + `sft_512.jsonl` + `sft_2048.jsonl` + `dpo.jsonl`
- **Training regime**: 2 epochs across full data pipeline
- **Total GPU-hours**: ≈ 38.2 hours
- **Total cost**: ≈ 49.6 ¥ (≈ $7 USD)

### MiniMind-Full 104M Pipeline (≈ 158.6 ¥)

The comprehensive training run for the full 104M parameter model:

- **Datasets**: Same full pipeline as MiniMind-Small
- **Total GPU-hours**: ≈ 122 hours
- **Total cost**: ≈ 158.6 ¥ (≈ $22.90 USD)
- **Hardware**: Single RTX 3090 (or equivalent VRAM)

## Hardware Requirements and Scaling Options

While the baseline calculations assume a **single NVIDIA RTX 3090**, the architecture supports flexible scaling strategies to optimize wall-clock time versus monetary cost.

Multi-GPU configurations can reduce wall-clock time significantly while maintaining similar total costs. For example, an **8 × RTX 4090** setup can compress the MiniMind-Zero training to approximately **10 minutes** of real time, as the hourly rental cost per GPU remains roughly equivalent to the single 3090 rate.

All training scripts in [`trainer/train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_pretrain.py), [`trainer/train_full_sft.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_full_sft.py), and [`trainer/train_dpo.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_dpo.py) support distributed training via `torchrun` with the `--nproc_per_node` parameter.

## Reproducing the Training Runs

The following command-line snippets demonstrate how to execute the three primary cost scenarios documented in the repository. Replace `./dataset` with the actual path to your downloaded data files.

### Training MiniMind-Zero

This two-stage pipeline produces the cheapest functional model:

```bash

# Stage 1: Pre-training on high-quality corpus

torchrun --nproc_per_node 1 trainer/train_pretrain.py \
    --data_dir ./dataset \
    --max_seq_len 512 \
    --epochs 1 \
    --learning_rate 3e-4 \
    --save_dir ./out

# Stage 2: Lightweight SFT on mini dataset

torchrun --nproc_per_node 1 trainer/train_full_sft.py \
    --data_dir ./dataset \
    --max_seq_len 512 \
    --epochs 1 \
    --learning_rate 2e-4 \
    --save_dir ./out

```

### Full Parameter Fine-Tuning with DPO

For the complete 38-hour pipeline including RLHF:

```bash

# Pre-training (2 epochs recommended for quality)

torchrun --nproc_per_node 1 trainer/train_pretrain.py \
    --data_dir ./dataset \
    --max_seq_len 512 \
    --epochs 2 \
    --learning_rate 3e-4

# Full SFT on mixed-length sequences

torchrun --nproc_per_node 1 trainer/train_full_sft.py \
    --data_dir ./dataset \
    --max_seq_len 1024 \
    --epochs 2 \
    --learning_rate 2e-4

# Direct Preference Optimization

torchrun --nproc_per_node 1 trainer/train_dpo.py \
    --data_dir ./dataset \
    --max_seq_len 1024 \
    --epochs 2 \
    --learning_rate 1e-5

```

### Switching to the 104M Configuration

To train the larger MiniMind2 variant (104M), modify the model selection via environment variable before running the same training commands:

```bash
export MINI_MODEL=MiniMind2   # Selects 104M config in model/model_minimind.py

torchrun --nproc_per_node 1 trainer/train_pretrain.py \
    --data_dir ./dataset \
    --epochs 2 \
    --learning_rate 3e-4

```

**Checkpointing note**: All scripts automatically save checkpoints every 100 steps and support resumption via `--from_resume 1`, ensuring that interrupted training runs do not waste GPU hours.

## Key Implementation Files

The cost and time estimates derive from the following source files in `jingyaogong/minimind`:

- [`README.md`](https://github.com/jingyaogong/minimind/blob/main/README.md) – Contains the definitive training cost tables and hardware assumptions
- [`trainer/train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_pretrain.py) – Implements the causal language modeling pre-training loop
- [`trainer/train_full_sft.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_full_sft.py) – Handles supervised fine-tuning on instruction datasets
- [`trainer/train_dpo.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_dpo.py) – Implements Direct Preference Optimization for alignment
- [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) – Defines the transformer architecture and parameter counts
- [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py) – Provides checkpoint management and data loading utilities
- [`dataset/dataset.md`](https://github.com/jingyaogong/minimind/blob/main/dataset/dataset.md) – Documents dataset sizes and recommended combinations for cost calculation

## Summary

- **MiniMind-Zero** delivers a chat-capable 26M model for **≈ 3 ¥** in **2 hours** on a single GPU
- **Full 26M pipeline** costs under **50 ¥** (≈ $7 USD) and completes in **38 hours**
- **104M full training** requires **≈ 160 ¥** (≈ $23 USD) and **122 hours** wall time
- All calculations assume **RTX 3090** at **1.3 ¥/hour**; multi-GPU setups reduce time proportionally while maintaining similar total costs
- Training scripts support resumption via checkpointing to optimize GPU-hour utilization

## Frequently Asked Questions

### How much does it cost to train a MiniMind model from scratch?

Training costs range from **2.73 ¥ (≈ $0.40 USD)** for the minimal MiniMind-Zero recipe to **158.6 ¥ (≈ $23 USD)** for the complete 104M parameter model with full SFT and DPO training. These figures assume single RTX 3090 GPU rental at approximately 1.3 ¥ per hour according to the repository's cost analysis.

### What GPU do I need to train MiniMind?

The reference hardware is a **single NVIDIA RTX 3090** (24GB VRAM). The training scripts in [`trainer/train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_pretrain.py) and related files use `torchrun` with `--nproc_per_node 1` by default, though the codebase supports multi-GPU configurations to accelerate training without increasing total cost significantly.

### How long does MiniMind pre-training take?

Pre-training takes **≈ 1.1 hours** for the 26M MiniMind2-Small model and **≈ 3.9 hours** for the 104M MiniMind2 model when training on the `pretrain_hq.jsonl` dataset for one epoch. Running 2 epochs for better quality doubles these times to approximately 2.2 and 7.8 hours respectively.

### Can I reduce training time with multiple GPUs?

Yes. While the baseline assumes single-GPU training, multi-GPU setups such as **8 × RTX 4090** can reduce wall-clock time to approximately **10 minutes** for the MiniMind-Zero recipe. The total GPU-hour cost remains roughly equivalent because the hourly rental rate per GPU is similar to the single 3090 rate, making this a time-efficient rather than cost-saving optimization.