MiniMind Training Cost and Time Requirements: Complete Hardware & Budget Guide

Training a MiniMind model from scratch costs between 3–160 Yuan (approximately $0.40–$23 USD) and requires 2–122 hours on a single NVIDIA RTX 3090 GPU, depending on the parameter count and training pipeline selected.

The MiniMind repository by jingyaogong/minimind is an open-source implementation of lightweight large language models designed specifically for resource-constrained environments. According to the project's training documentation, these cost and time requirements for training a MiniMind model make it one of the most affordable options for researchers and developers looking to experiment with full LLM training pipelines without multi-million dollar budgets.

Training Costs by Model Size

MiniMind provides detailed cost breakdowns for two primary model variants, assuming cloud GPU rental rates of ≈ 1.3 ¥/hour (≈ 0.18 USD/hour) for a single RTX 3090.

MiniMind2-Small (26M Parameters)

The smallest variant offers the fastest iteration cycle for experimentation:

  • Pre-training: ≈ 1.1 hours costing ≈ 1.43 ¥
  • Lightweight SFT (sft_mini_512): ≈ 1 hour costing ≈ 1.3 ¥
  • Full SFT (sft_512): ≈ 6 hours costing ≈ 7.8 ¥
  • RLHF/DPO (dpo): ≈ 1 hour costing ≈ 1.3 ¥

MiniMind2 (104M Parameters)

The full-sized variant requires proportionally more compute but remains under $25:

  • Pre-training: ≈ 3.9 hours costing ≈ 5.07 ¥
  • Lightweight SFT: ≈ 3.3 hours costing ≈ 4.29 ¥
  • Full SFT: ≈ 20 hours costing ≈ 26 ¥
  • RLHF/DPO: ≈ 3 hours costing ≈ 3.9 ¥

End-to-End Training Recipes and Total Investment

The repository defines three standard training recipes that combine different data pipelines and model sizes. These represent complete workflows from random initialization to instruction-tuned model.

MiniMind-Zero: The 2-Hour Chatbot (≈ 2.73 ¥)

The fastest path to a functional conversational model uses minimal datasets:

  • Datasets: pretrain_hq.jsonl + sft_mini_512.jsonl
  • Total GPU-hours: ≈ 2.1 hours
  • Total cost: ≈ 2.73 ¥ (≈ $0.38 USD)
  • Output: 26M parameter model capable of basic chat

Full MiniMind-Small Pipeline (≈ 49.6 ¥)

A complete production-quality training run for the 26M model:

  • Datasets: pretrain_hq.jsonl + sft_512.jsonl + sft_2048.jsonl + dpo.jsonl
  • Training regime: 2 epochs across full data pipeline
  • Total GPU-hours: ≈ 38.2 hours
  • Total cost: ≈ 49.6 ¥ (≈ $7 USD)

MiniMind-Full 104M Pipeline (≈ 158.6 ¥)

The comprehensive training run for the full 104M parameter model:

  • Datasets: Same full pipeline as MiniMind-Small
  • Total GPU-hours: ≈ 122 hours
  • Total cost: ≈ 158.6 ¥ (≈ $22.90 USD)
  • Hardware: Single RTX 3090 (or equivalent VRAM)

Hardware Requirements and Scaling Options

While the baseline calculations assume a single NVIDIA RTX 3090, the architecture supports flexible scaling strategies to optimize wall-clock time versus monetary cost.

Multi-GPU configurations can reduce wall-clock time significantly while maintaining similar total costs. For example, an 8 × RTX 4090 setup can compress the MiniMind-Zero training to approximately 10 minutes of real time, as the hourly rental cost per GPU remains roughly equivalent to the single 3090 rate.

All training scripts in trainer/train_pretrain.py, trainer/train_full_sft.py, and trainer/train_dpo.py support distributed training via torchrun with the --nproc_per_node parameter.

Reproducing the Training Runs

The following command-line snippets demonstrate how to execute the three primary cost scenarios documented in the repository. Replace ./dataset with the actual path to your downloaded data files.

Training MiniMind-Zero

This two-stage pipeline produces the cheapest functional model:


# Stage 1: Pre-training on high-quality corpus

torchrun --nproc_per_node 1 trainer/train_pretrain.py \
    --data_dir ./dataset \
    --max_seq_len 512 \
    --epochs 1 \
    --learning_rate 3e-4 \
    --save_dir ./out

# Stage 2: Lightweight SFT on mini dataset

torchrun --nproc_per_node 1 trainer/train_full_sft.py \
    --data_dir ./dataset \
    --max_seq_len 512 \
    --epochs 1 \
    --learning_rate 2e-4 \
    --save_dir ./out

Full Parameter Fine-Tuning with DPO

For the complete 38-hour pipeline including RLHF:


# Pre-training (2 epochs recommended for quality)

torchrun --nproc_per_node 1 trainer/train_pretrain.py \
    --data_dir ./dataset \
    --max_seq_len 512 \
    --epochs 2 \
    --learning_rate 3e-4

# Full SFT on mixed-length sequences

torchrun --nproc_per_node 1 trainer/train_full_sft.py \
    --data_dir ./dataset \
    --max_seq_len 1024 \
    --epochs 2 \
    --learning_rate 2e-4

# Direct Preference Optimization

torchrun --nproc_per_node 1 trainer/train_dpo.py \
    --data_dir ./dataset \
    --max_seq_len 1024 \
    --epochs 2 \
    --learning_rate 1e-5

Switching to the 104M Configuration

To train the larger MiniMind2 variant (104M), modify the model selection via environment variable before running the same training commands:

export MINI_MODEL=MiniMind2   # Selects 104M config in model/model_minimind.py

torchrun --nproc_per_node 1 trainer/train_pretrain.py \
    --data_dir ./dataset \
    --epochs 2 \
    --learning_rate 3e-4

Checkpointing note: All scripts automatically save checkpoints every 100 steps and support resumption via --from_resume 1, ensuring that interrupted training runs do not waste GPU hours.

Key Implementation Files

The cost and time estimates derive from the following source files in jingyaogong/minimind:

Summary

  • MiniMind-Zero delivers a chat-capable 26M model for ≈ 3 ¥ in 2 hours on a single GPU
  • Full 26M pipeline costs under 50 ¥ (≈ $7 USD) and completes in 38 hours
  • 104M full training requires ≈ 160 ¥ (≈ $23 USD) and 122 hours wall time
  • All calculations assume RTX 3090 at 1.3 ¥/hour; multi-GPU setups reduce time proportionally while maintaining similar total costs
  • Training scripts support resumption via checkpointing to optimize GPU-hour utilization

Frequently Asked Questions

How much does it cost to train a MiniMind model from scratch?

Training costs range from 2.73 ¥ (≈ $0.40 USD) for the minimal MiniMind-Zero recipe to 158.6 ¥ (≈ $23 USD) for the complete 104M parameter model with full SFT and DPO training. These figures assume single RTX 3090 GPU rental at approximately 1.3 ¥ per hour according to the repository's cost analysis.

What GPU do I need to train MiniMind?

The reference hardware is a single NVIDIA RTX 3090 (24GB VRAM). The training scripts in trainer/train_pretrain.py and related files use torchrun with --nproc_per_node 1 by default, though the codebase supports multi-GPU configurations to accelerate training without increasing total cost significantly.

How long does MiniMind pre-training take?

Pre-training takes ≈ 1.1 hours for the 26M MiniMind2-Small model and ≈ 3.9 hours for the 104M MiniMind2 model when training on the pretrain_hq.jsonl dataset for one epoch. Running 2 epochs for better quality doubles these times to approximately 2.2 and 7.8 hours respectively.

Can I reduce training time with multiple GPUs?

Yes. While the baseline assumes single-GPU training, multi-GPU setups such as 8 × RTX 4090 can reduce wall-clock time to approximately 10 minutes for the MiniMind-Zero recipe. The total GPU-hour cost remains roughly equivalent because the hourly rental rate per GPU is similar to the single 3090 rate, making this a time-efficient rather than cost-saving optimization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →