Where to Find the Recommended Pre-Training Data for MiniMind: Complete Download and Setup Guide

The recommended pre-training data for MiniMind is contained in the pretrain_hq.jsonl file, a high-quality 1.6 GB subset of the JiangShu large-model dataset available on both ModelScope and Hugging Face.

MiniMind is an open-source, lightweight large language model designed for rapid reproduction and educational fine-tuning. To replicate the official "Zero" chat model, you must use the recommended pre-training data for MiniMind, specifically the curated pretrain_hq.jsonl file that the authors extracted for optimal training efficiency. This guide provides the exact download locations, setup procedures, and integration steps required to begin training immediately.

According to the README_en.md at line 301, the primary pre-training corpus is pretrain_hq.jsonl. This file represents a filtered, high-quality subset of the JiangShu large-model dataset, weighing approximately 1.6 GB.

The dataset contains exclusively Chinese text samples with lengths under 512 characters, making it ideal for the MiniMind architecture's sequence length constraints. For the fastest reproduction of the "Zero" chat model, the authors recommend combining pretrain_hq.jsonl with the supervised fine-tuning file sft_mini_512.jsonl.

Official Download Sources

The MiniMind project hosts the pre-training data on two major model hub platforms. Both mirrors contain identical files, so choose based on your regional connectivity.

ModelScope Mirror

Access the dataset at the ModelScope repository:

https://www.modelscope.cn/datasets/gongjy/minimind_dataset/files

Look for the entry named pretrain_hq.jsonl (marked with a ✨ icon) in the file list. Download this file along with sft_mini_512.jsonl if you intend to complete the full training pipeline.

Hugging Face Hub

Alternatively, fetch the data from the Hugging Face dataset hub:

https://huggingface.co/datasets/jingyaogong/minimind_dataset

The repository structure mirrors the ModelScope layout, containing the same pretrain_hq.jsonl file required for the initial pre-training phase.

Step-by-Step Setup Guide

After downloading the files, you must place them in the correct directory structure for the training scripts to locate them automatically.

1. Create the Dataset Directory

From the repository root, execute:

mkdir -p ./dataset

This directory path is hardcoded in the training scripts' default configurations.

2. Place the Downloaded Files

Move pretrain_hq.jsonl (and optionally sft_mini_512.jsonl) into the ./dataset folder:

mv ~/Downloads/pretrain_hq.jsonl ./dataset/
mv ~/Downloads/sft_mini_512.jsonl ./dataset/  # Optional, for SFT stage

3. Verify the Source and Format

As documented in README_en.md at line 460, verify that your pretrain_hq.jsonl contains Chinese text samples with fewer than 512 characters per entry. Each line must be a valid JSON object with a "text" field:

{"text": "这是一个示例文本..."}

Loading and Training with the Dataset

Once the files are in place, you can immediately begin the training process using the provided trainer scripts or load the data programmatically.

Running Pre-Training

Execute the pre-training script with the explicit dataset path:

python trainer/train_pretrain.py \
  --model_name minimind \
  --dataset_path ./dataset/pretrain_hq.jsonl \
  --max_seq_len 320

The train_pretrain.py script automatically parses the JSONL format and handles the JiangShu dataset's text-only structure.

Loading with the Hugging Face Datasets Library

For custom preprocessing or inspection, load the data using the 🤗 Datasets library:

from datasets import load_dataset

# Load the pre-training corpus

pretrain = load_dataset(
    "json", 
    data_files="./dataset/pretrain_hq.jsonl", 
    split="train"
)

# Verify the contents

print(f"Number of examples: {len(pretrain)}")
print(pretrain[0]["text"][:200])

This approach validates that your pretrain_hq.jsonl follows the expected schema before launching expensive training runs.

Complete Zero Model Reproduction Pipeline

To reproduce the official "Zero" chat model exactly as the authors intended, execute the two-stage training sequence:


# Stage 1: Pre-training on high-quality corpus

python trainer/train_pretrain.py \
  --dataset_path ./dataset/pretrain_hq.jsonl \
  --max_seq_len 320

# Stage 2: Supervised fine-tuning for chat capabilities

python trainer/train_full_sft.py \
  --dataset_path ./dataset/sft_mini_512.jsonl \
  --max_seq_len 512

This combination of pretrain_hq.jsonl followed by sft_mini_512.jsonl yields the fastest path to the functional chat model according to the repository documentation.

Dataset Format and Technical Details

The pretrain_hq.jsonl file adheres to strict formatting requirements enforced by the training scripts:

  • Format: JSON Lines (JSONL), with one JSON object per line
  • Schema: Each object must contain a "text" string field containing the raw training text
  • Language: Chinese text only (sourced from the JiangShu dataset)
  • Length constraint: Text length must not exceed 512 characters to align with the model's context window
  • File size: Approximately 1.6 GB uncompressed

These constraints ensure compatibility with the data collation logic inside trainer/train_pretrain.py and prevent tokenization errors during the training loop.

Summary

  • pretrain_hq.jsonl is the single recommended pre-training file for MiniMind, extracted from the JiangShu large-model dataset.
  • Download the 1.6 GB file from ModelScope or Hugging Face mirrors listed in the official repository.
  • Place the file in the ./dataset directory at the repository root for automatic detection by training scripts.
  • Use trainer/train_pretrain.py with --dataset_path ./dataset/pretrain_hq.jsonl to initiate training.
  • Combine with sft_mini_512.jsonl using trainer/train_full_sft.py to reproduce the complete "Zero" chat model.

Frequently Asked Questions

The exact filename is pretrain_hq.jsonl. This specific file is referenced throughout README_en.md and dataset/dataset.md as the high-quality subset required for reproducing the baseline model. Do not substitute with the full JiangShu dataset unless you intend to conduct extended training experiments.

Where does the MiniMind pre-training data originate from?

The data originates from the JiangShu large-model dataset, a comprehensive Chinese text corpus. The authors filtered this source to create pretrain_hq.jsonl, retaining only high-quality samples under 512 characters to optimize for MiniMind's architectural constraints and training efficiency.

Can I use the pre-training data for non-Chinese language models?

No, pretrain_hq.jsonl contains exclusively Chinese text as explicitly stated in the documentation at line 460 of README_en.md. Attempting to use this data for English or other language pre-training will result in monolingual Chinese output. You would need to source alternative JSONL-formatted corpora in your target language and match the {"text": "..."} schema expected by the training scripts.

How do I verify the pre-training data is correctly formatted before starting training?

Load the file using the 🤗 Datasets library with load_dataset("json", data_files="./dataset/pretrain_hq.jsonl", split="train") and inspect the first few entries. Confirm that each record contains a "text" field with string content and that no line exceeds the 512-character limit mentioned in the dataset documentation. This validation prevents runtime errors in train_pretrain.py during the data loading phase.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →