# Where to Find the Recommended Pre-Training Data for MiniMind: Complete Download and Setup Guide

> Find the recommended pre-training data for MiniMind. Download the pretrain_hq.jsonl file from ModelScope or Hugging Face with our complete guide.

- Repository: [jingyaogong/minimind](https://github.com/jingyaogong/minimind)
- Tags: how-to-guide
- Published: 2026-03-24

---

**The recommended pre-training data for MiniMind is contained in the `pretrain_hq.jsonl` file, a high-quality 1.6 GB subset of the JiangShu large-model dataset available on both ModelScope and Hugging Face.**

MiniMind is an open-source, lightweight large language model designed for rapid reproduction and educational fine-tuning. To replicate the official "Zero" chat model, you must use the **recommended pre-training data for MiniMind**, specifically the curated `pretrain_hq.jsonl` file that the authors extracted for optimal training efficiency. This guide provides the exact download locations, setup procedures, and integration steps required to begin training immediately.

## What Is the Recommended MiniMind Pre-Training Dataset?

According to the [`README_en.md`](https://github.com/jingyaogong/minimind/blob/main/README_en.md) at line 301, the primary pre-training corpus is **`pretrain_hq.jsonl`**. This file represents a filtered, high-quality subset of the **JiangShu large-model dataset**, weighing approximately 1.6 GB.

The dataset contains exclusively Chinese text samples with lengths under 512 characters, making it ideal for the MiniMind architecture's sequence length constraints. For the fastest reproduction of the "Zero" chat model, the authors recommend combining `pretrain_hq.jsonl` with the supervised fine-tuning file `sft_mini_512.jsonl`.

## Official Download Sources

The MiniMind project hosts the pre-training data on two major model hub platforms. Both mirrors contain identical files, so choose based on your regional connectivity.

### ModelScope Mirror

Access the dataset at the ModelScope repository:

```text
https://www.modelscope.cn/datasets/gongjy/minimind_dataset/files

```

Look for the entry named `pretrain_hq.jsonl` (marked with a ✨ icon) in the file list. Download this file along with `sft_mini_512.jsonl` if you intend to complete the full training pipeline.

### Hugging Face Hub

Alternatively, fetch the data from the Hugging Face dataset hub:

```text
https://huggingface.co/datasets/jingyaogong/minimind_dataset

```

The repository structure mirrors the ModelScope layout, containing the same `pretrain_hq.jsonl` file required for the initial pre-training phase.

## Step-by-Step Setup Guide

After downloading the files, you must place them in the correct directory structure for the training scripts to locate them automatically.

### 1. Create the Dataset Directory

From the repository root, execute:

```bash
mkdir -p ./dataset

```

This directory path is hardcoded in the training scripts' default configurations.

### 2. Place the Downloaded Files

Move `pretrain_hq.jsonl` (and optionally `sft_mini_512.jsonl`) into the `./dataset` folder:

```bash
mv ~/Downloads/pretrain_hq.jsonl ./dataset/
mv ~/Downloads/sft_mini_512.jsonl ./dataset/  # Optional, for SFT stage

```

### 3. Verify the Source and Format

As documented in [`README_en.md`](https://github.com/jingyaogong/minimind/blob/main/README_en.md) at line 460, verify that your `pretrain_hq.jsonl` contains Chinese text samples with fewer than 512 characters per entry. Each line must be a valid JSON object with a `"text"` field:

```json
{"text": "这是一个示例文本..."}

```

## Loading and Training with the Dataset

Once the files are in place, you can immediately begin the training process using the provided trainer scripts or load the data programmatically.

### Running Pre-Training

Execute the pre-training script with the explicit dataset path:

```bash
python trainer/train_pretrain.py \
  --model_name minimind \
  --dataset_path ./dataset/pretrain_hq.jsonl \
  --max_seq_len 320

```

The [`train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/train_pretrain.py) script automatically parses the JSONL format and handles the JiangShu dataset's text-only structure.

### Loading with the Hugging Face Datasets Library

For custom preprocessing or inspection, load the data using the 🤗 Datasets library:

```python
from datasets import load_dataset

# Load the pre-training corpus

pretrain = load_dataset(
    "json", 
    data_files="./dataset/pretrain_hq.jsonl", 
    split="train"
)

# Verify the contents

print(f"Number of examples: {len(pretrain)}")
print(pretrain[0]["text"][:200])

```

This approach validates that your `pretrain_hq.jsonl` follows the expected schema before launching expensive training runs.

### Complete Zero Model Reproduction Pipeline

To reproduce the official "Zero" chat model exactly as the authors intended, execute the two-stage training sequence:

```bash

# Stage 1: Pre-training on high-quality corpus

python trainer/train_pretrain.py \
  --dataset_path ./dataset/pretrain_hq.jsonl \
  --max_seq_len 320

# Stage 2: Supervised fine-tuning for chat capabilities

python trainer/train_full_sft.py \
  --dataset_path ./dataset/sft_mini_512.jsonl \
  --max_seq_len 512

```

This combination of `pretrain_hq.jsonl` followed by `sft_mini_512.jsonl` yields the fastest path to the functional chat model according to the repository documentation.

## Dataset Format and Technical Details

The `pretrain_hq.jsonl` file adheres to strict formatting requirements enforced by the training scripts:

- **Format**: JSON Lines (JSONL), with one JSON object per line
- **Schema**: Each object must contain a `"text"` string field containing the raw training text
- **Language**: Chinese text only (sourced from the JiangShu dataset)
- **Length constraint**: Text length must not exceed 512 characters to align with the model's context window
- **File size**: Approximately 1.6 GB uncompressed

These constraints ensure compatibility with the data collation logic inside [`trainer/train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_pretrain.py) and prevent tokenization errors during the training loop.

## Summary

- **`pretrain_hq.jsonl`** is the single recommended pre-training file for MiniMind, extracted from the JiangShu large-model dataset.
- Download the 1.6 GB file from **ModelScope** or **Hugging Face** mirrors listed in the official repository.
- Place the file in the `./dataset` directory at the repository root for automatic detection by training scripts.
- Use [`trainer/train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_pretrain.py) with `--dataset_path ./dataset/pretrain_hq.jsonl` to initiate training.
- Combine with `sft_mini_512.jsonl` using [`trainer/train_full_sft.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_full_sft.py) to reproduce the complete "Zero" chat model.

## Frequently Asked Questions

### What is the exact filename of the recommended pre-training data for MiniMind?

The exact filename is **`pretrain_hq.jsonl`**. This specific file is referenced throughout [`README_en.md`](https://github.com/jingyaogong/minimind/blob/main/README_en.md) and [`dataset/dataset.md`](https://github.com/jingyaogong/minimind/blob/main/dataset/dataset.md) as the high-quality subset required for reproducing the baseline model. Do not substitute with the full JiangShu dataset unless you intend to conduct extended training experiments.

### Where does the MiniMind pre-training data originate from?

The data originates from the **JiangShu large-model dataset**, a comprehensive Chinese text corpus. The authors filtered this source to create `pretrain_hq.jsonl`, retaining only high-quality samples under 512 characters to optimize for MiniMind's architectural constraints and training efficiency.

### Can I use the pre-training data for non-Chinese language models?

No, `pretrain_hq.jsonl` contains **exclusively Chinese text** as explicitly stated in the documentation at line 460 of [`README_en.md`](https://github.com/jingyaogong/minimind/blob/main/README_en.md). Attempting to use this data for English or other language pre-training will result in monolingual Chinese output. You would need to source alternative JSONL-formatted corpora in your target language and match the `{"text": "..."}` schema expected by the training scripts.

### How do I verify the pre-training data is correctly formatted before starting training?

Load the file using the 🤗 Datasets library with `load_dataset("json", data_files="./dataset/pretrain_hq.jsonl", split="train")` and inspect the first few entries. Confirm that each record contains a `"text"` field with string content and that no line exceeds the 512-character limit mentioned in the dataset documentation. This validation prevents runtime errors in [`train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/train_pretrain.py) during the data loading phase.