# GPT-SoVITS Dataset Format and .list File Parsing: Complete Training Guide

> Learn the GPT-SoVITS dataset format and how to parse the .list file for effective TTS training. Understand vocal path, speaker name, language, and text mapping.

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: how-to-guide
- Published: 2026-03-07

---

**GPT-SoVITS requires a pipe-delimited `.list` annotation file where each line contains exactly four fields—`vocal_path|speaker_name|language|text`—to map audio files with their transcriptions for TTS training.**

The RVC-Boss/GPT-SoVITS repository implements a zero-shot text-to-speech synthesis system that relies on strict dataset conventions. Before fine-tuning the SoVITS or GPT modules, you must provide a structured annotation file that the preprocessing pipeline parses to locate audio assets and align text transcriptions. Understanding the exact format and parsing logic in [`tools/my_utils.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/my_utils.py) prevents common path resolution errors during dataset preparation.

## Expected Dataset Format

GPT-SoVITS consumes a **text-to-speech (TTS) annotation file** with the `.list` extension. The repository documentation explicitly defines this format in the README, specifying that each utterance occupies one line with four mandatory fields separated by the pipe character (`|`).

### The Four-Field Structure

Every line in the `.list` file must follow this exact order:

```text
vocal_path|speaker_name|language|text

```

- **vocal_path**: Full or relative filesystem path to the source WAV audio file
- **speaker_name**: Arbitrary string identifier for the speaker (supports multi-speaker datasets)
- **language**: Two-letter language code from the supported set—`zh` (Chinese), `ja` (Japanese), `en` (English), `ko` (Korean), or `yue` (Cantonese)
- **text**: The verbatim transcription that the model should learn to synthesize

As documented in the repository README, a valid entry appears as:

```text
D:\GPT-SoVITS\xxx/xxx.wav|xxx|en|I like playing Genshin.

```

## How the .list File is Parsed

When you initiate dataset processing through the WebUI or CLI, the system validates and ingests the `.list` file through a specific entry point before generating training-ready artifacts.

### Initial Validation in check_details

The function **`check_details`** in **[`tools/my_utils.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/my_utils.py)** (lines 90–104) performs the first-pass parsing. This utility opens the annotation file, reads the first line to validate the format, and extracts the audio path using the pipe delimiter:

```python
if is_dataset_processing:
    list_path, audio_path = path_list
    # ... validation omitted ...

    with open(list_path, "r", encoding="utf8") as f:
        line = f.readline().strip("\n").split("\n")
    wav_name, _, __, ___ = line[0].split("|")   # split on '|'

    wav_name = clean_path(wav_name)

    if audio_path not in ("", None):
        wav_name = os.path.basename(wav_name)    # keep only filename

        wav_path = f"{audio_path}/{wav_name}"
    else:
        wav_path = wav_name

```

**Key parsing behaviors:**
- The code splits each line on the `|` character to isolate the four fields
- If you supply a separate **audio directory** during processing, `check_details` discards the original directory structure and prepends the supplied root to the filename only
- The function handles path cleaning to ensure cross-platform compatibility

### Preprocessing Pipeline Output

The `.list` file itself is not consumed directly by the training loops. Instead, the dataset preparation stage generates intermediate files derived from the annotation data:

- **[`2-name2text.txt`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/2-name2text.txt)**: Maps audio identifiers to cleaned text transcriptions
- **`4-cnhubert/`**: Directory containing HuBERT features extracted from the audio paths
- **`5-wav32k/`**: Resampled 32kHz audio cache

These artifacts are built using the path and text information extracted during the initial `.list` parsing phase.

## Training Data Consumption

During the actual training phase, the data loader operates on the preprocessed derivatives rather than the raw annotation file.

### TextAudioSpeakerLoader Implementation

The **`TextAudioSpeakerLoader`** class in **[`GPT_SoVITS/module/data_utils.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/module/data_utils.py)** loads the prepared dataset. This loader reads the per-speaker text mappings (from [`2-name2text.txt`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/2-name2text.txt)) and aligns them with the corresponding audio, HuBERT, and semantic token data that were generated based on the original `.list` entries.

This architecture separates the **annotation parsing** (handled in [`tools/my_utils.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/my_utils.py)) from the **training ingestion** (handled in [`module/data_utils.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/module/data_utils.py)), ensuring that path resolution and text normalization occur once during dataset preparation rather than during every training epoch.

## Practical Implementation

### Manual Parsing Example

To programmatically validate or preview your dataset before launching the full pipeline, use this Python snippet that mirrors the logic in `check_details`:

```python
import os

def parse_list(list_path: str, audio_root: str = "") -> dict:
    """Parse GPT-SoVITS .list format and return first entry details."""
    with open(list_path, "r", encoding="utf8") as f:
        first_line = f.readline().strip()
    
    wav_path, speaker, lang, text = first_line.split("|")
    wav_path = wav_path.strip()
    
    if audio_root:
        # Replicate the UI's path rewriting behavior

        wav_path = os.path.join(audio_root, os.path.basename(wav_path))
    
    return {
        "wav_path": wav_path,
        "speaker": speaker,
        "language": lang,
        "text": text
    }

# Usage

entry = parse_list("dataset.list", audio_root="raw/audio")
print(f"Audio: {entry['wav_path']}, Text: {entry['text']}")

```

### Command-Line Workflow

1. **Prepare the annotation** following the `vocal_path|speaker_name|language|text` convention
2. **Launch dataset processing** via the WebUI "Dataset Processing" tab or by invoking the underlying utilities that call `check_details`
3. **Execute training** after preprocessing completes:

```bash
python GPT_SoVITS/s1_train.py --config configs/config.json

```

## Summary

- **GPT-SoVITS requires a `.list` file** with four pipe-delimited fields: audio path, speaker ID, language code, and transcription text
- **Path resolution occurs in [`tools/my_utils.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/my_utils.py)** via the `check_details` function, which optionally rewrites paths against a user-supplied audio root directory
- **Training does not read the `.list` file directly**; instead, `TextAudioSpeakerLoader` consumes preprocessed artifacts ([`2-name2text.txt`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/2-name2text.txt), HuBERT features, and resampled audio) generated from the original annotation
- **Supported language codes** are strictly limited to `zh`, `ja`, `en`, `ko`, and `yue` as defined in the repository documentation

## Frequently Asked Questions

### What delimiter does GPT-SoVITS use in the .list annotation file?

GPT-SoVITS uses the **pipe character (`|`)** as the field delimiter. Each line must contain exactly three pipes separating the four required fields: `vocal_path|speaker_name|language|text`. The parsing logic in [`tools/my_utils.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/my_utils.py) explicitly calls `split("|")` to tokenize each line.

### Which language codes are supported in the dataset annotation?

According to the GPT-SoVITS README and source validation, the supported two-letter language codes are **zh** (Chinese), **ja** (Japanese), **en** (English), **ko** (Korean), and **yue** (Cantonese). Using unsupported codes will cause errors during the text processing stage.

### Why does my audio path get rewritten during dataset processing?

If you specify an **audio directory** path in the WebUI or CLI arguments, the `check_details` function in [`tools/my_utils.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/my_utils.py) strips the original directory structure from the `vocal_path` field and prepends your specified root directory. This behavior ensures that audio files can be relocated without editing the `.list` file, but it requires that all WAV files exist directly in the supplied root folder.

### Does the training script read the .list file directly?

No. The training scripts ([`s1_train.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/s1_train.py), [`s2_train.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/s2_train.py)) and the `TextAudioSpeakerLoader` class in [`GPT_SoVITS/module/data_utils.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/module/data_utils.py) operate on **preprocessed cache files** (specifically [`2-name2text.txt`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/2-name2text.txt) and feature directories) rather than the raw `.list` annotation. The `.list` file is only parsed during the initial dataset preparation phase to generate these intermediate artifacts.