GPT-SoVITS Dataset Format and .list File Parsing: Complete Training Guide
GPT-SoVITS requires a pipe-delimited .list annotation file where each line contains exactly four fields—vocal_path|speaker_name|language|text—to map audio files with their transcriptions for TTS training.
The RVC-Boss/GPT-SoVITS repository implements a zero-shot text-to-speech synthesis system that relies on strict dataset conventions. Before fine-tuning the SoVITS or GPT modules, you must provide a structured annotation file that the preprocessing pipeline parses to locate audio assets and align text transcriptions. Understanding the exact format and parsing logic in tools/my_utils.py prevents common path resolution errors during dataset preparation.
Expected Dataset Format
GPT-SoVITS consumes a text-to-speech (TTS) annotation file with the .list extension. The repository documentation explicitly defines this format in the README, specifying that each utterance occupies one line with four mandatory fields separated by the pipe character (|).
The Four-Field Structure
Every line in the .list file must follow this exact order:
vocal_path|speaker_name|language|text
- vocal_path: Full or relative filesystem path to the source WAV audio file
- speaker_name: Arbitrary string identifier for the speaker (supports multi-speaker datasets)
- language: Two-letter language code from the supported set—
zh(Chinese),ja(Japanese),en(English),ko(Korean), oryue(Cantonese) - text: The verbatim transcription that the model should learn to synthesize
As documented in the repository README, a valid entry appears as:
D:\GPT-SoVITS\xxx/xxx.wav|xxx|en|I like playing Genshin.
How the .list File is Parsed
When you initiate dataset processing through the WebUI or CLI, the system validates and ingests the .list file through a specific entry point before generating training-ready artifacts.
Initial Validation in check_details
The function check_details in tools/my_utils.py (lines 90–104) performs the first-pass parsing. This utility opens the annotation file, reads the first line to validate the format, and extracts the audio path using the pipe delimiter:
if is_dataset_processing:
list_path, audio_path = path_list
# ... validation omitted ...
with open(list_path, "r", encoding="utf8") as f:
line = f.readline().strip("\n").split("\n")
wav_name, _, __, ___ = line[0].split("|") # split on '|'
wav_name = clean_path(wav_name)
if audio_path not in ("", None):
wav_name = os.path.basename(wav_name) # keep only filename
wav_path = f"{audio_path}/{wav_name}"
else:
wav_path = wav_name
Key parsing behaviors:
- The code splits each line on the
|character to isolate the four fields - If you supply a separate audio directory during processing,
check_detailsdiscards the original directory structure and prepends the supplied root to the filename only - The function handles path cleaning to ensure cross-platform compatibility
Preprocessing Pipeline Output
The .list file itself is not consumed directly by the training loops. Instead, the dataset preparation stage generates intermediate files derived from the annotation data:
2-name2text.txt: Maps audio identifiers to cleaned text transcriptions4-cnhubert/: Directory containing HuBERT features extracted from the audio paths5-wav32k/: Resampled 32kHz audio cache
These artifacts are built using the path and text information extracted during the initial .list parsing phase.
Training Data Consumption
During the actual training phase, the data loader operates on the preprocessed derivatives rather than the raw annotation file.
TextAudioSpeakerLoader Implementation
The TextAudioSpeakerLoader class in GPT_SoVITS/module/data_utils.py loads the prepared dataset. This loader reads the per-speaker text mappings (from 2-name2text.txt) and aligns them with the corresponding audio, HuBERT, and semantic token data that were generated based on the original .list entries.
This architecture separates the annotation parsing (handled in tools/my_utils.py) from the training ingestion (handled in module/data_utils.py), ensuring that path resolution and text normalization occur once during dataset preparation rather than during every training epoch.
Practical Implementation
Manual Parsing Example
To programmatically validate or preview your dataset before launching the full pipeline, use this Python snippet that mirrors the logic in check_details:
import os
def parse_list(list_path: str, audio_root: str = "") -> dict:
"""Parse GPT-SoVITS .list format and return first entry details."""
with open(list_path, "r", encoding="utf8") as f:
first_line = f.readline().strip()
wav_path, speaker, lang, text = first_line.split("|")
wav_path = wav_path.strip()
if audio_root:
# Replicate the UI's path rewriting behavior
wav_path = os.path.join(audio_root, os.path.basename(wav_path))
return {
"wav_path": wav_path,
"speaker": speaker,
"language": lang,
"text": text
}
# Usage
entry = parse_list("dataset.list", audio_root="raw/audio")
print(f"Audio: {entry['wav_path']}, Text: {entry['text']}")
Command-Line Workflow
- Prepare the annotation following the
vocal_path|speaker_name|language|textconvention - Launch dataset processing via the WebUI "Dataset Processing" tab or by invoking the underlying utilities that call
check_details - Execute training after preprocessing completes:
python GPT_SoVITS/s1_train.py --config configs/config.json
Summary
- GPT-SoVITS requires a
.listfile with four pipe-delimited fields: audio path, speaker ID, language code, and transcription text - Path resolution occurs in
tools/my_utils.pyvia thecheck_detailsfunction, which optionally rewrites paths against a user-supplied audio root directory - Training does not read the
.listfile directly; instead,TextAudioSpeakerLoaderconsumes preprocessed artifacts (2-name2text.txt, HuBERT features, and resampled audio) generated from the original annotation - Supported language codes are strictly limited to
zh,ja,en,ko, andyueas defined in the repository documentation
Frequently Asked Questions
What delimiter does GPT-SoVITS use in the .list annotation file?
GPT-SoVITS uses the pipe character (|) as the field delimiter. Each line must contain exactly three pipes separating the four required fields: vocal_path|speaker_name|language|text. The parsing logic in tools/my_utils.py explicitly calls split("|") to tokenize each line.
Which language codes are supported in the dataset annotation?
According to the GPT-SoVITS README and source validation, the supported two-letter language codes are zh (Chinese), ja (Japanese), en (English), ko (Korean), and yue (Cantonese). Using unsupported codes will cause errors during the text processing stage.
Why does my audio path get rewritten during dataset processing?
If you specify an audio directory path in the WebUI or CLI arguments, the check_details function in tools/my_utils.py strips the original directory structure from the vocal_path field and prepends your specified root directory. This behavior ensures that audio files can be relocated without editing the .list file, but it requires that all WAV files exist directly in the supplied root folder.
Does the training script read the .list file directly?
No. The training scripts (s1_train.py, s2_train.py) and the TextAudioSpeakerLoader class in GPT_SoVITS/module/data_utils.py operate on preprocessed cache files (specifically 2-name2text.txt and feature directories) rather than the raw .list annotation. The .list file is only parsed during the initial dataset preparation phase to generate these intermediate artifacts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →