# How to Perform Few-Shot Fine-Tuning with Only 1 Minute of Training Data in GPT-SoVITS

> Learn to perform few-shot fine-tuning with GPT-SoVITS using just 1 minute of audio. Discover the efficient workflow for high-quality voice cloning in minutes. Get started now!

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: how-to-guide
- Published: 2026-03-07

---

**GPT-SoVITS enables high-quality text-to-speech cloning with just 60 seconds of audio by slicing the input into short chunks and leveraging a pre-trained speaker encoder, requiring only minutes of GPU training via the built-in WebUI.**

The RVC-Boss/GPT-SoVITS repository provides a production-ready pipeline for few-shot fine-tuning with only 1 minute of training data, separating the speaker encoder from the text-to-speech decoder to minimize data requirements. This architecture allows the model to adapt to new voices using minimal samples while maintaining synthesis quality. The following workflow processes raw audio through automated slicing, optional denoising, and ASR transcription before launching a streamlined training interface.

## Step 1: Prepare Your Dataset

### Record and Organize the Source Audio

Collect approximately **60 seconds** of clear, single-speaker speech recorded at **24 kHz** for optimal quality. Place the raw audio files in a dedicated folder structure such as `data/raw/` to maintain organization throughout the preprocessing pipeline.

### Create the Dataset List File

Create a plain-text `.list` file where each line follows the strict format:

```text
vocal_path|speaker_name|language|text

```

For few-shot fine-tuning with only 1 minute of training data, you can leave the `text` column empty; the pipeline will populate it automatically during the ASR stage. This file serves as the manifest that maps audio paths to speaker identities and language codes.

## Step 2: Preprocess the Audio

### Slice Audio into Training Chunks

Run the built-in slicer to segment the 1-minute file into **~2-second clips**. The model trains on sub-second samples, so chunking exponentially increases the number of training examples derived from limited data. Execute the script located at **[`tools/audio_slicer.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/audio_slicer.py)**:

```bash
python tools/audio_slicer.py \
    --input_path data/raw \
    --output_root data/sliced \
    --threshold 0.05 \
    --min_length 1.0 \
    --min_interval 0.5 \
    --hop_size 0.01

```

### Optional: Denoise the Samples

If your recordings contain background noise, process the sliced audio through the denoiser before training. The utility at **[`tools/cmd-denoise.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/cmd-denoise.py)** applies signal processing to isolate clean speech:

```bash
python tools/cmd-denoise.py \
    --input_dir data/sliced \
    --output_dir data/denoised

```

## Step 3: Generate Transcriptions with ASR

Accurate text transcriptions are required for supervised training. GPT-SoVITS provides language-specific ASR wrappers in the **`tools/asr/`** directory.

### Chinese Speech Recognition

For Mandarin audio, use the FunASR wrapper:

```bash
python tools/asr/funasr_asr.py -i data/denoised -o data/asr_output

```

### English and Japanese Speech Recognition

For English or Japanese, use the Faster-Whisper implementation with appropriate language flags:

```bash
python tools/asr/fasterwhisper_asr.py \
    -i data/denoised \
    -o data/asr_output \
    -l en \
    -p medium

```

After generation, proofread the `.txt` files to correct any ASR errors, as transcription accuracy directly impacts synthesis quality.

## Step 4: Launch the WebUI and Fine-Tune

Start the training interface by executing **[`webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/webui.py)**:

```bash
python webui.py

```

### Configure Training Parameters

In the **Fine-tune** tab, auto-fill the paths to your preprocessed audio folder (`data/denoised`) and the corresponding `.list` file. For few-shot fine-tuning with only 1 minute of training data, maintain these critical settings from **[`GPT_SoVITS/configs/train.yaml`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/configs/train.yaml)**:

- **`max_sec`**: Set to **54** seconds to limit the maximum sample length per batch
- **`save_every_n_epoch`**: Set to **1** to preserve checkpoints frequently
- **`precision`**: Use **16-mixed** for stable mixed-precision training
- **`gradient_clip`**: Set to **1.0** to prevent gradient explosion on small datasets

The default learning rate of **0.01** and built-in learning rate schedule work optimally for limited data without manual tuning.

### Start the Few-Shot Training

Click **Start Training** to begin the process. Due to the **few-shot architecture**—which freezes the speaker encoder trained on large multi-speaker corpora while fine-tuning only the decoder—convergence occurs within minutes on modern GPUs. The chunk-level training approach treats each 2-second slice as an independent example, effectively multiplying your dataset size.

## Step 5: Run Inference

After training completes, switch to the **Inference** tab or use the CLI fallback at **[`GPT_SoVITS/inference_webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/inference_webui.py)**:

```bash
python GPT_SoVITS/inference_webui.py

```

Select your newly fine-tuned speaker model, enter target text, and synthesize. The **`max_sec`** parameter ensures the model respects the training length constraints during generation.

## Summary

- **Collect** 60 seconds of 24 kHz audio and organize it with a `.list` manifest file
- **Preprocess** using [`tools/audio_slicer.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/audio_slicer.py) to create training chunks, optionally denoising with [`tools/cmd-denoise.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/cmd-denoise.py)
- **Transcribe** via [`tools/asr/funasr_asr.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/asr/funasr_asr.py) (Chinese) or [`tools/asr/fasterwhisper_asr.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/asr/fasterwhisper_asr.py) (English/Japanese)
- **Configure** the WebUI with `max_sec: 54` and mixed-precision settings from [`train.yaml`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/train.yaml)
- **Train** for a few minutes using the few-shot mode to adapt the decoder while keeping the speaker encoder frozen
- **Synthesize** speech by launching [`inference_webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/inference_webui.py) or using the WebUI inference tab

## Frequently Asked Questions

### Can I fine-tune GPT-SoVITS with less than 1 minute of data?

While the repository is optimized for few-shot fine-tuning with only 1 minute of training data, shorter clips may produce unstable results. The architecture relies on the speaker encoder's ability to extract features from approximately 60 seconds of varied phonetic content. Reducing this below 30 seconds often leads to degraded timbre consistency and pronunciation artifacts.

### Why must I slice the audio into 2-second chunks?

The training pipeline in [`tools/audio_slicer.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/audio_slicer.py) creates sub-second to 2-second clips to maximize the number of training iterations per epoch. When you slice 60 seconds of audio into 30 two-second segments, the model sees 30 distinct training examples rather than one long file. This chunk-level training improves convergence and prevents overfitting on temporal inconsistencies within a continuous recording.

### What hardware requirements are needed for 1-minute fine-tuning?

The few-shot fine-tuning workflow runs efficiently on a single modern GPU with 8GB+ VRAM. The mixed-precision training configuration (`precision: 16-mixed`) and gradient clipping (`gradient_clip: 1.0`) in [`GPT_SoVITS/configs/train.yaml`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/configs/train.yaml) keep memory usage low while maintaining stability. Training completes in under 10 minutes on an RTX 3060 or equivalent.

### Do I need to manually transcribe the audio if the ASR makes mistakes?

Yes, proofreading is essential. The ASR wrappers in `tools/asr/` generate initial transcriptions, but errors in the `text` column of your `.list` file directly propagate to synthesis quality. The WebUI includes a text-editing interface specifically for correcting these transcriptions before launching training. Accurate labels ensure the text-to-speech decoder learns proper alignment between phonetics and acoustic features.