How to Perform Few-Shot Fine-Tuning with Only 1 Minute of Training Data in GPT-SoVITS

GPT-SoVITS enables high-quality text-to-speech cloning with just 60 seconds of audio by slicing the input into short chunks and leveraging a pre-trained speaker encoder, requiring only minutes of GPU training via the built-in WebUI.

The RVC-Boss/GPT-SoVITS repository provides a production-ready pipeline for few-shot fine-tuning with only 1 minute of training data, separating the speaker encoder from the text-to-speech decoder to minimize data requirements. This architecture allows the model to adapt to new voices using minimal samples while maintaining synthesis quality. The following workflow processes raw audio through automated slicing, optional denoising, and ASR transcription before launching a streamlined training interface.

Step 1: Prepare Your Dataset

Record and Organize the Source Audio

Collect approximately 60 seconds of clear, single-speaker speech recorded at 24 kHz for optimal quality. Place the raw audio files in a dedicated folder structure such as data/raw/ to maintain organization throughout the preprocessing pipeline.

Create the Dataset List File

Create a plain-text .list file where each line follows the strict format:

vocal_path|speaker_name|language|text

For few-shot fine-tuning with only 1 minute of training data, you can leave the text column empty; the pipeline will populate it automatically during the ASR stage. This file serves as the manifest that maps audio paths to speaker identities and language codes.

Step 2: Preprocess the Audio

Slice Audio into Training Chunks

Run the built-in slicer to segment the 1-minute file into ~2-second clips. The model trains on sub-second samples, so chunking exponentially increases the number of training examples derived from limited data. Execute the script located at tools/audio_slicer.py:

python tools/audio_slicer.py \
    --input_path data/raw \
    --output_root data/sliced \
    --threshold 0.05 \
    --min_length 1.0 \
    --min_interval 0.5 \
    --hop_size 0.01

Optional: Denoise the Samples

If your recordings contain background noise, process the sliced audio through the denoiser before training. The utility at tools/cmd-denoise.py applies signal processing to isolate clean speech:

python tools/cmd-denoise.py \
    --input_dir data/sliced \
    --output_dir data/denoised

Step 3: Generate Transcriptions with ASR

Accurate text transcriptions are required for supervised training. GPT-SoVITS provides language-specific ASR wrappers in the tools/asr/ directory.

Chinese Speech Recognition

For Mandarin audio, use the FunASR wrapper:

python tools/asr/funasr_asr.py -i data/denoised -o data/asr_output

English and Japanese Speech Recognition

For English or Japanese, use the Faster-Whisper implementation with appropriate language flags:

python tools/asr/fasterwhisper_asr.py \
    -i data/denoised \
    -o data/asr_output \
    -l en \
    -p medium

After generation, proofread the .txt files to correct any ASR errors, as transcription accuracy directly impacts synthesis quality.

Step 4: Launch the WebUI and Fine-Tune

Start the training interface by executing webui.py:

python webui.py

Configure Training Parameters

In the Fine-tune tab, auto-fill the paths to your preprocessed audio folder (data/denoised) and the corresponding .list file. For few-shot fine-tuning with only 1 minute of training data, maintain these critical settings from GPT_SoVITS/configs/train.yaml:

  • max_sec: Set to 54 seconds to limit the maximum sample length per batch
  • save_every_n_epoch: Set to 1 to preserve checkpoints frequently
  • precision: Use 16-mixed for stable mixed-precision training
  • gradient_clip: Set to 1.0 to prevent gradient explosion on small datasets

The default learning rate of 0.01 and built-in learning rate schedule work optimally for limited data without manual tuning.

Start the Few-Shot Training

Click Start Training to begin the process. Due to the few-shot architecture—which freezes the speaker encoder trained on large multi-speaker corpora while fine-tuning only the decoder—convergence occurs within minutes on modern GPUs. The chunk-level training approach treats each 2-second slice as an independent example, effectively multiplying your dataset size.

Step 5: Run Inference

After training completes, switch to the Inference tab or use the CLI fallback at GPT_SoVITS/inference_webui.py:

python GPT_SoVITS/inference_webui.py

Select your newly fine-tuned speaker model, enter target text, and synthesize. The max_sec parameter ensures the model respects the training length constraints during generation.

Summary

Frequently Asked Questions

Can I fine-tune GPT-SoVITS with less than 1 minute of data?

While the repository is optimized for few-shot fine-tuning with only 1 minute of training data, shorter clips may produce unstable results. The architecture relies on the speaker encoder's ability to extract features from approximately 60 seconds of varied phonetic content. Reducing this below 30 seconds often leads to degraded timbre consistency and pronunciation artifacts.

Why must I slice the audio into 2-second chunks?

The training pipeline in tools/audio_slicer.py creates sub-second to 2-second clips to maximize the number of training iterations per epoch. When you slice 60 seconds of audio into 30 two-second segments, the model sees 30 distinct training examples rather than one long file. This chunk-level training improves convergence and prevents overfitting on temporal inconsistencies within a continuous recording.

What hardware requirements are needed for 1-minute fine-tuning?

The few-shot fine-tuning workflow runs efficiently on a single modern GPU with 8GB+ VRAM. The mixed-precision training configuration (precision: 16-mixed) and gradient clipping (gradient_clip: 1.0) in GPT_SoVITS/configs/train.yaml keep memory usage low while maintaining stability. Training completes in under 10 minutes on an RTX 3060 or equivalent.

Do I need to manually transcribe the audio if the ASR makes mistakes?

Yes, proofreading is essential. The ASR wrappers in tools/asr/ generate initial transcriptions, but errors in the text column of your .list file directly propagate to synthesis quality. The WebUI includes a text-editing interface specifically for correcting these transcriptions before launching training. Accurate labels ensure the text-to-speech decoder learns proper alignment between phonetics and acoustic features.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →