How to Debug Common Inference Failures in GPT-SoVITS: Fixes for Repetitive Text and Missing Audio

Repetitive text in GPT-SoVITS usually stems from a disabled repetition penalty or restrictive sampling parameters, while missing audio output typically results from empty prompt errors or missing model checkpoints.

GPT-SoVITS generates speech through a three-stage pipeline: text to semantic tokens, semantic tokens to acoustic features, and acoustic features to waveform. When you debug common inference failures such as repetitive text or missing audio output, you need to trace issues through the Text-to-Semantic (T2S) model, the VITS acoustic model, and the vocoder. The high-level entry point for both the Web UI and CLI is the inference() function in GPT_SoVITS/inference_webui_fast.py, which builds an inputs dictionary and passes it to tts_pipeline.run() implemented in GPT_SoVITS/TTS_infer_pack/TTS.py.

Understanding the Inference Pipeline

The synthesis pipeline follows a strict data flow:

  1. Text → Semantic Tokens: The T2S model generates discrete tokens from input text.
  2. Semantic Tokens → Acoustic Features: The VITS (or VITS-Pro) model converts tokens into a mel-spectrogram.
  3. Acoustic Features → Waveform: BigVGAN or Griffin-Lim vocoder synthesizes the final audio.

In GPT_SoVITS/inference_webui_fast.py (lines 50-71), the inference() function constructs the parameter dictionary and yields results:

for item in tts_pipeline.run(inputs):
    yield item, actual_seed

Fixing Repetitive Text Output

How Repetition Penalty Works

Repetition is controlled by the repetition penalty applied to logits before sampling. The implementation lives in GPT_SoVITS/AR/models/utils.py (lines 59-71) inside the logits_to_probs function:

if previous_tokens is not None and repetition_penalty != 1.0:
    previous_tokens = previous_tokens.long()
    score = torch.gather(logits, dim=1, index=previous_tokens)
    score = torch.where(
        score < 0,
        score * repetition_penalty,
        score / repetition_penalty,
    )
    logits.scatter_(dim=1, index=previous_tokens, src=score)

A penalty greater than 1.0 reduces probabilities for tokens that have already appeared, while a value of 1.0 disables the effect entirely and causes loops. The default value is 1.35 according to api_v2.py (lines 41-42), exposed via the UI slider and API parameter.

Adjusting Sampling Parameters

If increasing the repetition penalty does not resolve looping, check the top-k and top-p (nucleus) sampling constraints in GPT_SoVITS/AR/models/utils.py (lines 92-100). Overly restrictive values limit the token pool to only a few candidates, forcing repetition regardless of penalty.

To eliminate repetitive output:

  • Increase repetition_penalty to 1.2–1.5 (default 1.35)
  • Raise top_k to 50–200 (default is often lower)
  • Set top_p closer to 1.0 for broader sampling
  • Adjust temperature to 0.8–1.2 (values near 0 produce bland, repetitive speech)

API Example: Tuning Sampling Parameters

When calling the REST API, explicitly set these parameters to prevent loops:

import requests

payload = {
    "text": "Hello world, this is a test.",
    "text_lang": "en",
    "ref_audio_path": "ref.wav",
    "prompt_text": "Hello world",
    "prompt_lang": "en",
    "top_k": 100,
    "top_p": 0.95,
    "temperature": 0.9,
    "repetition_penalty": 1.4,
}
r = requests.post("http://localhost:9872/inference", json=payload)

Resolving Missing Audio Output

Empty Prompt Errors (NO_PROMPT_ERROR)

For SoVITS-V3, the pipeline requires a non-empty prompt_text. In GPT_SoVITS/TTS_infer_pack/TTS.py (lines 145-149), the code raises NO_PROMPT_ERROR if this condition is violated:

if self.configs.version == "v3" and not prompt_text:
    raise NO_PROMPT_ERROR("prompt_text cannot be empty when using SoVITS_V3")

The Web UI catches this exception in inference_webui_fast.py (line 200) and returns no audio. If the UI shows "Generating…" forever but produces silence, check the console for this specific error message.

Fix: Provide a non-empty prompt_text or set ref_text_free=True to disable the prompt requirement.

Model Loading Failures

If checkpoint paths are invalid, tts_pipeline.run() aborts silently. The initialization code in GPT_SoVITS/TTS_infer_pack/TTS.py (lines 1000-1005) logs the missing path in init_vits_weights() but may return None for audio fragments later.

Verification: Check the log output for "Loading VITS weights from…" and confirm files exist in ./GPT_SoVITS/pretrained_models/. Alternatively, use get_weights_names() from config.py (lines 15-18) to validate available checkpoints:

from config import get_weights_names
gpt, sovits = get_weights_names()
print("Available GPT models:", gpt)
print("Available SoVITS models:", sovits)

Configuration Edge Cases

Empty audio arrays (audio.shape == (0,)) occur when return_fragment or streaming_mode are enabled while split_bucket=False. This triggers an early return in TTS.run() (lines 1159-1164 of TTS.py). Additionally, if use_vocoder=False but the model expects VITS-Pro, the vocoder remains uninitialized and fails silently.

Fix: Disable fragment/streaming mode or enable split_bucket to yield full-length audio. Ensure VITS-Pro weights (containing enc_q) are loaded when using the vocoder path.

Debugging Tools and Logging

Enable verbose logging by setting GRADIO_DEBUG=True before launching the Web UI to view the full inputs dictionary in the console.

For CLI debugging, run:

python -m GPT_SoVITS.inference_cli \
    --text "重复的文字" \
    --text_lang zh \
    --ref_audio_path ref.wav \
    --prompt_text "提示" \
    --repetition_penalty 1.0

Add a temporary debug line in inference_cli.py before tts.run():

print("DEBUG INPUTS:", inputs)

Validate checkpoint integrity by loading weights in a Python REPL:

import torch
torch.load("path/to/model.pth")  # Should not raise exceptions

Summary

  • Repetitive text indicates insufficient repetition_penalty (aim for 1.2–1.5) or overly restrictive top_k/top_p sampling.
  • Missing audio typically stems from NO_PROMPT_ERROR in SoVITS-V3 (empty prompt text) or missing checkpoint files logged in init_vits_weights().
  • Empty arrays result from misconfigured return_fragment and streaming_mode settings when split_bucket=False.
  • Inspect the inputs dictionary passed to tts_pipeline.run() to verify parameters reach the sampling functions in GPT_SoVITS/AR/models/utils.py.

Frequently Asked Questions

Why does my GPT-SoVITS output repeat the same phrase infinitely?

This occurs when the repetition_penalty parameter is set to 1.0 or lower, effectively disabling the penalty mechanism in logits_to_probs(). Increase the value to 1.2–1.5 via the UI slider or API call. If repetition persists, raise top_k to 50–200 and increase temperature to introduce more token diversity.

What causes the "prompt_text cannot be empty" error during inference?

SoVITS-V3 strictly requires a reference prompt text. According to GPT_SoVITS/TTS_infer_pack/TTS.py (line 145-149), the pipeline raises NO_PROMPT_ERROR if prompt_text is empty when using version "v3". Provide a non-empty prompt or set ref_text_free=True in your API payload to bypass this requirement.

Why is my inference returning an empty audio array with no errors?

Check your return_fragment and streaming_mode settings. In GPT_SoVITS/TTS_infer_pack/TTS.py (lines 1159-1164), enabling these options with split_bucket=False causes the function to return before generating the final waveform. Disable fragment mode or enable split_bucket to ensure the pipeline yields complete audio.

How do I verify that my model checkpoints are loading correctly?

Inspect the console logs for the "Loading VITS weights from…" message emitted by init_vits_weights() in TTS.py. You can also run get_weights_names() from config.py to list available models, or manually test checkpoint integrity with torch.load() in a Python REPL. If paths are incorrect, set the gpt_path and sovits_path environment variables to absolute paths of your .pth files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →