How to Debug Common Inference Failures in GPT-SoVITS: Fixes for Repetitive Text and Missing Audio
Repetitive text in GPT-SoVITS usually stems from a disabled repetition penalty or restrictive sampling parameters, while missing audio output typically results from empty prompt errors or missing model checkpoints.
GPT-SoVITS generates speech through a three-stage pipeline: text to semantic tokens, semantic tokens to acoustic features, and acoustic features to waveform. When you debug common inference failures such as repetitive text or missing audio output, you need to trace issues through the Text-to-Semantic (T2S) model, the VITS acoustic model, and the vocoder. The high-level entry point for both the Web UI and CLI is the inference() function in GPT_SoVITS/inference_webui_fast.py, which builds an inputs dictionary and passes it to tts_pipeline.run() implemented in GPT_SoVITS/TTS_infer_pack/TTS.py.
Understanding the Inference Pipeline
The synthesis pipeline follows a strict data flow:
- Text → Semantic Tokens: The T2S model generates discrete tokens from input text.
- Semantic Tokens → Acoustic Features: The VITS (or VITS-Pro) model converts tokens into a mel-spectrogram.
- Acoustic Features → Waveform: BigVGAN or Griffin-Lim vocoder synthesizes the final audio.
In GPT_SoVITS/inference_webui_fast.py (lines 50-71), the inference() function constructs the parameter dictionary and yields results:
for item in tts_pipeline.run(inputs):
yield item, actual_seed
Fixing Repetitive Text Output
How Repetition Penalty Works
Repetition is controlled by the repetition penalty applied to logits before sampling. The implementation lives in GPT_SoVITS/AR/models/utils.py (lines 59-71) inside the logits_to_probs function:
if previous_tokens is not None and repetition_penalty != 1.0:
previous_tokens = previous_tokens.long()
score = torch.gather(logits, dim=1, index=previous_tokens)
score = torch.where(
score < 0,
score * repetition_penalty,
score / repetition_penalty,
)
logits.scatter_(dim=1, index=previous_tokens, src=score)
A penalty greater than 1.0 reduces probabilities for tokens that have already appeared, while a value of 1.0 disables the effect entirely and causes loops. The default value is 1.35 according to api_v2.py (lines 41-42), exposed via the UI slider and API parameter.
Adjusting Sampling Parameters
If increasing the repetition penalty does not resolve looping, check the top-k and top-p (nucleus) sampling constraints in GPT_SoVITS/AR/models/utils.py (lines 92-100). Overly restrictive values limit the token pool to only a few candidates, forcing repetition regardless of penalty.
To eliminate repetitive output:
- Increase
repetition_penaltyto 1.2–1.5 (default 1.35) - Raise
top_kto 50–200 (default is often lower) - Set
top_pcloser to 1.0 for broader sampling - Adjust
temperatureto 0.8–1.2 (values near 0 produce bland, repetitive speech)
API Example: Tuning Sampling Parameters
When calling the REST API, explicitly set these parameters to prevent loops:
import requests
payload = {
"text": "Hello world, this is a test.",
"text_lang": "en",
"ref_audio_path": "ref.wav",
"prompt_text": "Hello world",
"prompt_lang": "en",
"top_k": 100,
"top_p": 0.95,
"temperature": 0.9,
"repetition_penalty": 1.4,
}
r = requests.post("http://localhost:9872/inference", json=payload)
Resolving Missing Audio Output
Empty Prompt Errors (NO_PROMPT_ERROR)
For SoVITS-V3, the pipeline requires a non-empty prompt_text. In GPT_SoVITS/TTS_infer_pack/TTS.py (lines 145-149), the code raises NO_PROMPT_ERROR if this condition is violated:
if self.configs.version == "v3" and not prompt_text:
raise NO_PROMPT_ERROR("prompt_text cannot be empty when using SoVITS_V3")
The Web UI catches this exception in inference_webui_fast.py (line 200) and returns no audio. If the UI shows "Generating…" forever but produces silence, check the console for this specific error message.
Fix: Provide a non-empty prompt_text or set ref_text_free=True to disable the prompt requirement.
Model Loading Failures
If checkpoint paths are invalid, tts_pipeline.run() aborts silently. The initialization code in GPT_SoVITS/TTS_infer_pack/TTS.py (lines 1000-1005) logs the missing path in init_vits_weights() but may return None for audio fragments later.
Verification: Check the log output for "Loading VITS weights from…" and confirm files exist in ./GPT_SoVITS/pretrained_models/. Alternatively, use get_weights_names() from config.py (lines 15-18) to validate available checkpoints:
from config import get_weights_names
gpt, sovits = get_weights_names()
print("Available GPT models:", gpt)
print("Available SoVITS models:", sovits)
Configuration Edge Cases
Empty audio arrays (audio.shape == (0,)) occur when return_fragment or streaming_mode are enabled while split_bucket=False. This triggers an early return in TTS.run() (lines 1159-1164 of TTS.py). Additionally, if use_vocoder=False but the model expects VITS-Pro, the vocoder remains uninitialized and fails silently.
Fix: Disable fragment/streaming mode or enable split_bucket to yield full-length audio. Ensure VITS-Pro weights (containing enc_q) are loaded when using the vocoder path.
Debugging Tools and Logging
Enable verbose logging by setting GRADIO_DEBUG=True before launching the Web UI to view the full inputs dictionary in the console.
For CLI debugging, run:
python -m GPT_SoVITS.inference_cli \
--text "重复的文字" \
--text_lang zh \
--ref_audio_path ref.wav \
--prompt_text "提示" \
--repetition_penalty 1.0
Add a temporary debug line in inference_cli.py before tts.run():
print("DEBUG INPUTS:", inputs)
Validate checkpoint integrity by loading weights in a Python REPL:
import torch
torch.load("path/to/model.pth") # Should not raise exceptions
Summary
- Repetitive text indicates insufficient
repetition_penalty(aim for 1.2–1.5) or overly restrictivetop_k/top_psampling. - Missing audio typically stems from
NO_PROMPT_ERRORin SoVITS-V3 (empty prompt text) or missing checkpoint files logged ininit_vits_weights(). - Empty arrays result from misconfigured
return_fragmentandstreaming_modesettings whensplit_bucket=False. - Inspect the
inputsdictionary passed totts_pipeline.run()to verify parameters reach the sampling functions inGPT_SoVITS/AR/models/utils.py.
Frequently Asked Questions
Why does my GPT-SoVITS output repeat the same phrase infinitely?
This occurs when the repetition_penalty parameter is set to 1.0 or lower, effectively disabling the penalty mechanism in logits_to_probs(). Increase the value to 1.2–1.5 via the UI slider or API call. If repetition persists, raise top_k to 50–200 and increase temperature to introduce more token diversity.
What causes the "prompt_text cannot be empty" error during inference?
SoVITS-V3 strictly requires a reference prompt text. According to GPT_SoVITS/TTS_infer_pack/TTS.py (line 145-149), the pipeline raises NO_PROMPT_ERROR if prompt_text is empty when using version "v3". Provide a non-empty prompt or set ref_text_free=True in your API payload to bypass this requirement.
Why is my inference returning an empty audio array with no errors?
Check your return_fragment and streaming_mode settings. In GPT_SoVITS/TTS_infer_pack/TTS.py (lines 1159-1164), enabling these options with split_bucket=False causes the function to return before generating the final waveform. Disable fragment mode or enable split_bucket to ensure the pipeline yields complete audio.
How do I verify that my model checkpoints are loading correctly?
Inspect the console logs for the "Loading VITS weights from…" message emitted by init_vits_weights() in TTS.py. You can also run get_weights_names() from config.py to list available models, or manually test checkpoint integrity with torch.load() in a Python REPL. If paths are incorrect, set the gpt_path and sovits_path environment variables to absolute paths of your .pth files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →