How to Integrate GPT-SoVITS as a Backend Service with External Applications via the API

Start the FastAPI server by running api.py, then send HTTP POST or GET requests to http://localhost:9880/ with text, reference audio paths, and language parameters to receive synthesized audio in WAV, OGG, or AAC format.

The RVC-Boss/GPT-SoVITS repository provides a self-contained HTTP API that transforms the text-to-speech engine into a backend service. By launching api.py, you expose a FastAPI server that handles voice cloning requests from any external application without modifying the core inference code.

Bootstrapping the API Service

Configuration and Model Loading

The service initialization begins with the Config class in [config.py](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py#L98-L104) (lines 98–104), which reads environment variables, detects available GPUs, and configures half-precision settings. Key configuration fields include:

  • sovits_path / gpt_path – Paths to model checkpoints (fallback to pre-trained defaults)
  • infer_device – Computation device (cuda or cpu)
  • api_port – Default listening port 9880 (override with -p flag)

When api.py starts, it initializes the full model stack:


# Lines 81-89 in api.py

cnhubert.cnhubert_base_path = cnhubert_base_path
tokenizer = AutoTokenizer.from_pretrained(bert_path)
bert_model = AutoModelForMaskedLM.from_pretrained(bert_path)
ssl_model = cnhubert.get_model()

The GPT and SoVITS weights load via change_gpt_sovits_weights (lines 92–94), preparing the acoustic and language models for inference.

Starting the FastAPI Server

The server instantiation occurs at line 99:

app = FastAPI()

This creates the application instance that routes all subsequent TTS requests. The server supports both streaming and buffered response modes, enabling integration with real-time applications or batch processing systems.

API Endpoints and Request Parameters

The GPT-SoVITS API exposes four primary endpoints defined in [api.py](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py):

Endpoint Method Purpose
/ GET / POST Text-to-speech synthesis (streaming or file response)
/set_model GET / POST Hot-swap SoVITS and GPT checkpoints without restarting
/change_refer GET / POST Update the default reference audio for inference
/control GET / POST Server lifecycle commands (restart, exit)

Core Synthesis Parameters

All endpoints accept parameters as either query strings (GET) or JSON body (POST):

  • text – Target text to synthesize (required)
  • text_language – Language code of target text: zh, en, ja, etc.
  • refer_wav_path – Path to reference audio file (optional, falls back to default)
  • prompt_text – Text spoken in the reference audio
  • prompt_language – Language code of the reference audio
  • media_type – Output format: wav, ogg, or aac
  • stream_mode – normal for chunked streaming or close for buffered response
  • top_k, top_p, temperature – GPT sampling controls
  • speed – Acoustic decoder speed factor
  • if_sr – Enable super-resolution (V3 models only)
  • sample_steps – Diffusion steps for V3/V4 models
  • inp_refs – List of extra reference wavs for multi-reference inference
  • cut_punc – Custom punctuation set for sentence splitting

The Synthesis Pipeline

The get_tts_wav function ([api.py lines 330–530](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py#L330-L530)) implements the full inference pipeline:

  1. Reference Processing – Loads audio via librosa.load and pads with silence (zero_wav)
  2. Hubert Feature Extraction – Extracts semantic tokens using ssl_model from the reference audio
  3. Text Encoding – Converts text to phoneme IDs (phones1) via clean_text_inf, with BERT features for Chinese (get_bert_inf)
  4. GPT Inference – Generates semantic tokens using t2s_model.model.infer_panel
  5. Acoustic Decoding – Converts semantics to mel-spectrogram via vq_model.decode (V1/V2) or vq_model.decode_encp (V3/V4)
  6. Vocoding – Generates waveform using BigVGAN (V3) or HiFi-GAN (V4)
  7. Post-Processing – Optional super-resolution (audio_sr) and audio packing via pack_audio

When stream_mode is set to "normal", the function yields audio chunks immediately as they become available, enabling low-latency playback.

Client Integration Examples

Python Client with Requests

import requests

url = "http://localhost:9880/"
payload = {
    "text": "你好,世界!",
    "text_language": "zh",
    "refer_wav_path": "samples/reference.wav",
    "prompt_text": "这是参考音频的文字。",
    "prompt_language": "zh",
    "stream_mode": "close",
    "media_type": "ogg",
    "top_k": 15,
    "temperature": 0.6,
}

response = requests.post(url, json=payload)
with open("output.ogg", "wb") as f:
    f.write(response.content)

cURL One-Shot Synthesis

curl -X POST http://127.0.0.1:9880/ \
     -H "Content-Type: application/json" \
     -d '{
          "text":"Hello, GPT-SoVITS!",
          "text_language":"en",
          "refer_wav_path":"ref.wav",
          "prompt_text":"Reference speech.",
          "prompt_language":"en",
          "media_type":"wav"
        }' \
     --output hello.wav

Streaming Audio in Real-Time

curl "http://127.0.0.1:9880/?text=こんにちは&text_language=ja&stream_mode=normal" \
     --output hello.wav

The server maintains the connection open, transmitting audio chunks as the get_tts_wav generator produces them.

Runtime Model Swapping

Change checkpoints without restarting the service:

curl -X POST http://127.0.0.1:9880/set_model \
     -H "Content-Type: application/json" \
     -d '{
          "gpt_model_path":"MyGPT.ckpt",
          "sovits_model_path":"MySoVITS.pth"
        }'

Key Source Files Reference

File Role
[api.py](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py) FastAPI server, request handling (handle), and synthesis flow (get_tts_wav)
[config.py](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py) Global configuration, device selection, and model-path defaults
[module/models.py](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/module/models.py) Acoustic model definitions (Generator, SynthesizerTrn, SynthesizerTrnV3)
[feature_extractor/cnhubert.py](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/feature_extractor/cnhubert.py) HuBERT feature extractor (ssl_model)
[text/cleaner.py](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/cleaner.py) Text-to-phoneme conversion
[tools/audio_sr.py](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/audio_sr.py) Super-resolution module (AP_BWE) for if_sr

Summary

  • Launch Command – Execute python api.py (default port 9880) to start the FastAPI backend
  • Primary Endpoint – Send POST requests to / with text, text_language, and optional reference audio parameters
  • Output Formats – Configure media_type for wav, ogg, or aac containers
  • Streaming Support – Set stream_mode to "normal" for real-time chunk delivery suitable for voice chat applications
  • Hot Reloading – Use /set_model to swap GPT and SoVITS checkpoints without service interruption
  • Multi-Language – Support for Chinese, English, Japanese, and other languages via BERT and phoneme processing

Frequently Asked Questions

What is the default port for the GPT-SoVITS API server?

The default listening port is 9880, defined in the Config class within [config.py](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py). You can override this at startup using the -p flag, for example: python api.py -p 9999.

Can I stream audio responses instead of waiting for the full file?

Yes. Set the query parameter stream_mode=normal in your GET request or include "stream_mode": "normal" in your POST JSON body. This activates the generator in get_tts_wav, which yields audio chunks as they are synthesized, enabling playback to begin before generation completes.

Do I need to provide a reference audio file with every request?

No. While you can pass refer_wav_path and prompt_text per request, you can also set a default reference using the /change_refer endpoint. If no reference is provided in a request, the API falls back to the globally configured default reference audio.

How do I switch between different voice models without stopping the server?

Send a POST request to the /set_model endpoint with gpt_model_path and sovits_model_path parameters. This triggers change_gpt_sovits_weights to reload the checkpoints in memory, allowing you to serve different voices or languages from the same running process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →