# How to Integrate GPT-SoVITS as a Backend Service with External Applications via the API

> Integrate GPT-SoVITS as a backend service using its API. Send requests to the FastAPI server to receive synthesized audio in WAV OGG or AAC format. Learn how to connect external applications.

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: how-to-guide
- Published: 2026-03-07

---

**Start the FastAPI server by running [`api.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py), then send HTTP POST or GET requests to `http://localhost:9880/` with text, reference audio paths, and language parameters to receive synthesized audio in WAV, OGG, or AAC format.**

The RVC-Boss/GPT-SoVITS repository provides a self-contained HTTP API that transforms the text-to-speech engine into a backend service. By launching [`api.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py), you expose a FastAPI server that handles voice cloning requests from any external application without modifying the core inference code.

## Bootstrapping the API Service

### Configuration and Model Loading

The service initialization begins with the `Config` class in [[`config.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py)](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py#L98-L104) (lines 98–104), which reads environment variables, detects available GPUs, and configures half-precision settings. Key configuration fields include:

- **sovits_path** / **gpt_path** – Paths to model checkpoints (fallback to pre-trained defaults)
- **infer_device** – Computation device (`cuda` or `cpu`)
- **api_port** – Default listening port **9880** (override with `-p` flag)

When [`api.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py) starts, it initializes the full model stack:

```python

# Lines 81-89 in api.py

cnhubert.cnhubert_base_path = cnhubert_base_path
tokenizer = AutoTokenizer.from_pretrained(bert_path)
bert_model = AutoModelForMaskedLM.from_pretrained(bert_path)
ssl_model = cnhubert.get_model()

```

The **GPT** and **SoVITS** weights load via `change_gpt_sovits_weights` (lines 92–94), preparing the acoustic and language models for inference.

### Starting the FastAPI Server

The server instantiation occurs at line 99:

```python
app = FastAPI()

```

This creates the application instance that routes all subsequent TTS requests. The server supports both streaming and buffered response modes, enabling integration with real-time applications or batch processing systems.

## API Endpoints and Request Parameters

The GPT-SoVITS API exposes four primary endpoints defined in [[`api.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py)](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py):

| Endpoint | Method | Purpose |
|----------|--------|---------|
| `/` | GET / POST | Text-to-speech synthesis (streaming or file response) |
| `/set_model` | GET / POST | Hot-swap SoVITS and GPT checkpoints without restarting |
| `/change_refer` | GET / POST | Update the default reference audio for inference |
| `/control` | GET / POST | Server lifecycle commands (`restart`, `exit`) |

### Core Synthesis Parameters

All endpoints accept parameters as either **query strings** (GET) or **JSON body** (POST):

- **text** – Target text to synthesize (required)
- **text_language** – Language code of target text: `zh`, `en`, `ja`, etc.
- **refer_wav_path** – Path to reference audio file (optional, falls back to default)
- **prompt_text** – Text spoken in the reference audio
- **prompt_language** – Language code of the reference audio
- **media_type** – Output format: `wav`, `ogg`, or `aac`
- **stream_mode** – `normal` for chunked streaming or `close` for buffered response
- **top_k**, **top_p**, **temperature** – GPT sampling controls
- **speed** – Acoustic decoder speed factor
- **if_sr** – Enable super-resolution (V3 models only)
- **sample_steps** – Diffusion steps for V3/V4 models
- **inp_refs** – List of extra reference wavs for multi-reference inference
- **cut_punc** – Custom punctuation set for sentence splitting

## The Synthesis Pipeline

The **`get_tts_wav`** function ([[`api.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py) lines 330–530](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py#L330-L530)) implements the full inference pipeline:

1. **Reference Processing** – Loads audio via `librosa.load` and pads with silence (`zero_wav`)
2. **Hubert Feature Extraction** – Extracts semantic tokens using `ssl_model` from the reference audio
3. **Text Encoding** – Converts text to phoneme IDs (`phones1`) via `clean_text_inf`, with BERT features for Chinese (`get_bert_inf`)
4. **GPT Inference** – Generates semantic tokens using `t2s_model.model.infer_panel`
5. **Acoustic Decoding** – Converts semantics to mel-spectrogram via `vq_model.decode` (V1/V2) or `vq_model.decode_encp` (V3/V4)
6. **Vocoding** – Generates waveform using BigVGAN (V3) or HiFi-GAN (V4)
7. **Post-Processing** – Optional super-resolution (`audio_sr`) and audio packing via `pack_audio`

When `stream_mode` is set to `"normal"`, the function yields audio chunks immediately as they become available, enabling low-latency playback.

## Client Integration Examples

### Python Client with Requests

```python
import requests

url = "http://localhost:9880/"
payload = {
    "text": "你好，世界！",
    "text_language": "zh",
    "refer_wav_path": "samples/reference.wav",
    "prompt_text": "这是参考音频的文字。",
    "prompt_language": "zh",
    "stream_mode": "close",
    "media_type": "ogg",
    "top_k": 15,
    "temperature": 0.6,
}

response = requests.post(url, json=payload)
with open("output.ogg", "wb") as f:
    f.write(response.content)

```

### cURL One-Shot Synthesis

```bash
curl -X POST http://127.0.0.1:9880/ \
     -H "Content-Type: application/json" \
     -d '{
          "text":"Hello, GPT-SoVITS!",
          "text_language":"en",
          "refer_wav_path":"ref.wav",
          "prompt_text":"Reference speech.",
          "prompt_language":"en",
          "media_type":"wav"
        }' \
     --output hello.wav

```

### Streaming Audio in Real-Time

```bash
curl "http://127.0.0.1:9880/?text=こんにちは&text_language=ja&stream_mode=normal" \
     --output hello.wav

```

The server maintains the connection open, transmitting audio chunks as the `get_tts_wav` generator produces them.

### Runtime Model Swapping

Change checkpoints without restarting the service:

```bash
curl -X POST http://127.0.0.1:9880/set_model \
     -H "Content-Type: application/json" \
     -d '{
          "gpt_model_path":"MyGPT.ckpt",
          "sovits_model_path":"MySoVITS.pth"
        }'

```

## Key Source Files Reference

| File | Role |
|------|------|
| [[`api.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py)](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/api.py) | FastAPI server, request handling (`handle`), and synthesis flow (`get_tts_wav`) |
| [[`config.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py)](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py) | Global configuration, device selection, and model-path defaults |
| [[`module/models.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/module/models.py)](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/module/models.py) | Acoustic model definitions (`Generator`, `SynthesizerTrn`, `SynthesizerTrnV3`) |
| [[`feature_extractor/cnhubert.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/feature_extractor/cnhubert.py)](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/feature_extractor/cnhubert.py) | HuBERT feature extractor (`ssl_model`) |
| [[`text/cleaner.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/cleaner.py)](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/cleaner.py) | Text-to-phoneme conversion |
| [[`tools/audio_sr.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/audio_sr.py)](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/audio_sr.py) | Super-resolution module (`AP_BWE`) for `if_sr` |

## Summary

- **Launch Command** – Execute `python api.py` (default port 9880) to start the FastAPI backend
- **Primary Endpoint** – Send POST requests to `/` with `text`, `text_language`, and optional reference audio parameters
- **Output Formats** – Configure `media_type` for `wav`, `ogg`, or `aac` containers
- **Streaming Support** – Set `stream_mode` to `"normal"` for real-time chunk delivery suitable for voice chat applications
- **Hot Reloading** – Use `/set_model` to swap GPT and SoVITS checkpoints without service interruption
- **Multi-Language** – Support for Chinese, English, Japanese, and other languages via BERT and phoneme processing

## Frequently Asked Questions

### What is the default port for the GPT-SoVITS API server?

The default listening port is **9880**, defined in the `Config` class within [[`config.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py)](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/config.py). You can override this at startup using the `-p` flag, for example: `python api.py -p 9999`.

### Can I stream audio responses instead of waiting for the full file?

Yes. Set the query parameter `stream_mode=normal` in your GET request or include `"stream_mode": "normal"` in your POST JSON body. This activates the generator in `get_tts_wav`, which yields audio chunks as they are synthesized, enabling playback to begin before generation completes.

### Do I need to provide a reference audio file with every request?

No. While you can pass `refer_wav_path` and `prompt_text` per request, you can also set a default reference using the `/change_refer` endpoint. If no reference is provided in a request, the API falls back to the globally configured default reference audio.

### How do I switch between different voice models without stopping the server?

Send a POST request to the `/set_model` endpoint with `gpt_model_path` and `sovits_model_path` parameters. This triggers `change_gpt_sovits_weights` to reload the checkpoints in memory, allowing you to serve different voices or languages from the same running process.