# How to Switch Between CPU and GPU Execution Providers in Supertonic

> Learn to switch between CPU and GPU execution providers in Supertonic. Install onnxruntime-gpu, update helper.py, and set use_gpu=True for faster model loading.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-05-14

---

**To switch between CPU and GPU execution providers in Supertonic, install `onnxruntime-gpu`, replace the `CPUExecutionProvider` string with `CUDAExecutionProvider` in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), and pass `use_gpu=True` when loading the model.**

Supertonic by supertone-inc uses **ONNX Runtime** as its inference engine for text-to-speech synthesis. While the repository defaults to CPU execution, the architecture supports GPU acceleration through CUDA-enabled execution providers. This guide explains how to modify the Python source code to switch between CPU and GPU execution providers.

## Understanding the Execution Provider Architecture

Supertonic delegates all neural network inference to ONNX Runtime sessions created in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py). The entry point `load_text_to_speech()` controls which execution provider the runtime uses:

```python

# py/helper.py

def load_text_to_speech(onnx_dir: str, use_gpu: bool = False) -> TextToSpeech:
    if use_gpu:
        raise NotImplementedError("GPU mode is not fully tested")
    else:
        providers = ["CPUExecutionProvider"]

```

The `providers` list is a standard ONNX Runtime parameter that determines hardware acceleration. When you call `load_text_to_speech()`, this list propagates to `load_onnx_all()`, which constructs four separate inference sessions:

1. **Duration predictor** (`dp_ort`)
2. **Text encoder** (`text_enc_ort`)
3. **Vector estimator** (`vector_est_ort`)
4. **Vocoder** (`vocoder_ort`)

Each session receives the same `providers` configuration, ensuring consistent hardware usage across the entire inference pipeline.

## Prerequisites for GPU Execution

Before switching to GPU mode, verify your environment meets these requirements:

- **Install the GPU-enabled ONNX Runtime package:**
  
  ```bash
  pip install onnxruntime-gpu
  ```

  
  Windows users may alternatively use `onnxruntime-directml` for DirectML support.

- **NVIDIA GPU with CUDA support** and compatible drivers installed
- **CUDA Toolkit** matching your ONNX Runtime version

## Modifying the Source to Enable GPU Providers

The repository currently raises `NotImplementedError` when `use_gpu=True`. To enable GPU execution, you must edit [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) to replace the exception with the CUDA provider string.

### Step 1: Locate the Provider Assignment

In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) around line 22-27, find the provider selection logic inside `load_text_to_speech()`:

```python
if use_gpu:
    raise NotImplementedError("GPU mode is not fully tested")
else:
    providers = ["CPUExecutionProvider"]

```

### Step 2: Replace with GPU Provider

Modify the code to use `CUDAExecutionProvider` when the GPU flag is active:

```python
if use_gpu:
    providers = ["CUDAExecutionProvider"]
    print("Using GPU for inference")
else:
    providers = ["CPUExecutionProvider"]
    print("Using CPU for inference")

```

This change allows the `providers` list containing `"CUDAExecutionProvider"` to flow into `load_onnx_all()`, where it initializes the four model sessions with GPU acceleration.

## Programmatic Execution Provider Switching

If you prefer not to modify the library source, create a custom loader function that bypasses the `load_text_to_speech()` guard:

```python
import onnxruntime as ort
from helper import load_onnx_all, load_cfgs, load_text_processor, TextToSpeech

def load_tts_gpu(onnx_dir: str):
    """Load Supertonic TTS with GPU execution provider."""
    opts = ort.SessionOptions()
    providers = ["CUDAExecutionProvider"]  # GPU execution

    
    # Load configurations and create sessions directly

    cfgs = load_cfgs(onnx_dir)
    dp_ort, text_enc_ort, vector_est_ort, vocoder_ort = load_onnx_all(
        onnx_dir, opts, providers
    )
    
    text_proc = load_text_processor(onnx_dir)
    return TextToSpeech(cfgs, text_proc, dp_ort, text_enc_ort, vector_est_ort, vocoder_ort)

# Usage

tts = load_tts_gpu("./assets/onnx")

```

This approach directly constructs the `TextToSpeech` object with GPU-enabled sessions while maintaining the CPU fallback capability in the original helper functions.

## Using the Command-Line Interface

The [`example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/example_onnx.py) script includes a `--use-gpu` argument that passes the GPU flag to `load_text_to_speech()`. After applying the patch to [`helper.py`](https://github.com/supertone-inc/supertonic/blob/main/helper.py), run:

```bash

# CPU execution (default)

python example_onnx.py --onnx-dir ./assets/onnx

# GPU execution

python example_onnx.py --use-gpu --onnx-dir ./assets/onnx

```

The script parses the flag and forwards it to the loader:

```python

# py/example_onnx.py

text_to_speech = load_text_to_speech(args.onnx_dir, args.use_gpu)

```

## How Provider Configuration Propagates Through the Pipeline

Understanding the data flow helps debug execution provider issues:

1. **`load_text_to_speech()`** selects the provider list based on the `use_gpu` boolean and passes it to `load_onnx_all()`.

2. **`load_onnx_all()`** (defined around line 90 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)) creates four `ort.InferenceSession` objects using the signature:
   
   ```python
   ort.InferenceSession(model_path, sess_options, providers=providers)
   ```

3. **`TextToSpeech` class** stores these sessions and calls `.run()` on them during inference in the `_infer()` method.

4. **ONNX Runtime** automatically schedules kernel execution on the GPU when `CUDAExecutionProvider` is specified, falling back to CPU only if GPU initialization fails.

## Summary

- **Supertonic** uses ONNX Runtime with execution providers specified in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py).
- **Default execution** uses `CPUExecutionProvider` defined in `load_text_to_speech()`.
- **GPU activation** requires installing `onnxruntime-gpu` and modifying the provider list to `["CUDAExecutionProvider"]`.
- **Four model sessions** (duration, encoder, vector, vocoder) share the same execution provider configuration through `load_onnx_all()`.
- **CLI support** exists via `--use-gpu` in [`example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/example_onnx.py), but requires removing the `NotImplementedError` guard in the source.

## Frequently Asked Questions

### What execution providers does Supertonic support?

Supertonic supports any ONNX Runtime execution provider, including `CPUExecutionProvider`, `CUDAExecutionProvider` for NVIDIA GPUs, and `DirectMLExecutionProvider` for Windows DirectML-compatible hardware. The provider string is passed directly to `ort.InferenceSession` in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), so you can specify any provider supported by your ONNX Runtime installation.

### Why does `load_text_to_speech()` raise NotImplementedError for GPU mode?

The `NotImplementedError` in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) indicates that GPU execution has not been fully tested by the maintainers for the current release. The underlying ONNX Runtime infrastructure fully supports GPU acceleration, but you must manually enable it by replacing the exception with the appropriate provider list as shown in the modification steps above.

### Do all four models use the same execution provider?

Yes. The `load_onnx_all()` function applies the same `providers` list to all four inference sessions (duration predictor, text encoder, vector estimator, and vocoder). This ensures consistent hardware acceleration across the entire synthesis pipeline. If you need mixed CPU/GPU execution, you would need to modify `load_onnx_all()` to accept separate provider lists for each model component.

### How do I verify that GPU inference is actually running?

If the provider list contains `CUDAExecutionProvider` and initialization succeeds, ONNX Runtime automatically uses the GPU. You can verify by monitoring GPU utilization during inference using `nvidia-smi` or by checking that `CUDAExecutionProvider` appears in the available providers list returned by `ort.get_available_providers()`. If the GPU is unavailable, ONNX Runtime will raise an error at session creation rather than silently falling back to CPU.