How to Switch Between CPU and GPU Execution Providers in Supertonic

To switch between CPU and GPU execution providers in Supertonic, install onnxruntime-gpu, replace the CPUExecutionProvider string with CUDAExecutionProvider in py/helper.py, and pass use_gpu=True when loading the model.

Supertonic by supertone-inc uses ONNX Runtime as its inference engine for text-to-speech synthesis. While the repository defaults to CPU execution, the architecture supports GPU acceleration through CUDA-enabled execution providers. This guide explains how to modify the Python source code to switch between CPU and GPU execution providers.

Understanding the Execution Provider Architecture

Supertonic delegates all neural network inference to ONNX Runtime sessions created in py/helper.py. The entry point load_text_to_speech() controls which execution provider the runtime uses:


# py/helper.py

def load_text_to_speech(onnx_dir: str, use_gpu: bool = False) -> TextToSpeech:
    if use_gpu:
        raise NotImplementedError("GPU mode is not fully tested")
    else:
        providers = ["CPUExecutionProvider"]

The providers list is a standard ONNX Runtime parameter that determines hardware acceleration. When you call load_text_to_speech(), this list propagates to load_onnx_all(), which constructs four separate inference sessions:

  1. Duration predictor (dp_ort)
  2. Text encoder (text_enc_ort)
  3. Vector estimator (vector_est_ort)
  4. Vocoder (vocoder_ort)

Each session receives the same providers configuration, ensuring consistent hardware usage across the entire inference pipeline.

Prerequisites for GPU Execution

Before switching to GPU mode, verify your environment meets these requirements:

  • Install the GPU-enabled ONNX Runtime package:

    pip install onnxruntime-gpu

    Windows users may alternatively use onnxruntime-directml for DirectML support.

  • NVIDIA GPU with CUDA support and compatible drivers installed

  • CUDA Toolkit matching your ONNX Runtime version

Modifying the Source to Enable GPU Providers

The repository currently raises NotImplementedError when use_gpu=True. To enable GPU execution, you must edit py/helper.py to replace the exception with the CUDA provider string.

Step 1: Locate the Provider Assignment

In py/helper.py around line 22-27, find the provider selection logic inside load_text_to_speech():

if use_gpu:
    raise NotImplementedError("GPU mode is not fully tested")
else:
    providers = ["CPUExecutionProvider"]

Step 2: Replace with GPU Provider

Modify the code to use CUDAExecutionProvider when the GPU flag is active:

if use_gpu:
    providers = ["CUDAExecutionProvider"]
    print("Using GPU for inference")
else:
    providers = ["CPUExecutionProvider"]
    print("Using CPU for inference")

This change allows the providers list containing "CUDAExecutionProvider" to flow into load_onnx_all(), where it initializes the four model sessions with GPU acceleration.

Programmatic Execution Provider Switching

If you prefer not to modify the library source, create a custom loader function that bypasses the load_text_to_speech() guard:

import onnxruntime as ort
from helper import load_onnx_all, load_cfgs, load_text_processor, TextToSpeech

def load_tts_gpu(onnx_dir: str):
    """Load Supertonic TTS with GPU execution provider."""
    opts = ort.SessionOptions()
    providers = ["CUDAExecutionProvider"]  # GPU execution

    
    # Load configurations and create sessions directly

    cfgs = load_cfgs(onnx_dir)
    dp_ort, text_enc_ort, vector_est_ort, vocoder_ort = load_onnx_all(
        onnx_dir, opts, providers
    )
    
    text_proc = load_text_processor(onnx_dir)
    return TextToSpeech(cfgs, text_proc, dp_ort, text_enc_ort, vector_est_ort, vocoder_ort)

# Usage

tts = load_tts_gpu("./assets/onnx")

This approach directly constructs the TextToSpeech object with GPU-enabled sessions while maintaining the CPU fallback capability in the original helper functions.

Using the Command-Line Interface

The example_onnx.py script includes a --use-gpu argument that passes the GPU flag to load_text_to_speech(). After applying the patch to helper.py, run:


# CPU execution (default)

python example_onnx.py --onnx-dir ./assets/onnx

# GPU execution

python example_onnx.py --use-gpu --onnx-dir ./assets/onnx

The script parses the flag and forwards it to the loader:


# py/example_onnx.py

text_to_speech = load_text_to_speech(args.onnx_dir, args.use_gpu)

How Provider Configuration Propagates Through the Pipeline

Understanding the data flow helps debug execution provider issues:

  1. load_text_to_speech() selects the provider list based on the use_gpu boolean and passes it to load_onnx_all().

  2. load_onnx_all() (defined around line 90 in py/helper.py) creates four ort.InferenceSession objects using the signature:

    ort.InferenceSession(model_path, sess_options, providers=providers)
  3. TextToSpeech class stores these sessions and calls .run() on them during inference in the _infer() method.

  4. ONNX Runtime automatically schedules kernel execution on the GPU when CUDAExecutionProvider is specified, falling back to CPU only if GPU initialization fails.

Summary

  • Supertonic uses ONNX Runtime with execution providers specified in py/helper.py.
  • Default execution uses CPUExecutionProvider defined in load_text_to_speech().
  • GPU activation requires installing onnxruntime-gpu and modifying the provider list to ["CUDAExecutionProvider"].
  • Four model sessions (duration, encoder, vector, vocoder) share the same execution provider configuration through load_onnx_all().
  • CLI support exists via --use-gpu in example_onnx.py, but requires removing the NotImplementedError guard in the source.

Frequently Asked Questions

What execution providers does Supertonic support?

Supertonic supports any ONNX Runtime execution provider, including CPUExecutionProvider, CUDAExecutionProvider for NVIDIA GPUs, and DirectMLExecutionProvider for Windows DirectML-compatible hardware. The provider string is passed directly to ort.InferenceSession in py/helper.py, so you can specify any provider supported by your ONNX Runtime installation.

Why does load_text_to_speech() raise NotImplementedError for GPU mode?

The NotImplementedError in py/helper.py indicates that GPU execution has not been fully tested by the maintainers for the current release. The underlying ONNX Runtime infrastructure fully supports GPU acceleration, but you must manually enable it by replacing the exception with the appropriate provider list as shown in the modification steps above.

Do all four models use the same execution provider?

Yes. The load_onnx_all() function applies the same providers list to all four inference sessions (duration predictor, text encoder, vector estimator, and vocoder). This ensures consistent hardware acceleration across the entire synthesis pipeline. If you need mixed CPU/GPU execution, you would need to modify load_onnx_all() to accept separate provider lists for each model component.

How do I verify that GPU inference is actually running?

If the provider list contains CUDAExecutionProvider and initialization succeeds, ONNX Runtime automatically uses the GPU. You can verify by monitoring GPU utilization during inference using nvidia-smi or by checking that CUDAExecutionProvider appears in the available providers list returned by ort.get_available_providers(). If the GPU is unavailable, ONNX Runtime will raise an error at session creation rather than silently falling back to CPU.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →