How to Switch Between CPU and GPU Execution Providers in Supertonic
To switch between CPU and GPU execution providers in Supertonic, install onnxruntime-gpu, replace the CPUExecutionProvider string with CUDAExecutionProvider in py/helper.py, and pass use_gpu=True when loading the model.
Supertonic by supertone-inc uses ONNX Runtime as its inference engine for text-to-speech synthesis. While the repository defaults to CPU execution, the architecture supports GPU acceleration through CUDA-enabled execution providers. This guide explains how to modify the Python source code to switch between CPU and GPU execution providers.
Understanding the Execution Provider Architecture
Supertonic delegates all neural network inference to ONNX Runtime sessions created in py/helper.py. The entry point load_text_to_speech() controls which execution provider the runtime uses:
# py/helper.py
def load_text_to_speech(onnx_dir: str, use_gpu: bool = False) -> TextToSpeech:
if use_gpu:
raise NotImplementedError("GPU mode is not fully tested")
else:
providers = ["CPUExecutionProvider"]
The providers list is a standard ONNX Runtime parameter that determines hardware acceleration. When you call load_text_to_speech(), this list propagates to load_onnx_all(), which constructs four separate inference sessions:
- Duration predictor (
dp_ort) - Text encoder (
text_enc_ort) - Vector estimator (
vector_est_ort) - Vocoder (
vocoder_ort)
Each session receives the same providers configuration, ensuring consistent hardware usage across the entire inference pipeline.
Prerequisites for GPU Execution
Before switching to GPU mode, verify your environment meets these requirements:
-
Install the GPU-enabled ONNX Runtime package:
pip install onnxruntime-gpuWindows users may alternatively use
onnxruntime-directmlfor DirectML support. -
NVIDIA GPU with CUDA support and compatible drivers installed
-
CUDA Toolkit matching your ONNX Runtime version
Modifying the Source to Enable GPU Providers
The repository currently raises NotImplementedError when use_gpu=True. To enable GPU execution, you must edit py/helper.py to replace the exception with the CUDA provider string.
Step 1: Locate the Provider Assignment
In py/helper.py around line 22-27, find the provider selection logic inside load_text_to_speech():
if use_gpu:
raise NotImplementedError("GPU mode is not fully tested")
else:
providers = ["CPUExecutionProvider"]
Step 2: Replace with GPU Provider
Modify the code to use CUDAExecutionProvider when the GPU flag is active:
if use_gpu:
providers = ["CUDAExecutionProvider"]
print("Using GPU for inference")
else:
providers = ["CPUExecutionProvider"]
print("Using CPU for inference")
This change allows the providers list containing "CUDAExecutionProvider" to flow into load_onnx_all(), where it initializes the four model sessions with GPU acceleration.
Programmatic Execution Provider Switching
If you prefer not to modify the library source, create a custom loader function that bypasses the load_text_to_speech() guard:
import onnxruntime as ort
from helper import load_onnx_all, load_cfgs, load_text_processor, TextToSpeech
def load_tts_gpu(onnx_dir: str):
"""Load Supertonic TTS with GPU execution provider."""
opts = ort.SessionOptions()
providers = ["CUDAExecutionProvider"] # GPU execution
# Load configurations and create sessions directly
cfgs = load_cfgs(onnx_dir)
dp_ort, text_enc_ort, vector_est_ort, vocoder_ort = load_onnx_all(
onnx_dir, opts, providers
)
text_proc = load_text_processor(onnx_dir)
return TextToSpeech(cfgs, text_proc, dp_ort, text_enc_ort, vector_est_ort, vocoder_ort)
# Usage
tts = load_tts_gpu("./assets/onnx")
This approach directly constructs the TextToSpeech object with GPU-enabled sessions while maintaining the CPU fallback capability in the original helper functions.
Using the Command-Line Interface
The example_onnx.py script includes a --use-gpu argument that passes the GPU flag to load_text_to_speech(). After applying the patch to helper.py, run:
# CPU execution (default)
python example_onnx.py --onnx-dir ./assets/onnx
# GPU execution
python example_onnx.py --use-gpu --onnx-dir ./assets/onnx
The script parses the flag and forwards it to the loader:
# py/example_onnx.py
text_to_speech = load_text_to_speech(args.onnx_dir, args.use_gpu)
How Provider Configuration Propagates Through the Pipeline
Understanding the data flow helps debug execution provider issues:
-
load_text_to_speech()selects the provider list based on theuse_gpuboolean and passes it toload_onnx_all(). -
load_onnx_all()(defined around line 90 inpy/helper.py) creates fourort.InferenceSessionobjects using the signature:ort.InferenceSession(model_path, sess_options, providers=providers) -
TextToSpeechclass stores these sessions and calls.run()on them during inference in the_infer()method. -
ONNX Runtime automatically schedules kernel execution on the GPU when
CUDAExecutionProvideris specified, falling back to CPU only if GPU initialization fails.
Summary
- Supertonic uses ONNX Runtime with execution providers specified in
py/helper.py. - Default execution uses
CPUExecutionProviderdefined inload_text_to_speech(). - GPU activation requires installing
onnxruntime-gpuand modifying the provider list to["CUDAExecutionProvider"]. - Four model sessions (duration, encoder, vector, vocoder) share the same execution provider configuration through
load_onnx_all(). - CLI support exists via
--use-gpuinexample_onnx.py, but requires removing theNotImplementedErrorguard in the source.
Frequently Asked Questions
What execution providers does Supertonic support?
Supertonic supports any ONNX Runtime execution provider, including CPUExecutionProvider, CUDAExecutionProvider for NVIDIA GPUs, and DirectMLExecutionProvider for Windows DirectML-compatible hardware. The provider string is passed directly to ort.InferenceSession in py/helper.py, so you can specify any provider supported by your ONNX Runtime installation.
Why does load_text_to_speech() raise NotImplementedError for GPU mode?
The NotImplementedError in py/helper.py indicates that GPU execution has not been fully tested by the maintainers for the current release. The underlying ONNX Runtime infrastructure fully supports GPU acceleration, but you must manually enable it by replacing the exception with the appropriate provider list as shown in the modification steps above.
Do all four models use the same execution provider?
Yes. The load_onnx_all() function applies the same providers list to all four inference sessions (duration predictor, text encoder, vector estimator, and vocoder). This ensures consistent hardware acceleration across the entire synthesis pipeline. If you need mixed CPU/GPU execution, you would need to modify load_onnx_all() to accept separate provider lists for each model component.
How do I verify that GPU inference is actually running?
If the provider list contains CUDAExecutionProvider and initialization succeeds, ONNX Runtime automatically uses the GPU. You can verify by monitoring GPU utilization during inference using nvidia-smi or by checking that CUDAExecutionProvider appears in the available providers list returned by ort.get_available_providers(). If the GPU is unavailable, ONNX Runtime will raise an error at session creation rather than silently falling back to CPU.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →