How to Configure CPU vs GPU Execution Providers in ONNX Runtime in Supertonic
Supertonic configures ONNX Runtime execution providers by passing a Python list of provider strings—such as ["CPUExecutionProvider"] or ["CUDAExecutionProvider"]—to the providers parameter of ort.InferenceSession within the load_onnx helper function.
The supertone-inc/supertonic repository implements a text-to-speech pipeline that relies on ONNX Runtime for model inference. Understanding how to configure CPU vs GPU execution providers in ONNX Runtime allows you to optimize inference latency across different hardware setups, from standard CPU-only servers to CUDA-enabled workstations.
Where Execution Providers Are Configured in Supertonic
In py/helper.py, the load_text_to_speech function (around line 22) determines execution provider selection based on the use_gpu boolean flag passed from the CLI. The actual session creation occurs in load_onnx (lines 84-86), which forwards the provider list unchanged to ort.InferenceSession(onnx_path, sess_options=opts, providers=providers).
All four model components—duration predictor, text encoder, vector estimator, and vocoder—share the same provider list via the load_onnx_all wrapper function.
The CPU Execution Path
When use_gpu is False, the code constructs a provider list containing only the CPU execution provider:
providers = ["CPUExecutionProvider"]
This configuration works out-of-the-box on any host and is the default behavior in example_onnx.py when running without the --use-gpu flag.
The GPU Execution Path (Currently Placeholder)
The current implementation raises NotImplementedError when use_gpu is True because GPU support has not been fully vetted. However, the architecture is already in place to support GPU-specific providers such as "CUDAExecutionProvider" for NVIDIA CUDA or "DirectMLExecutionProvider" on Windows.
Implementing CPU Inference
To run inference on CPU (the default configuration), use the standard CLI invocation:
python py/example_onnx.py \
--onnx-dir ./assets/onnx \
--text "Hello world!" \
--lang en
The script will print "Using CPU for inference" (see line 28 of helper.py) and instantiate four InferenceSession objects, each using ["CPUExecutionProvider"].
Enabling GPU Acceleration (CUDA/DirectML)
To enable GPU acceleration, modify the load_text_to_speech function in py/helper.py to replace the NotImplementedError with a GPU provider configuration:
# In py/helper.py – modify the GPU branch
if use_gpu:
# For NVIDIA CUDA
providers = ["CUDAExecutionProvider"]
# Optional: Configure session options for GPU optimization
opts = ort.SessionOptions()
opts.enable_mem_pattern = True
else:
providers = ["CPUExecutionProvider"]
Then run the CLI with the --use-gpu flag:
python py/example_onnx.py \
--onnx-dir ./assets/onnx \
--use-gpu \
--text "Hello from the GPU!" \
--lang en
Using Provider Fallback Chains
Because the provider list is a plain Python list, you can chain multiple providers to create automatic fallback behavior. ONNX Runtime will attempt the first provider and fall back to the next if the first is unavailable:
providers = ["CUDAExecutionProvider", "CPUExecutionProvider"]
This pattern ensures that if CUDA is not available on the host system, the models will automatically execute on the CPU without crashing.
Complete Configuration Example
Here is a complete example showing how to manually override the provider list when loading models programmatically:
import onnxruntime as ort
from helper import load_text_to_speech, load_onnx_all, load_cfgs, load_text_processor
# Define custom provider chain with fallback
providers = ["CUDAExecutionProvider", "CPUExecutionProvider"]
# Configure session options for graph optimization
opts = ort.SessionOptions()
opts.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL
# Load the configuration and text processor
cfg = load_cfgs("./assets/onnx")
text_processor = load_text_processor(cfg.text, cfg.text_tokenizer)
# Instantiate TextToSpeech with custom providers
# Note: This requires modifying load_onnx_all to accept external opts/providers
t2s = load_text_to_speech(
onnx_dir="./assets/onnx",
use_gpu=False, # Ignored when we manually set providers below
)
# Override individual model sessions with GPU providers
t2s.dp_ort = ort.InferenceSession(
"./assets/onnx/duration_predictor.onnx",
sess_options=opts,
providers=providers
)
t2s.text_enc_ort = ort.InferenceSession(
"./assets/onnx/text_encoder.onnx",
sess_options=opts,
providers=providers
)
t2s.vector_est_ort = ort.InferenceSession(
"./assets/onnx/vector_estimator.onnx",
sess_options=opts,
providers=providers
)
t2s.vocoder_ort = ort.InferenceSession(
"./assets/onnx/vocoder.onnx",
sess_options=opts,
providers=providers
)
Key Files for Execution Provider Configuration
| File | Purpose |
|---|---|
py/helper.py |
Contains load_text_to_speech (provider selection logic) and load_onnx (session instantiation). Lines 84-86 handle the InferenceSession creation. |
py/example_onnx.py |
CLI entry point that parses --use-gpu (lines 12-16) and passes the flag to load_text_to_speech. |
assets/onnx/ |
Directory containing the four ONNX model files (duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, vocoder.onnx). |
Summary
- Provider selection happens in
load_text_to_speechvia theuse_gpuflag, which builds a Python list of provider strings. - CPU execution uses
["CPUExecutionProvider"]and works universally without additional dependencies. - GPU execution requires replacing the
NotImplementedErrorinhelper.pywith provider strings like["CUDAExecutionProvider"]or["DirectMLExecutionProvider"]. - Fallback chains allow you to specify
["CUDAExecutionProvider", "CPUExecutionProvider"]so ONNX Runtime automatically falls back to CPU if GPU is unavailable. - SessionOptions can be configured alongside providers to enable graph optimizations and memory patterns for better GPU performance.
Frequently Asked Questions
What execution providers does Supertonic currently support out of the box?
Supertonic currently supports CPUExecutionProvider without modification. The GPU code path exists in the architecture but raises NotImplementedError until explicitly implemented by the user.
How do I enable CUDA GPU acceleration in Supertonic?
Modify py/helper.py in the load_text_to_speech function to replace the raise NotImplementedError line with providers = ["CUDAExecutionProvider"], then run example_onnx.py with the --use-gpu flag. Ensure you have the CUDA-enabled ONNX Runtime package installed (onnxruntime-gpu).
Can I configure multiple execution providers for automatic fallback?
Yes. Pass a list of providers in order of preference, such as ["CUDAExecutionProvider", "CPUExecutionProvider"]. ONNX Runtime will attempt CUDA first and automatically fall back to CPU if the CUDA provider is unavailable on the system.
Where is the execution provider list passed to the ONNX Runtime session?
The provider list is passed to ort.InferenceSession in the load_onnx function (lines 84-86 of py/helper.py). This function is called by load_onnx_all, which initializes all four model components (duration predictor, text encoder, vector estimator, and vocoder) with the same provider configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →