# How to Configure CPU vs GPU Execution Providers in ONNX Runtime in Supertonic

> Learn how to configure CPU vs GPU execution providers in ONNX Runtime with Supertonic. Easily optimize your model performance by selecting the right provider for your hardware.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-14

---

**Supertonic configures ONNX Runtime execution providers by passing a Python list of provider strings—such as `["CPUExecutionProvider"]` or `["CUDAExecutionProvider"]`—to the `providers` parameter of `ort.InferenceSession` within the `load_onnx` helper function.**

The supertone-inc/supertonic repository implements a text-to-speech pipeline that relies on ONNX Runtime for model inference. Understanding how to configure CPU vs GPU execution providers in ONNX Runtime allows you to optimize inference latency across different hardware setups, from standard CPU-only servers to CUDA-enabled workstations.

## Where Execution Providers Are Configured in Supertonic

In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), the `load_text_to_speech` function (around line 22) determines execution provider selection based on the `use_gpu` boolean flag passed from the CLI. The actual session creation occurs in `load_onnx` (lines 84-86), which forwards the provider list unchanged to `ort.InferenceSession(onnx_path, sess_options=opts, providers=providers)`.

All four model components—duration predictor, text encoder, vector estimator, and vocoder—share the same provider list via the `load_onnx_all` wrapper function.

### The CPU Execution Path

When `use_gpu` is `False`, the code constructs a provider list containing only the CPU execution provider:

```python
providers = ["CPUExecutionProvider"]

```

This configuration works out-of-the-box on any host and is the default behavior in [`example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/example_onnx.py) when running without the `--use-gpu` flag.

### The GPU Execution Path (Currently Placeholder)

The current implementation raises `NotImplementedError` when `use_gpu` is `True` because GPU support has not been fully vetted. However, the architecture is already in place to support GPU-specific providers such as `"CUDAExecutionProvider"` for NVIDIA CUDA or `"DirectMLExecutionProvider"` on Windows.

## Implementing CPU Inference

To run inference on CPU (the default configuration), use the standard CLI invocation:

```bash
python py/example_onnx.py \
  --onnx-dir ./assets/onnx \
  --text "Hello world!" \
  --lang en

```

The script will print "Using CPU for inference" (see line 28 of [`helper.py`](https://github.com/supertone-inc/supertonic/blob/main/helper.py)) and instantiate four `InferenceSession` objects, each using `["CPUExecutionProvider"]`.

## Enabling GPU Acceleration (CUDA/DirectML)

To enable GPU acceleration, modify the `load_text_to_speech` function in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) to replace the `NotImplementedError` with a GPU provider configuration:

```python

# In py/helper.py – modify the GPU branch

if use_gpu:
    # For NVIDIA CUDA

    providers = ["CUDAExecutionProvider"]
    
    # Optional: Configure session options for GPU optimization

    opts = ort.SessionOptions()
    opts.enable_mem_pattern = True
else:
    providers = ["CPUExecutionProvider"]

```

Then run the CLI with the `--use-gpu` flag:

```bash
python py/example_onnx.py \
  --onnx-dir ./assets/onnx \
  --use-gpu \
  --text "Hello from the GPU!" \
  --lang en

```

## Using Provider Fallback Chains

Because the provider list is a plain Python list, you can chain multiple providers to create automatic fallback behavior. ONNX Runtime will attempt the first provider and fall back to the next if the first is unavailable:

```python
providers = ["CUDAExecutionProvider", "CPUExecutionProvider"]

```

This pattern ensures that if CUDA is not available on the host system, the models will automatically execute on the CPU without crashing.

## Complete Configuration Example

Here is a complete example showing how to manually override the provider list when loading models programmatically:

```python
import onnxruntime as ort
from helper import load_text_to_speech, load_onnx_all, load_cfgs, load_text_processor

# Define custom provider chain with fallback

providers = ["CUDAExecutionProvider", "CPUExecutionProvider"]

# Configure session options for graph optimization

opts = ort.SessionOptions()
opts.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL

# Load the configuration and text processor

cfg = load_cfgs("./assets/onnx")
text_processor = load_text_processor(cfg.text, cfg.text_tokenizer)

# Instantiate TextToSpeech with custom providers

# Note: This requires modifying load_onnx_all to accept external opts/providers

t2s = load_text_to_speech(
    onnx_dir="./assets/onnx",
    use_gpu=False,  # Ignored when we manually set providers below

)

# Override individual model sessions with GPU providers

t2s.dp_ort = ort.InferenceSession(
    "./assets/onnx/duration_predictor.onnx", 
    sess_options=opts, 
    providers=providers
)
t2s.text_enc_ort = ort.InferenceSession(
    "./assets/onnx/text_encoder.onnx", 
    sess_options=opts, 
    providers=providers
)
t2s.vector_est_ort = ort.InferenceSession(
    "./assets/onnx/vector_estimator.onnx", 
    sess_options=opts, 
    providers=providers
)
t2s.vocoder_ort = ort.InferenceSession(
    "./assets/onnx/vocoder.onnx", 
    sess_options=opts, 
    providers=providers
)

```

## Key Files for Execution Provider Configuration

| File | Purpose |
|------|---------|
| [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) | Contains `load_text_to_speech` (provider selection logic) and `load_onnx` (session instantiation). Lines 84-86 handle the `InferenceSession` creation. |
| [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) | CLI entry point that parses `--use-gpu` (lines 12-16) and passes the flag to `load_text_to_speech`. |
| `assets/onnx/` | Directory containing the four ONNX model files (`duration_predictor.onnx`, `text_encoder.onnx`, `vector_estimator.onnx`, `vocoder.onnx`). |

## Summary

- **Provider selection** happens in `load_text_to_speech` via the `use_gpu` flag, which builds a Python list of provider strings.
- **CPU execution** uses `["CPUExecutionProvider"]` and works universally without additional dependencies.
- **GPU execution** requires replacing the `NotImplementedError` in [`helper.py`](https://github.com/supertone-inc/supertonic/blob/main/helper.py) with provider strings like `["CUDAExecutionProvider"]` or `["DirectMLExecutionProvider"]`.
- **Fallback chains** allow you to specify `["CUDAExecutionProvider", "CPUExecutionProvider"]` so ONNX Runtime automatically falls back to CPU if GPU is unavailable.
- **SessionOptions** can be configured alongside providers to enable graph optimizations and memory patterns for better GPU performance.

## Frequently Asked Questions

### What execution providers does Supertonic currently support out of the box?

Supertonic currently supports **CPUExecutionProvider** without modification. The GPU code path exists in the architecture but raises `NotImplementedError` until explicitly implemented by the user.

### How do I enable CUDA GPU acceleration in Supertonic?

Modify [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) in the `load_text_to_speech` function to replace the `raise NotImplementedError` line with `providers = ["CUDAExecutionProvider"]`, then run [`example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/example_onnx.py) with the `--use-gpu` flag. Ensure you have the CUDA-enabled ONNX Runtime package installed (`onnxruntime-gpu`).

### Can I configure multiple execution providers for automatic fallback?

Yes. Pass a list of providers in order of preference, such as `["CUDAExecutionProvider", "CPUExecutionProvider"]`. ONNX Runtime will attempt CUDA first and automatically fall back to CPU if the CUDA provider is unavailable on the system.

### Where is the execution provider list passed to the ONNX Runtime session?

The provider list is passed to `ort.InferenceSession` in the `load_onnx` function (lines 84-86 of [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)). This function is called by `load_onnx_all`, which initializes all four model components (duration predictor, text encoder, vector estimator, and vocoder) with the same provider configuration.