# Supertonic ONNX Runtime Configuration and Execution Providers: A Cross-Platform Setup Guide

> Master Supertonic ONNX Runtime configuration and execution providers for cross-platform TTS. Our guide simplifies setup across Python, Web, Swift, and Rust, focusing on CPU optimization.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-05-14

---

**Supertonic configures ONNX Runtime with CPU-only execution by default across Python, Web, Swift, and Rust, loading four TTS model components while explicitly disabling GPU support through `NotImplementedError` exceptions until future releases.**

Supertonic is a unified text-to-speech (TTS) pipeline that leverages ONNX Runtime to run inference across multiple languages and platforms. Understanding the Supertonic ONNX Runtime configuration and execution providers is essential for optimizing inference performance, as the library defaults to CPU execution while providing clear extension points for GPU acceleration. The runtime initializes four ONNX sessions—Duration Predictor, Text Encoder, Vector Estimator, and Vocoder—using platform-specific provider selections defined in helper modules for each language binding.

## Platform-Specific Execution Provider Configuration

Supertonic implements ONNX Runtime differently across each supported platform, with distinct approaches to session creation and execution provider selection. The following sections detail the configuration patterns found in the source code.

### Python: Explicit CPU Default with GPU Placeholder

In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), the `load_text_to_speech()` function configures the ONNX Runtime session using `CPUExecutionProvider` by default. The implementation explicitly raises a `NotImplementedError` when `use_gpu=True` is passed, indicating that GPU acceleration remains untested.

```python
def load_text_to_speech(onnx_dir: str, use_gpu: bool = False) -> TextToSpeech:
    opts = ort.SessionOptions()
    if use_gpu:
        raise NotImplementedError("GPU mode is not fully tested")
    else:
        providers = ["CPUExecutionProvider"]
        print("Using CPU for inference")
    cfgs = load_cfgs(onnx_dir)
    dp_ort, text_enc_ort, vector_est_ort, vocoder_ort = load_onnx_all(
        onnx_dir, opts, providers
    )

```

The provider list `["CPUExecutionProvider"]` is passed to `load_onnx_all()`, which creates the four inference sessions for the ONNX models stored in the specified directory.

### Web: WebAssembly and WebGPU Auto-Detection

The JavaScript implementation in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) utilizes ONNX Runtime Web, which automatically selects between `WebAssemblyExecutionProvider` and `WebGPUExecutionProvider` based on browser capabilities. The `loadTextToSpeech()` function accepts optional `sessionOptions` that propagate to the underlying `InferenceSession.create()` calls.

```javascript
export async function loadTextToSpeech(onnxDir, sessionOptions = {}, progressCallback = null) {
    console.log('Using WebAssembly/WebGPU for inference');
    const cfgs = await loadCfgs(onnxDir);
    const modelPaths = [
        { name: 'Duration Predictor', path: `${onnxDir}/duration_predictor.onnx` },
        { name: 'Text Encoder', path: `${onnxDir}/text_encoder.onnx` },
        { name: 'Vector Estimator', path: `${onnxDir}/vector_estimator.onnx` },
        { name: 'Vocoder', path: `${onnxDir}/vocoder.onnx` }
    ];
    const sessions = [];
    for (let i = 0; i < modelPaths.length; i++) {
        if (progressCallback) {
            progressCallback(modelPaths[i].name, i + 1, modelPaths.length);
        }
        const session = await loadOnnx(modelPaths[i].path, sessionOptions);
        sessions.push(session);
    }
    return { textToSpeech: new TextToSpeech(sessions, cfgs), cfgs };
}

```

WebGPU is selected automatically when the compiled binary includes support and the browser exposes the WebGPU API; otherwise, the runtime falls back to WebAssembly.

### Swift: CPU-Only with Explicit Error Handling

The Swift implementation in [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)硬codes CPU execution and throws a descriptive `NSError` when GPU mode is requested. The `loadTextToSpeech(_, useGpu, env)` function checks the `useGpu` boolean before proceeding with CPU session initialization.

```swift
func loadTextToSpeech(_ onnxDir: String, _ useGpu: Bool, _ env: ORTEnv) throws -> TextToSpeech {
    if useGpu {
        throw NSError(domain: "TTS", code: 1,
                      userInfo: [NSLocalizedDescriptionKey: "GPU mode is not supported yet"])
    }
    print("Using CPU for inference\n")
    let cfgs = try loadCfgs(onnxDir)
    // Session creation continues with CPU provider...
}

```

### Rust and Other Languages

The Rust implementation in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) mirrors the Python approach, defaulting to CPU sessions without exposing GPU configuration in the reference implementation. Go, C++, and Java examples follow similar patterns, utilizing CPU execution providers unless platform-specific code adds GPU support externally.

## Execution Provider Selection Logic

Supertonic's execution provider hierarchy follows a deliberate CPU-first strategy:

- **`CPUExecutionProvider`**: Used in Python (explicit), Swift (implicit via ORTEnv), and Rust (default SessionOptions)
- **`WebAssemblyExecutionProvider`**: Fallback for web browsers without WebGPU support
- **`WebGPUExecutionProvider`**: Automatically selected in modern browsers when available
- **Future GPU providers**: Placeholder code exists for `CUDAExecutionProvider` and similar, but raises `NotImplementedError` or throws errors when invoked

The CPU-only default is strategic—the four ONNX models total approximately 70MB and run efficiently on modern desktop CPUs, eliminating GPU dependencies for basic use cases.

## ONNX Runtime Session Creation Flow

All Supertonic language implementations follow an identical five-step initialization sequence:

1. **Load configuration** from [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json), which contains `sample_rate`, `base_chunk_size`, and hyperparameters
2. **Create SessionOptions** with provider-specific flags (e.g., `graph_optimization_level`, `execution_mode`)
3. **Instantiate four InferenceSession objects** for `duration_predictor.onnx`, `text_encoder.onnx`, `vector_estimator.onnx`, and `vocoder.onnx`
4. **Initialize UnicodeProcessor** using [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) for text normalization
5. **Bundle into TextToSpeech class**, exposing `call()` for single utterances and `batch()` for multi-utterance processing

## Practical Code Examples

### Single Utterance Inference in Python

The following example demonstrates loading a TTS pipeline and synthesizing speech using the CPU provider:

```python
from example_onnx import parse_args, load_text_to_speech, load_voice_style
import soundfile as sf, os

args = parse_args()
tts = load_text_to_speech(args.onnx_dir, use_gpu=False)

style = load_voice_style(args.voice_style, verbose=True)
wav, duration = tts(
    text=args.text[0],
    lang=args.lang[0],
    style=style,
    total_step=args.total_step,
    speed=args.speed,
)

os.makedirs(args.save_dir, exist_ok=True)
sf.write(os.path.join(args.save_dir, "out.wav"), wav, tts.sample_rate)

```

### Browser-Based TTS with WebGPU

For web applications, Supertonic automatically leverages WebGPU when available:

```html
<script type="module">
import { loadTextToSpeech, loadVoiceStyle, writeWavFile } from './helper.js';

async function runTTS() {
  const { textToSpeech, cfgs } = await loadTextToSpeech('../assets/onnx');
  const style = await loadVoiceStyle(['../assets/voice_styles/M1.json']);

  const { wav, duration } = await textToSpeech.call(
    "Hello world!", "en", style, 8
  );

  const wavBlob = new Blob([writeWavFile(wav, cfgs.ae.sample_rate)], {type: 'audio/wav'});
  const url = URL.createObjectURL(wavBlob);
  const audio = new Audio(url);
  audio.play();
}
runTTS();
</script>

```

### Swift Command-Line Usage

For macOS or iOS applications, implement the TTS pipeline using the CPU provider as follows:

```swift
import Foundation
import OnnxRuntimeBindings

let args = CommandLineArguments()
let env = try OrtEnv()
let tts = try loadTextToSpeech(args.onnxDir, false, env)

let style = try loadVoiceStyle(args.voiceStylePaths)
let (wav, duration) = try tts.call(
    "Hello Swift TTS!", "en", style, args.totalStep
)

try writeWavFile("result.wav", wav, tts.sampleRate)

```

## Summary

- **Supertonic defaults to CPU execution** across all platforms, utilizing `CPUExecutionProvider` in Python, native CPU sessions in Swift, and WebAssembly/WebGPU auto-detection in browsers
- **GPU support is explicitly disabled** in the reference implementation, with `load_text_to_speech()` and `loadTextToSpeech()` functions raising errors when GPU mode is requested
- **Four ONNX sessions** are created for each TTS pipeline: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder
- **Configuration files** ([`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json) and [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json)) are loaded alongside the ONNX models during session initialization
- **Extension points** are clearly marked in the source code for future implementation of CUDA and other GPU execution providers

## Frequently Asked Questions

### Why does Supertonic disable GPU execution by default?

Supertonic disables GPU execution by default because the ONNX models are optimized for CPU inference and total only approximately 70MB in size. According to the source code in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) and [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift), the maintainers have explicitly marked GPU support as "not fully tested" and "not supported yet" to ensure stability, while leaving clear extension points for future GPU implementations.

### How does the Web implementation choose between WebAssembly and WebGPU?

The Web implementation automatically selects `WebGPUExecutionProvider` when the browser supports the WebGPU API and the ONNX Runtime binary includes WebGPU support; otherwise, it falls back to `WebAssemblyExecutionProvider`. This selection happens transparently within the `loadOnnx()` function calls in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js), requiring no manual provider configuration from the developer.

### What files are required to initialize the ONNX Runtime sessions?

Supertonic requires eight files to initialize the TTS pipeline: four ONNX model files (`duration_predictor.onnx`, `text_encoder.onnx`, `vector_estimator.onnx`, `vocoder.onnx`), the configuration file [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json) containing sample rates and chunk sizes, and [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) for the UnicodeProcessor text normalization component.

### Can I modify the execution providers in the Rust or Go implementations?

Yes, while the reference implementations in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) and the Go examples default to CPU execution, you can modify the `SessionOptions` or equivalent configuration objects to include GPU providers like `CUDAExecutionProvider`. However, you must ensure that the ONNX Runtime binary for your target platform is built with the appropriate execution provider support, as the reference code does not include these configurations by default.