Supertonic ONNX Runtime Configuration and Execution Providers: A Cross-Platform Setup Guide
Supertonic configures ONNX Runtime with CPU-only execution by default across Python, Web, Swift, and Rust, loading four TTS model components while explicitly disabling GPU support through NotImplementedError exceptions until future releases.
Supertonic is a unified text-to-speech (TTS) pipeline that leverages ONNX Runtime to run inference across multiple languages and platforms. Understanding the Supertonic ONNX Runtime configuration and execution providers is essential for optimizing inference performance, as the library defaults to CPU execution while providing clear extension points for GPU acceleration. The runtime initializes four ONNX sessions—Duration Predictor, Text Encoder, Vector Estimator, and Vocoder—using platform-specific provider selections defined in helper modules for each language binding.
Platform-Specific Execution Provider Configuration
Supertonic implements ONNX Runtime differently across each supported platform, with distinct approaches to session creation and execution provider selection. The following sections detail the configuration patterns found in the source code.
Python: Explicit CPU Default with GPU Placeholder
In py/helper.py, the load_text_to_speech() function configures the ONNX Runtime session using CPUExecutionProvider by default. The implementation explicitly raises a NotImplementedError when use_gpu=True is passed, indicating that GPU acceleration remains untested.
def load_text_to_speech(onnx_dir: str, use_gpu: bool = False) -> TextToSpeech:
opts = ort.SessionOptions()
if use_gpu:
raise NotImplementedError("GPU mode is not fully tested")
else:
providers = ["CPUExecutionProvider"]
print("Using CPU for inference")
cfgs = load_cfgs(onnx_dir)
dp_ort, text_enc_ort, vector_est_ort, vocoder_ort = load_onnx_all(
onnx_dir, opts, providers
)
The provider list ["CPUExecutionProvider"] is passed to load_onnx_all(), which creates the four inference sessions for the ONNX models stored in the specified directory.
Web: WebAssembly and WebGPU Auto-Detection
The JavaScript implementation in web/helper.js utilizes ONNX Runtime Web, which automatically selects between WebAssemblyExecutionProvider and WebGPUExecutionProvider based on browser capabilities. The loadTextToSpeech() function accepts optional sessionOptions that propagate to the underlying InferenceSession.create() calls.
export async function loadTextToSpeech(onnxDir, sessionOptions = {}, progressCallback = null) {
console.log('Using WebAssembly/WebGPU for inference');
const cfgs = await loadCfgs(onnxDir);
const modelPaths = [
{ name: 'Duration Predictor', path: `${onnxDir}/duration_predictor.onnx` },
{ name: 'Text Encoder', path: `${onnxDir}/text_encoder.onnx` },
{ name: 'Vector Estimator', path: `${onnxDir}/vector_estimator.onnx` },
{ name: 'Vocoder', path: `${onnxDir}/vocoder.onnx` }
];
const sessions = [];
for (let i = 0; i < modelPaths.length; i++) {
if (progressCallback) {
progressCallback(modelPaths[i].name, i + 1, modelPaths.length);
}
const session = await loadOnnx(modelPaths[i].path, sessionOptions);
sessions.push(session);
}
return { textToSpeech: new TextToSpeech(sessions, cfgs), cfgs };
}
WebGPU is selected automatically when the compiled binary includes support and the browser exposes the WebGPU API; otherwise, the runtime falls back to WebAssembly.
Swift: CPU-Only with Explicit Error Handling
The Swift implementation in swift/Sources/Helper.swift硬codes CPU execution and throws a descriptive NSError when GPU mode is requested. The loadTextToSpeech(_, useGpu, env) function checks the useGpu boolean before proceeding with CPU session initialization.
func loadTextToSpeech(_ onnxDir: String, _ useGpu: Bool, _ env: ORTEnv) throws -> TextToSpeech {
if useGpu {
throw NSError(domain: "TTS", code: 1,
userInfo: [NSLocalizedDescriptionKey: "GPU mode is not supported yet"])
}
print("Using CPU for inference\n")
let cfgs = try loadCfgs(onnxDir)
// Session creation continues with CPU provider...
}
Rust and Other Languages
The Rust implementation in rust/src/helper.rs mirrors the Python approach, defaulting to CPU sessions without exposing GPU configuration in the reference implementation. Go, C++, and Java examples follow similar patterns, utilizing CPU execution providers unless platform-specific code adds GPU support externally.
Execution Provider Selection Logic
Supertonic's execution provider hierarchy follows a deliberate CPU-first strategy:
CPUExecutionProvider: Used in Python (explicit), Swift (implicit via ORTEnv), and Rust (default SessionOptions)WebAssemblyExecutionProvider: Fallback for web browsers without WebGPU supportWebGPUExecutionProvider: Automatically selected in modern browsers when available- Future GPU providers: Placeholder code exists for
CUDAExecutionProviderand similar, but raisesNotImplementedErroror throws errors when invoked
The CPU-only default is strategic—the four ONNX models total approximately 70MB and run efficiently on modern desktop CPUs, eliminating GPU dependencies for basic use cases.
ONNX Runtime Session Creation Flow
All Supertonic language implementations follow an identical five-step initialization sequence:
- Load configuration from
tts.json, which containssample_rate,base_chunk_size, and hyperparameters - Create SessionOptions with provider-specific flags (e.g.,
graph_optimization_level,execution_mode) - Instantiate four InferenceSession objects for
duration_predictor.onnx,text_encoder.onnx,vector_estimator.onnx, andvocoder.onnx - Initialize UnicodeProcessor using
unicode_indexer.jsonfor text normalization - Bundle into TextToSpeech class, exposing
call()for single utterances andbatch()for multi-utterance processing
Practical Code Examples
Single Utterance Inference in Python
The following example demonstrates loading a TTS pipeline and synthesizing speech using the CPU provider:
from example_onnx import parse_args, load_text_to_speech, load_voice_style
import soundfile as sf, os
args = parse_args()
tts = load_text_to_speech(args.onnx_dir, use_gpu=False)
style = load_voice_style(args.voice_style, verbose=True)
wav, duration = tts(
text=args.text[0],
lang=args.lang[0],
style=style,
total_step=args.total_step,
speed=args.speed,
)
os.makedirs(args.save_dir, exist_ok=True)
sf.write(os.path.join(args.save_dir, "out.wav"), wav, tts.sample_rate)
Browser-Based TTS with WebGPU
For web applications, Supertonic automatically leverages WebGPU when available:
<script type="module">
import { loadTextToSpeech, loadVoiceStyle, writeWavFile } from './helper.js';
async function runTTS() {
const { textToSpeech, cfgs } = await loadTextToSpeech('../assets/onnx');
const style = await loadVoiceStyle(['../assets/voice_styles/M1.json']);
const { wav, duration } = await textToSpeech.call(
"Hello world!", "en", style, 8
);
const wavBlob = new Blob([writeWavFile(wav, cfgs.ae.sample_rate)], {type: 'audio/wav'});
const url = URL.createObjectURL(wavBlob);
const audio = new Audio(url);
audio.play();
}
runTTS();
</script>
Swift Command-Line Usage
For macOS or iOS applications, implement the TTS pipeline using the CPU provider as follows:
import Foundation
import OnnxRuntimeBindings
let args = CommandLineArguments()
let env = try OrtEnv()
let tts = try loadTextToSpeech(args.onnxDir, false, env)
let style = try loadVoiceStyle(args.voiceStylePaths)
let (wav, duration) = try tts.call(
"Hello Swift TTS!", "en", style, args.totalStep
)
try writeWavFile("result.wav", wav, tts.sampleRate)
Summary
- Supertonic defaults to CPU execution across all platforms, utilizing
CPUExecutionProviderin Python, native CPU sessions in Swift, and WebAssembly/WebGPU auto-detection in browsers - GPU support is explicitly disabled in the reference implementation, with
load_text_to_speech()andloadTextToSpeech()functions raising errors when GPU mode is requested - Four ONNX sessions are created for each TTS pipeline: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder
- Configuration files (
tts.jsonandunicode_indexer.json) are loaded alongside the ONNX models during session initialization - Extension points are clearly marked in the source code for future implementation of CUDA and other GPU execution providers
Frequently Asked Questions
Why does Supertonic disable GPU execution by default?
Supertonic disables GPU execution by default because the ONNX models are optimized for CPU inference and total only approximately 70MB in size. According to the source code in py/helper.py and swift/Sources/Helper.swift, the maintainers have explicitly marked GPU support as "not fully tested" and "not supported yet" to ensure stability, while leaving clear extension points for future GPU implementations.
How does the Web implementation choose between WebAssembly and WebGPU?
The Web implementation automatically selects WebGPUExecutionProvider when the browser supports the WebGPU API and the ONNX Runtime binary includes WebGPU support; otherwise, it falls back to WebAssemblyExecutionProvider. This selection happens transparently within the loadOnnx() function calls in web/helper.js, requiring no manual provider configuration from the developer.
What files are required to initialize the ONNX Runtime sessions?
Supertonic requires eight files to initialize the TTS pipeline: four ONNX model files (duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, vocoder.onnx), the configuration file tts.json containing sample rates and chunk sizes, and unicode_indexer.json for the UnicodeProcessor text normalization component.
Can I modify the execution providers in the Rust or Go implementations?
Yes, while the reference implementations in rust/src/helper.rs and the Go examples default to CPU execution, you can modify the SessionOptions or equivalent configuration objects to include GPU providers like CUDAExecutionProvider. However, you must ensure that the ONNX Runtime binary for your target platform is built with the appropriate execution provider support, as the reference code does not include these configurations by default.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →