Key Differences Between Supertonic's Python, Swift, Rust, and Go Implementations
Supertonic provides identical TTS inference pipelines across Python, Swift, Rust, and Go, but each implementation leverages language-specific ecosystems for ONNX bindings, Unicode normalization, tensor operations, and error handling.
Supertonic is a cross-language text-to-speech (TTS) inference library by supertone-inc that implements the same five-step pipeline—Unicode preprocessing, ID conversion, ONNX model inference, latent sampling, and audio post-processing—across four major programming languages. While the high-level algorithm remains constant, the concrete implementations in py/helper.py, swift/Sources/Helper.swift, rust/src/helper.rs, and go/helper.go diverge significantly to accommodate each language's idioms, memory models, and performance characteristics.
ONNX Runtime Bindings by Language
Each implementation uses a different ONNX Runtime binding suited to its ecosystem.
Python: Official ONNX Runtime Wheels
In py/helper.py, Python loads models using the official onnxruntime PyPI package. The load_onnx_all function creates InferenceSession objects for the four required models (duration predictor, text encoder, latent denoiser, and vocoder), reusing them across calls. Tensors are created from NumPy arrays using np.int64 and np.float32 types before being passed to session.run().
Swift: Apple Framework Bindings
The Swift implementation in swift/Sources/Helper.swift uses OnnxRuntimeBindings, an Apple-framework wrapper around ONNX Runtime. The loadTextToSpeech function initializes ORTEnv and creates sessions stored in a TextToSpeech struct. Tensors are constructed via ORTValue using MutableData buffers with explicit Int64 and Float element type annotations.
Rust: The Ort Crate
Rust leverages the ort crate in rust/src/helper.rs to bind to ONNX Runtime. The load_text_to_speech function returns a TextToSpeech struct containing pre-initialized sessions. Tensors are created from ndarray::Array and ndarray::Array3 types, converted to ort::value::Value for inference.
Go: Native CGO Bindings
Go uses github.com/yalue/onnxruntime_go as implemented in go/helper.go. The LoadTextToSpeech function stores sessions in the TextToSpeech struct. Helper functions IntArrayToTensor and ArrayToTensor manually convert Go slices to ort.NewTensor objects, requiring explicit memory management through Destroy calls.
Unicode Text Preprocessing
Text normalization diverges based on each language's Unicode capabilities.
Python employs unicodedata.normalize("NFKD") combined with regex-based emoji removal. The chunk_text function uses regex patterns with an abbreviation list to split sentences, returning Python lists.
Swift uses String.decomposedStringWithCompatibilityMapping for normalization and manual unicodeScalars.filter for emoji removal, as Swift lacks full Unicode regex support. The chunkText function utilizes NSRegularExpression with an abbreviation array.
Rust implements normalization via unicode_normalization::UnicodeNormalization and emoji filtering with the Regex crate. The implementation in rust/src/helper.rs follows the same regex strategy as Python but with Rust's ownership model.
Go relies on golang.org/x/text/unicode/norm for NFKD normalization and compiled regex for emoji ranges. The chunkText function returns Go string slices ([]string).
Tensor Operations and Memory Management
Memory models vary significantly across implementations.
| Language | Tensor Implementation | Memory Model |
|---|---|---|
| Python | NumPy arrays (np.int64, np.float32) |
Heap-allocated, garbage collected |
| Swift | [[[Float]]] 3D arrays and ORTValue |
Value-type structs, manual Data conversions |
| Rust | ndarray::Array3<f32> |
Zero-allocation, owned buffers with Result propagation |
| Go | [][][]float64 and ort.NewTensor |
Manual slice allocation, explicit tensor destruction |
The length_to_mask function illustrates these differences: Python returns a (B, 1, max_len) NumPy float array, Swift returns a [[[Float]]], Rust returns ndarray::Array3<f32>, and Go returns [][][]float64.
Latent Sampling and Random Number Generation
Gaussian noise generation for latent sampling uses language-specific random number generators.
Python uses np.random.randn from NumPy, generating standard normal distributions without explicit seeding in the examples.
Swift implements manual Box-Muller transformation using Float.random for each element individually, as Swift lacks a built-in normal distribution sampler.
Rust utilizes rand_distr::Normal with thread_rng in an explicit sample loop, providing deterministic randomness when needed.
Go employs rand.New(rand.NewSource(time.Now().UnixNano())) combined with the Box-Muller formula, manually calculating normally distributed random values.
Error Handling Patterns
Error handling reflects each language's philosophy.
Python raises exceptions using raise and panic, wrapped in try/except blocks at the caller level.
Swift calls fatalError for unrecoverable conditions (such as invalid language codes), terminating the application.
Rust returns anyhow::Result types from load_text_to_speech and tts.call, propagating errors up the stack with the ? operator.
Go returns explicit error values; callers check if err != nil after operations like LoadTextToSpeech or tts.Call.
Platform Support and Performance
Each implementation targets specific deployment scenarios.
Python offers cross-platform support (Linux, macOS, Windows) but remains single-threaded due to the Global Interpreter Lock (GIL). Best suited for rapid prototyping and research notebooks.
Swift supports only macOS and iOS, leveraging Apple's Accelerate framework and ARC memory management. Ideal for native iOS applications using SwiftUI.
Rust provides cross-platform compatibility with zero-copy memory handling and strict safety guarantees. The implementation in rust/src/helper.rs offers the best raw speed for high-performance back-ends.
Go delivers cross-platform binaries (requiring Go 1.22+) with a concurrency-ready design suitable for cloud micro-services, though the TTS inference itself remains CPU-bound and synchronous.
Implementation Examples
Below are minimal snippets demonstrating single TTS requests in each language, assuming ONNX assets reside in ../assets/onnx and voice styles in ../assets/voice_styles/.
Python Example
from supertonic import load_text_to_speech, load_voice_style
tts = load_text_to_speech("../assets/onnx")
style = load_voice_style(["../assets/voice_styles/M1.json"])[0]
wav, duration = tts("Hello world!", "en", style, total_step=8)
# wav is a NumPy array of float32 samples
Source: py/helper.py – functions load_text_to_speech, load_voice_style, and method TextToSpeech.__call__.
Swift Example
let cfg = try loadCfgs("../assets/onnx")
let env = try ORTEnv()
let tts = try loadTextToSpeech("../assets/onnx", false, env)
let style = try loadVoiceStyle(["../assets/voice_styles/M1.json"])
let (wav, duration) = try tts.call(
"Hello world!", "en", style, totalStep: 8, speed: 1.05, silenceDuration: 0.3)
Source: swift/Sources/Helper.swift – functions loadCfgs, loadTextToSpeech, loadVoiceStyle, and method TextToSpeech.call.
Rust Example
use supertonic::{load_text_to_speech, load_voice_style};
let cfg = load_cfgs("../assets/onnx")?;
let tts = load_text_to_speech("../assets/onnx")?;
let style = load_voice_style(&["../assets/voice_styles/M1.json"], false)?;
let (wav, dur) = tts.call(
"Hello world!", "en", &style, 8, 1.05, 0.3)?;
Source: rust/src/helper.rs – functions load_text_to_speech, load_voice_style, and method TextToSpeech.call.
Go Example
tts, err := LoadTextToSpeech("../assets/onnx", false, cfg)
if err != nil { log.Fatal(err) }
style, err := LoadVoiceStyle([]string{"../assets/voice_styles/M1.json"}, false)
if err != nil { log.Fatal(err) }
wav, dur, err := tts.Call(
"Hello world!", "en", style, 8, 1.05, 0.3)
Source: go/helper.go – functions LoadTextToSpeech, LoadVoiceStyle, and method TextToSpeech.Call.
Summary
- Supertonic implementations share identical TTS pipelines across Python, Swift, Rust, and Go, differing only in language-specific idioms.
- ONNX bindings vary by ecosystem: Python uses
onnxruntime, Swift usesOnnxRuntimeBindings, Rust uses theortcrate, and Go usesgithub.com/yalue/onnxruntime_go. - Memory management ranges from garbage collection (Python) to zero-copy ownership (Rust) to manual tensor lifecycle (Go).
- Platform support determines deployment: Python for research, Swift for iOS/macOS, Rust for high-performance services, and Go for cloud micro-services.
- Error handling spans exceptions (Python), fatal errors (Swift), Result types (Rust), and explicit error returns (Go).
Frequently Asked Questions
Which Supertonic implementation offers the best performance?
The Rust implementation provides the best raw performance due to zero-allocation ndarray buffers and strict memory safety without garbage collection overhead. Python is limited by the GIL, while Swift and Go offer intermediate performance suited for their respective platform constraints.
Can I use Supertonic on iOS devices?
Yes, but only through the Swift implementation. The Swift version in swift/Sources/Helper.swift targets macOS and iOS exclusively, leveraging Apple's Accelerate framework and native SwiftUI integration. The other languages do not support iOS deployment.
How do the error handling approaches differ between Python and Rust?
Python uses exception raising with raise statements that can be caught with try/except blocks. Rust uses anyhow::Result types that propagate errors through the ? operator, requiring explicit error handling at the call site without runtime exception unwinding.
Are the ONNX model files interchangeable between the different language implementations?
Yes, all four implementations use the same ONNX model files (duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, and vocoder.onnx) and JSON voice style files from the shared assets/ directory. The models are loaded differently in each language, but the underlying binary assets are identical.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →