Using Supertonic Across Different Programming Languages: A Cross-Platform TTS Guide
Supertonic exposes a uniform Text-to-Speech API across Python, Node.js, Rust, Java, Go, C++, C#, Swift, and Flutter by wrapping ONNX Runtime bindings around a shared set of inference models, enabling consistent on-device speech synthesis regardless of your technology stack.
Supertonic is a cross-platform TTS system developed by Supertone Inc. that achieves true language independence through an ONNX Runtime architecture. Whether you are building Python server applications, Node.js microservices, or Rust embedded tools, you can integrate the same high-quality voice synthesis pipeline using idiomatic code structures that all map to identical core inference steps.
The Language-Agnostic Architecture
Supertonic’s architecture is deliberately language-agnostic. The core inference pipeline lives in a set of ONNX models—specifically the duration predictor, text encoder, vector estimator, and vocoder—paired with a small runtime shim. This shim handles loading models, preparing input tensors, running the inference loop, and writing resulting audio. Because each language binding merely wraps these ONNX inference steps, the API surface remains uniform across all implementations.
The Three-Step Integration Pattern
Every language binding follows the same implementation pattern demonstrated in the reference implementations:
Step 1: Load Configuration and ONNX Models
The system reads tts.json to initialize model parameters such as sample rate and chunk sizes, then loads the four model files from the assets/onnx directory. In nodejs/helper.js, the loadCfgs function handles this configuration loading (lines 61-64), while Python implementations use py/helper.py to achieve the same result.
Step 2: Prepare Text and Style Tensors
A Unicode processor normalizes raw strings, validates language tags, and converts characters to indices using unicode_indexer.json. In the Node.js implementation, UnicodeProcessor._preprocessText (lines 86-98) handles this transformation. Voice styles stored as JSON blobs (style_ttl and style_dp) are packed into batched tensors by the loadVoiceStyle routine (lines 79-118).
Step 3: Run the Inference Loop
The TextToSpeech class orchestrates a per-step denoising loop that repeatedly calls the vector estimator ONNX model, applying a mask over the latent representation before passing it to the vocoder to obtain the final waveform. In nodejs/helper.js, this logic appears in the _infer method and surrounding step loop (lines 98-165).
Language-Specific Implementations
Supertonic provides reference implementations for ten distinct programming environments:
| Language | Core Shim File | Example Entry Point |
|---|---|---|
| Python | py/helper.py |
py/example_onnx.py |
| Node.js | nodejs/helper.js |
nodejs/example_onnx.js |
| Rust | rust/src/helper.rs |
rust/src/example_onnx.rs |
| Java | java/Helper.java |
java/ExampleONNX.java |
| Go | go/helper.go |
go/example_onnx.go |
| C++ | cpp/helper.cpp |
cpp/example_onnx.cpp |
| C# | csharp/Helper.cs |
csharp/ExampleONNX.cs |
| Swift | swift/Helper.swift |
swift/example_onnx.swift |
| Flutter | flutter/lib/helper.dart |
flutter/lib/main.dart |
| iOS | ios/ExampleiOSApp/TTSService.swift |
— |
All implementations expose the same API primitives: specify the ONNX directory, optionally enable GPU acceleration, provide a list of (text, language, voice-style) tuples, and call either a single-utterance method (call) or a batch method (batch). The runtime automatically chunks long sentences, inserts brief silences between concatenated chunks, and writes 16-bit WAV files at the model’s native sample rate.
Python
The Python API provides a high-level TTS class that encapsulates the three-step pattern. According to the source in py/helper.py and the reference implementation in py/example_onnx.py:
from supertonic import TTS
# Initialise the TTS engine (downloads the model on first run)
tts = TTS(auto_download=True)
# Load a voice style JSON (included in the assets)
style = tts.get_voice_style(voice_name="M1")
# Synthesize a single sentence
wav, duration = tts.synthesize(
"The quick brown fox jumps over the lazy dog.",
voice_style=style,
lang="en"
)
# Save the audio
tts.save_audio(wav, "output.wav")
print(f"Generated {duration:.2f}s of speech")
Node.js
The JavaScript implementation in nodejs/helper.js exposes asynchronous functions that mirror the Python synchronous API. The loadTextToSpeech function initializes the ONNX session, while loadVoiceStyle prepares the style tensors:
import { loadTextToSpeech, loadVoiceStyle } from "./helper.js";
async function main() {
const tts = await loadTextToSpeech("../assets/onnx", false);
const style = loadVoiceStyle(["../assets/voice_styles/M1.json"], true);
const { wav, duration } = await tts.call(
"Hello from Supertonic, running in Node.js!",
"en",
style,
8,
1.05
);
// Helper to write a WAV file (see writeWavFile in helper.js)
writeWavFile("node_output.wav", wav, tts.sampleRate);
console.log(`Audio saved – length ${duration[0].toFixed(2)} s`);
}
main();
Rust
The Rust bindings in rust/src/helper.rs expose the same functionality through strongly-typed functions, as demonstrated in rust/src/example_onnx.rs:
use supertonic::helper::{load_text_to_speech, load_voice_style, timer, write_wav_file, sanitize_filename};
fn main() -> anyhow::Result<()> {
// Load the inference pipeline
let mut tts = load_text_to_speech("../assets/onnx", false)?;
// Load a voice style
let style = load_voice_style(&["../assets/voice_styles/M1.json"], true)?;
// Synthesize
let (wav, duration) = timer("Generating speech", || {
tts.call(
"Rust can also use Supertonic!",
"en",
&style,
8,
1.05,
0.3,
)
})?;
// Save the result
let fname = format!("{}.wav", sanitize_filename("rust_example", 20));
write_wav_file(&std::path::PathBuf::from("rust_output.wav"), &wav, tts.sample_rate)?;
println!("Saved {} ({} s)", fname, duration[0]);
Ok(())
}
Summary
- Supertonic separates the TTS inference pipeline from language-specific implementations by building on ONNX Runtime.
- Every binding—from
py/helper.pytorust/src/helper.rstonodejs/helper.js—follows the same three-step pattern: load models, preprocess text and styles, then run the inference loop. - The API surface is intentionally uniform: initialize with an ONNX directory path, optionally enable GPU, and call
call()for single utterances orbatch()for multiple inputs. - All implementations automatically handle sentence chunking, cross-chunk silence insertion, and 16-bit WAV output at the native sample rate defined in
assets/onnx/tts.json.
Frequently Asked Questions
Which programming languages does Supertonic officially support?
Supertonic maintains official bindings for Python, Node.js, Rust, Java, Go, C++, C#, Swift, Flutter, and iOS. Each implementation resides in its own directory (e.g., py/, rust/, nodejs/) and provides equivalent functionality through idiomatic language patterns while sharing the same underlying ONNX models in assets/onnx/.
Do I need separate voice model files for each programming language?
No. All language bindings consume the same ONNX model files stored in assets/onnx/ and the same voice style JSONs in assets/voice_styles/. You can train or download models once and use them interchangeably across Python, Rust, Node.js, or any other supported language without conversion or replication.
How do I enable GPU acceleration in Supertonic?
Each language binding accepts a boolean GPU flag during initialization. For example, in Node.js you pass true as the second argument to loadTextToSpeech("../assets/onnx", true), while in Python and Rust the equivalent parameters enable ONNX Runtime’s CUDA or DirectML execution providers. The underlying inference logic in _infer (as implemented in nodejs/helper.js lines 98-165) remains identical regardless of the execution provider.
Can Supertonic handle long text inputs across all programming languages?
Yes. The TextToSpeech class automatically segments long sentences into chunks based on the configuration in tts.json, processes each chunk through the denoising loop, and concatenates the results with brief silences between them. This behavior is implemented identically in py/helper.py, rust/src/helper.rs, and nodejs/helper.js, ensuring consistent audio output length and quality regardless of your chosen language.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →