Is Supertonic Compatible with Other Audio Libraries? A Complete Integration Guide

Supertonic generates standard PCM-16 bit WAV files across all its language bindings, ensuring seamless interoperability with any audio-processing library that supports canonical RIFF-formatted WAV files.

Supertonic, maintained by the team at supertone-inc/supertonic, is an open-source text-to-speech engine designed with interoperability as a core principle. If you are evaluating whether Supertonic is compatible with other audio libraries, the answer centers on its architectural decision to output raw floating-point audio samples that every language binding converts to the universally supported PCM-16 WAV format. This standardization eliminates the need for proprietary codecs or conversion utilities.

How Supertonic Standardizes Audio Output

The core TTS engine in Supertonic outputs raw floating-point audio samples ([Float] or &[f32]). However, before any audio reaches your application code, each language binding converts this data to a standard PCM-16 bit WAV file with canonical RIFF formatting. This design choice ensures that downstream audio processing tools receive a universally readable format.

JavaScript WAV Generation

In the JavaScript binding, web/helper.js implements the writeWavFile function that constructs a proper RIFF header and writes 16-bit PCM samples. This implementation ensures browser-based applications can generate files compatible with the Web Audio API and Node.js audio processing modules.

Rust WAV Export

The Rust implementation in rust/src/helper.rs provides the write_wav_file function, which leverages the hound crate to emit properly formatted PCM-16 WAV files. This creates standard audio files that any Rust audio ecosystem tool can ingest without additional parsing logic.

Swift Audio Serialization

For iOS and macOS developers, swift/Sources/Helper.swift contains the writeWavFile method that manually constructs the WAV header and writes Int16 samples. This approach guarantees compatibility with Apple's AVAudioEngine and other platform-specific audio frameworks.

Python Audio Pipeline

The Python examples in py/example_onnx.py demonstrate integration with SoundFile (sf.write), while the project's py/pyproject.toml lists optional dependencies including Librosa (>=0.10) and SoundFile (>=0.12.1). These choices reflect the Python audio community's standard tools for analysis and file I/O.

Because Supertonic outputs regular WAV files with canonical formatting, you can feed generated audio into pydub, torchaudio, FFmpeg, Audacity, or the Web Audio API without conversion steps. The only requirement is support for 16-bit PCM WAV, the de-facto standard for audio software.

Loading with Librosa in Python

Librosa is the standard library for music and audio analysis in Python. Since Supertonic generates standard WAV files, loading them requires no special configuration:

import librosa

# Path to the WAV file produced by Supertonic

wav_path = "results/example_onnx.wav"

# Librosa loads the file as float32 samples (range -1…1)

audio, sr = librosa.load(wav_path, sr=None, mono=True)

print(f"Loaded {len(audio)} samples at {sr} Hz")

Converting Formats with Pydub

For format conversion or manipulation, Supertonic's output works seamlessly with Pydub, which wraps FFmpeg:

from pydub import AudioSegment

wav = AudioSegment.from_wav("results/example_onnx.wav")

# Export to MP3 (or any format pydub/ffmpeg supports)

wav.export("example.mp3", format="mp3")

Reading WAV Files in Rust

In Rust applications, you can read Supertonic-generated files using the same hound crate that Supertonic uses for writing:

use hound::WavReader;

let mut reader = WavReader::open("results/example.wav")?;
let spec = reader.spec();
let samples: Vec<i16> = reader.samples::<i16>().map(|s| s.unwrap()).collect();

println!("{} samples, {} Hz, {} channels", samples.len(), spec.sample_rate, spec.channels);

Playing Audio in Swift

iOS and macOS applications can play Supertonic-generated audio using native AVAudioPlayer without format conversion:

import AVFoundation

let url = URL(fileURLWithPath: "results/example.wav")
let player = try AVAudioPlayer(contentsOf: url)
player.play()

Web Audio API Integration

In browser environments, you can load Supertonic-generated audio buffers directly into the Web Audio API:

// Assume `wavBuffer` is a Float32Array from Supertonic
const audioCtx = new AudioContext();
const buffer = audioCtx.createBuffer(1, wavBuffer.length, sampleRate);
buffer.copyToChannel(wavBuffer, 0);
const source = audioCtx.createBufferSource();
source.buffer = buffer;
source.connect(audioCtx.destination);
source.start();

Summary

  • Supertonic outputs standard PCM-16 WAV files across all language bindings (Python, Rust, Swift, JavaScript), ensuring universal compatibility.
  • No proprietary formats are used; the core engine converts raw &[f32] samples to canonical RIFF-formatted WAV files before saving.
  • Verified integrations include Librosa, SoundFile, Pydub, FFmpeg, hound, AVAudioPlayer, and the Web Audio API.
  • Implementation files handling the conversion include web/helper.js, rust/src/helper.rs, swift/Sources/Helper.swift, and py/example_onnx.py.
  • Optional dependencies in the Python package specifically target standard audio libraries like Librosa (>=0.10) and SoundFile (>=0.12.1).

Frequently Asked Questions

Can Supertonic generate audio formats other than WAV?

Supertonic's language bindings currently export to PCM-16 bit WAV as the default format. While the core engine produces raw floating-point samples that could theoretically support any format, the official implementations in rust/src/helper.rs, web/helper.js, and swift/Sources/Helper.swift standardize on WAV for maximum compatibility. To create MP3, FLAC, or AAC files, use a secondary library like Pydub or FFmpeg to convert the generated WAV files.

Do I need to install specific codecs to use Supertonic with PyTorch or TensorFlow audio tools?

No additional codecs are required. Since Supertonic outputs standard PCM-16 WAV files, libraries like torchaudio can load them directly using standard I/O methods. The WAV format is universally supported across deep learning frameworks because it requires no decompression algorithms and provides lossless audio data suitable for model training pipelines.

Is Supertonic compatible with real-time audio streaming?

Yes, though with considerations. The core engine generates raw floating-point samples ([Float] or &[f32]) that can be streamed as buffers before WAV file finalization. For real-time applications, you can intercept the raw float buffers from the language bindings and pipe them directly to audio output APIs (like the Web Audio API or Core Audio) without waiting for file write operations to complete.

Which audio libraries does Supertonic explicitly depend on?

According to the source code in py/pyproject.toml and py/example_onnx.py, Supertonic lists Librosa (>=0.10) and SoundFile (>=0.12.1) as optional Python dependencies. The Rust implementation depends on the hound crate for WAV writing. These are standard, well-maintained libraries in the audio processing ecosystem, ensuring long-term compatibility and community support.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →