# Supertonic Audio Formats: Why It Only Outputs WAV (44.1 kHz 16-bit)

> Discover why Supertonic audio formats are limited to WAV 44.1 kHz 16-bit for studio-quality output. Learn about its SDK and platform compatibility and lack of MP3 support.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: api-reference
- Published: 2026-06-13

---

**Supertonic exclusively outputs studio-grade WAV files at 44.1 kHz 16-bit PCM across all language SDKs and platforms, with no native support for compressed formats like MP3 or OGG.**

The supertone-inc/supertonic text-to-speech (TTS) engine generates high-fidelity audio through a deterministic ONNX Runtime pipeline. According to the source code, all synthesis operations produce raw PCM audio wrapped in the **WAV** container format, ensuring consistent **Supertonic audio formats** across Python, JavaScript, Rust, Go, and other supported languages.

## The Core Audio Format Specification

Supertonic's architecture mandates a single audio output specification. The core library generates a float-array waveform (`wav`) that is subsequently encoded as a **16-bit PCM** stream at **44.1 kHz** sample rate within a standard WAV container.

This design choice appears in the repository's documentation at [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md) line 27, which explicitly states the engine "Outputs studio-grade **44.1kHz 16-bit WAV** directly." The TTS engine runs entirely on-device via ONNX Runtime, returning a float tensor (`wav_tts`) that helper utilities convert to the standardized WAV format.

### Why WAV Only?

The developers chose WAV for its lossless, self-describing properties. Unlike compressed formats such as MP3 or OGG, WAV provides **deterministic latency** and **universal compatibility** with browsers, media players, and native audio APIs. This ensures fidelity-critical applications receive uncompressed audio data without codec overhead or quality loss.

## Implementation Across Language SDKs

Every Supertonic SDK wraps the core WAV output consistently, varying only in how the binary data is delivered to the application.

### Python SDK ([`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py))

In the Python implementation, the `save_audio` function handles WAV serialization. Located at [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) lines 82-86, this utility extracts the synthesis tensor and writes standard WAV files to disk.

```python
from supertonic import TTS

tts = TTS(auto_download=True)
style = tts.get_voice_style(voice_name="M1")
wav, duration = tts.synthesize(
    text="Hello, Supertonic!",
    lang="en",
    voice_style=style,
    total_steps=8,
)
tts.save_audio(wav, "hello.wav")   # → 44.1 kHz, 16-bit PCM WAV

```

### JavaScript and Web ([`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) and [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js))

The browser implementation creates WAV Blobs for download or playback. The [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) file (lines 266-275) extracts the `wav_tts` tensor and prepares the buffer, while [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js) (lines 197-218) wraps this data with the MIME type `audio/wav`.

```javascript
const { wav, duration } = await textToSpeech.call("Hello, Supertonic!", "en", style, 8);
const wavBuffer = writeWavFile(wav, textToSpeech.sampleRate);
const blob = new Blob([wavBuffer], { type: "audio/wav" });
downloadAudio(URL.createObjectURL(blob), "hello.wav");

```

### Rust ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs))

The Rust SDK provides the `write_wav_file` function at [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) lines 294-301. This implementation follows the same 44.1 kHz 16-bit specification, ensuring binary compatibility with outputs from other languages.

### Go ([`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go))

Go developers utilize the `WriteWavFile` function in [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go), which leverages the `github.com/go-audio/wav` package to maintain format consistency across the ecosystem.

```go
wav, dur, err := tts.Synthesize("Hello, Supertonic!", "en", style, 8)
if err != nil { log.Fatal(err) }
if err := helper.WriteWavFile("hello.wav", wav, tts.SampleRate); err != nil {
    log.Fatal(err)
}

```

## Working with Supertonic Audio Output

Since Supertonic emits only uncompressed WAV data, applications requiring MP3, OGG, or AAC must perform conversion externally. The `sample_rate` constant (44100) and 16-bit depth remain consistent across all platforms, simplifying downstream processing pipelines.

## Summary

- Supertonic exclusively outputs **WAV format** (44.1 kHz, 16-bit PCM) across all SDKs
- The core engine generates float tensors (`wav_tts`) that helper utilities convert to WAV via functions like `save_audio` and `write_wav_file`
- **No native support** exists for MP3, OGG, or other compressed containers
- All language implementations (Python, JavaScript, Rust, Go) maintain identical audio specifications
- Source references include [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js), [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs), and [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go)

## Frequently Asked Questions

### Does Supertonic support MP3 or OGG output?

No. According to the source code analysis, Supertonic does not generate MP3, OGG, or any compressed audio formats internally. The synthesis pipeline in `supertone-inc/supertonic` produces only uncompressed WAV files. If your application requires compressed audio, you must convert the WAV output using external libraries like FFmpeg or pydub.

### What is the sample rate and bit depth of Supertonic output?

All Supertonic outputs use **44.1 kHz sample rate** and **16-bit PCM** depth. This specification is hardcoded across the Python ([`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)), JavaScript ([`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js)), Rust ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)), and Go ([`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go)) implementations, ensuring consistent audio quality regardless of programming language.

### Can I convert Supertonic's WAV output to other formats?

Yes. While Supertonic only emits WAV files, you can convert them to MP3, AAC, or OGG using standard audio processing tools. The 44.1 kHz 16-bit WAV format provides a high-quality source for any conversion process, though you must implement this conversion in your application code after receiving the output from `tts.save_audio()` or equivalent SDK methods.

### Is the audio format the same across all Supertonic SDKs?

Yes. Whether using Python, JavaScript (browser/Node), Go, Rust, Swift, C++, or C#, all Supertonic SDKs wrap the same core synthesis engine and output identical WAV formats. The `sample_rate` remains 44100 across all platforms, and helper functions like `writeWavFile` (JavaScript) and `write_wav_file` (Rust) implement the same PCM encoding logic.