Supertonic Text-to-Speech Output Audio Format: 44.1 kHz 16-bit WAV Configuration

Supertonic generates 16-bit PCM WAV files at a configurable sample rate (default 44.1 kHz) in mono channel format.

Supertonic is an open-source text-to-speech (TTS) engine developed by Supertone Inc. that produces standard audio files for maximum compatibility. The Supertonic output audio format follows the PCM WAV specification with 16-bit depth and a sample rate defined in the model configuration, defaulting to CD-quality 44.1 kHz.

Output Format Specifications

Supertonic outputs raw audio as PCM WAV files with specific characteristics hardcoded in the encoder and configurable via model settings.

Bit Depth and Channels

The encoder explicitly writes 16-bit integer samples in mono (single channel) format. In go/helper.go, the writeWavFile function initializes a wav.NewEncoder with bit depth set to 16 and channels set to 1, ensuring consistent output across all generated files.

Sample Rate Configuration

While the bit depth remains fixed at 16-bit, the sample rate is configurable through the model configuration file. The LoadCfgs function in go/helper.go parses tts.json and populates cfg.AE.SampleRate, which defaults to 44,100 Hz (44.1 kHz) in the reference models shipped with the repository.

Audio Generation Pipeline

The TTS pipeline follows a three-stage process from text input to WAV file output, as implemented in the Go core of the repository.

Configuration Loading

First, LoadCfgs reads the ONNX model directory and loads tts.json, extracting the sample rate into cfg.AE.SampleRate. This value determines the temporal resolution of the output audio.

Inference and Sample Generation

The TextToSpeech.Call method runs the neural inference pipeline, producing a slice of float32 audio samples (wav). These raw floating-point values represent the synthesized speech waveform before quantization.

WAV Encoding

Finally, writeWavFile creates the final output file. According to the source in go/helper.go (lines 78-100), the function creates a wav.NewEncoder with the configured sample rate, 16-bit depth, and 1 channel, then converts the float32 samples to 16-bit integers before writing to disk.

Practical Implementation Example

The following Go code demonstrates the end-to-end process of generating a 44.1 kHz 16-bit WAV file:

package main

import (
    "log"
    "path/filepath"
)

func main() {
    // Load configuration (includes cfg.AE.SampleRate)
    cfg, err := LoadCfgs("path/to/onnx")
    if err != nil {
        log.Fatalf("config error: %v", err)
    }

    // Initialize TTS engine
    tts, err := LoadTextToSpeech("path/to/onnx", false, cfg)
    if err != nil {
        log.Fatalf("load error: %v", err)
    }
    defer tts.Destroy()

    // Generate float32 audio samples
    wav, _, err := tts.Call("Hello world!", "en", nil, 50, 1.0, 0.0)
    if err != nil {
        log.Fatalf("synthesis error: %v", err)
    }

    // Write 16-bit PCM WAV (44.1 kHz by default)
    outPath := filepath.Join("out", "hello.wav")
    if err := writeWavFile(outPath, wav, cfg.AE.SampleRate); err != nil {
        log.Fatalf("write error: %v", err)
    }
}

Running this code creates hello.wav, a standard 44.1 kHz 16-bit mono WAV file compatible with any audio player.

Modifying Output Parameters

To change the output sample rate from the default 44.1 kHz, modify the SampleRate field in your model's tts.json configuration file before loading it with LoadCfgs. The writeWavFile function automatically respects this configuration when creating the encoder, though the bit depth remains locked at 16-bit for PCM compliance.

Summary

  • Supertonic outputs 16-bit PCM WAV files in mono format.
  • The default sample rate is 44.1 kHz, configurable via cfg.AE.SampleRate in tts.json.
  • The writeWavFile function in go/helper.go handles encoding using wav.NewEncoder with the specified sample rate and fixed 16-bit depth.
  • Raw audio passes through as float32 during inference before conversion to 16-bit integers at write time.

Frequently Asked Questions

What audio format does Supertonic use for TTS output?

Supertonic generates PCM WAV files with 16-bit sample depth and mono channel configuration. The format follows standard WAV specifications ensuring compatibility with standard media players and audio processing tools.

Is the sample rate configurable in Supertonic?

Yes. The sample rate is determined by the AE.SampleRate field in the model configuration file (tts.json). While the default is 44.1 kHz, you can modify this value in the configuration before loading the model with LoadCfgs in go/helper.go.

What bit depth does Supertonic use for WAV files?

Supertonic exclusively uses 16-bit depth for its WAV output. The writeWavFile function hardcodes this value when creating the wav.NewEncoder, converting internal float32 samples to 16-bit integers during the write process.

Can Supertonic output stereo audio instead of mono?

No. The current implementation in go/helper.go explicitly sets the channel count to 1 (mono) in the wav.NewEncoder call. To generate stereo output, you would need to modify the writeWavFile function or post-process the mono WAV file into dual-channel audio.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →