How the Supertonic Speed Parameter (0.7–2.0) Affects Latency vs. Audio Quality

Increasing the Supertonic speed parameter above 1.0 reduces inference latency by compressing phoneme durations into fewer latent frames, while values below 1.0 increase latency for higher fidelity; the recommended 0.9–1.5 range balances natural speech timing with computational speed.

The speed parameter in the supertone-inc/supertonic repository controls a fundamental trade-off in this neural TTS pipeline. Understanding how this scalar value modifies the duration predictor’s output reveals why latency scales inversely with speed settings, and why audio quality degrades when pushing beyond the model’s trained timing distribution.

How the Speed Parameter Works Internally

Supertonic’s architecture uses a duration predictor to estimate how long each phoneme should last in the final audio. The speed parameter acts as a post-processing scalar on these predictions across all language implementations.

In web/helper.js (lines 82–87), the scaling occurs immediately after the duration model inference:

const duration = Array.from(dpOutputs.duration.data);
for (let i = 0; i < duration.length; i++) {
    duration[i] /= speed;          // Higher speed → shorter duration
}

This division operation appears identically in swift/Sources/Helper.swift, rust/src/helper.rs, py/helper.py, java/Helper.java, and go/helper.go. Because the latent representation size (sampleNoisyLatent) depends directly on these duration values, modifying speed changes the tensor dimensions fed into the denoising diffusion model.

Latency Impact: Frame Reduction and Inference Time

The relationship between speed and latency is linear and deterministic, stemming from how diffusion TTS models process latent representations.

  • Higher speed (1.5–2.0): Dividing durations by larger values shrinks the latent tensor. Fewer frames require fewer denoising steps to process, reducing overall inference time by up to 50% or more at the maximum 2.0 setting.
  • Lower speed (0.7–0.9): Lengthening durations expands the latent representation. More frames trigger additional computational work in the denoising loop, increasing latency by approximately 10–30% compared to the baseline.

The denoising steps (controlled separately via parameters like total_step) process each latent frame individually. Therefore, frame count directly dictates the computational workload regardless of the acoustic model’s complexity.

Audio Quality Trade-offs

While latency improves at higher speeds, the duration predictor was trained on natural speech timing distributions. Rescaling these values beyond the training distribution introduces perceptual artifacts.

Compression artifacts (>1.5): When speed exceeds 1.5, phonemes become artificially compressed. The vocoder receives latent frames that no longer match the acoustic characteristics expected from the training data, resulting in choppy articulation or robotic-sounding speech.

Stretching artifacts (<0.7): Values below 0.9 (and especially below 0.7) stretch timings beyond natural limits, producing muddy, overly drawn-out speech with reduced clarity and definition.

The developers empirically recommend 0.9–1.5 as the operational sweet spot. Within this range, users gain noticeable latency improvements (20–40% faster at 1.5) while the model maintains natural-sounding prosody and articulation.

Implementation Across Language Bindings

Supertonic maintains consistent speed handling across its multi-language SDKs. Each implementation applies the same scalar division to duration tensors:

This uniformity ensures that a speed value of 1.3 produces identical audio timing and latency characteristics whether calling the Web API, Python bindings, or Go batch processor.

Practical Configuration Examples

Web/JavaScript API Usage

Configure speed during the textToSpeech call to balance responsiveness and clarity:

import { loadTextToSpeech, loadVoiceStyle } from './helper.js';

const { textToSpeech } = await loadTextToSpeech('onnx_models');
const style = await loadVoiceStyle(['style1.json']);

// 1.3x speed for faster synthesis with acceptable quality
const result = await textToSpeech.call(
    "Hello, Supertonic!",
    "en",
    style,
    12,      // denoising steps
    1.3,     // speed factor
    0.3      // silence between chunks
);

The duration reduction is applied internally at line 84 of helper.js.

Python CLI Synthesis

For higher fidelity at the cost of latency, use values below 1.0:

python example_onnx.py \
    --text "Supertonic demo" \
    --lang en \
    --style style1.json \
    --total_step 12 \
    --speed 0.85

This slower setting increases latency approximately 15% but preserves clearer intonation.

Go Batch Processing

For maximum throughput processing multiple utterances:

tts, _ := LoadTextToSpeech("onnx_models")
style, _ := LoadVoiceStyle([]string{"style1.json"})

speed := float32(1.7) // Aggressive speed for lowest latency
wav, durations, _ := tts.Batch(
    []string{"First sentence.", "Second sentence."},
    []string{"en", "en"},
    style,
    12,
    speed,
)

The division occurs in helper.go at line 704, reducing the duration tensor before latent generation.

Summary

  • Mechanism: The speed parameter divides predicted phoneme durations by the scalar value, directly controlling the size of the latent representation fed to the denoising model.
  • Latency: Values above 1.0 reduce frame count and inference time up to 50%; values below 1.0 increase latency by 10–30%.
  • Quality: The 0.9–1.5 range preserves natural speech characteristics; outside this range, compression (>1.5) creates choppy artifacts while stretching (<0.9) produces muddy audio.
  • Consistency: All Supertonic bindings (JS, Python, Go, Rust, Java, Swift) implement identical duration scaling logic, ensuring cross-platform behavioral consistency.

Frequently Asked Questions

What happens if I set the speed parameter to 2.0?

Setting speed to 2.0 maximizes inference speed by halving all predicted durations, but typically produces noticeable robotic artifacts and choppy articulation. The developers recommend staying below 1.5 for production applications where audio quality matters.

Does the speed parameter affect the denoising step count?

No, speed only modifies the duration tensor size; the total_step or equivalent parameter controls denoising iterations independently. However, fewer duration frames mean fewer total operations per denoising step, which indirectly reduces wall-clock time.

Why does lower speed increase latency?

Lower speed values (0.7–0.9) increase the number of latent frames proportionally. Since the diffusion model processes each frame through its denoising loop, more frames require more computation, extending inference time while potentially improving prosodic naturalness.

Is the speed parameter applied before or after the acoustic model prediction?

The scaling applies after the duration predictor but before the vocoder and denoising stages. As implemented in web/helper.js and equivalent files, the duration tensor is modified immediately after the duration model inference, ensuring the acoustic model receives properly timed latent representations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →