# How the Supertonic Speed Parameter (0.7–2.0) Affects Latency vs. Audio Quality

> Discover how the Supertonic speed parameter (0.7-2.0) impacts audio latency and quality. Learn how to balance natural speech with computational speed for optimal results.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: performance
- Published: 2026-06-15

---

**Increasing the Supertonic `speed` parameter above 1.0 reduces inference latency by compressing phoneme durations into fewer latent frames, while values below 1.0 increase latency for higher fidelity; the recommended 0.9–1.5 range balances natural speech timing with computational speed.**

The `speed` parameter in the [supertone-inc/supertonic](https://github.com/supertone-inc/supertonic) repository controls a fundamental trade-off in this neural TTS pipeline. Understanding how this scalar value modifies the duration predictor’s output reveals why latency scales inversely with speed settings, and why audio quality degrades when pushing beyond the model’s trained timing distribution.

## How the Speed Parameter Works Internally

Supertonic’s architecture uses a duration predictor to estimate how long each phoneme should last in the final audio. The `speed` parameter acts as a post-processing scalar on these predictions across all language implementations.

In [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) (lines 82–87), the scaling occurs immediately after the duration model inference:

```javascript
const duration = Array.from(dpOutputs.duration.data);
for (let i = 0; i < duration.length; i++) {
    duration[i] /= speed;          // Higher speed → shorter duration
}

```

This division operation appears identically in [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift), [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs), [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), [`java/Helper.java`](https://github.com/supertone-inc/supertonic/blob/main/java/Helper.java), and [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go). Because the latent representation size (`sampleNoisyLatent`) depends directly on these duration values, modifying `speed` changes the tensor dimensions fed into the denoising diffusion model.

## Latency Impact: Frame Reduction and Inference Time

The relationship between `speed` and latency is linear and deterministic, stemming from how diffusion TTS models process latent representations.

- **Higher speed (1.5–2.0):** Dividing durations by larger values shrinks the latent tensor. Fewer frames require fewer denoising steps to process, reducing overall inference time by up to 50% or more at the maximum 2.0 setting.
- **Lower speed (0.7–0.9):** Lengthening durations expands the latent representation. More frames trigger additional computational work in the denoising loop, increasing latency by approximately 10–30% compared to the baseline.

The denoising steps (controlled separately via parameters like `total_step`) process each latent frame individually. Therefore, frame count directly dictates the computational workload regardless of the acoustic model’s complexity.

## Audio Quality Trade-offs

While latency improves at higher speeds, the duration predictor was trained on natural speech timing distributions. Rescaling these values beyond the training distribution introduces perceptual artifacts.

**Compression artifacts (>1.5):** When `speed` exceeds 1.5, phonemes become artificially compressed. The vocoder receives latent frames that no longer match the acoustic characteristics expected from the training data, resulting in choppy articulation or robotic-sounding speech.

**Stretching artifacts (<0.7):** Values below 0.9 (and especially below 0.7) stretch timings beyond natural limits, producing muddy, overly drawn-out speech with reduced clarity and definition.

The developers empirically recommend **0.9–1.5** as the operational sweet spot. Within this range, users gain noticeable latency improvements (20–40% faster at 1.5) while the model maintains natural-sounding prosody and articulation.

## Implementation Across Language Bindings

Supertonic maintains consistent `speed` handling across its multi-language SDKs. Each implementation applies the same scalar division to duration tensors:

- **JavaScript/Web:** [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) at line 84
- **Python:** [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) (lines 183–188) using `dur_onnx = dur_onnx / speed`
- **Go:** [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) at line 704 with `durOnnx[i] /= speed`
- **Rust:** [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) applying `*dur /= speed;`
- **Java:** [`java/Helper.java`](https://github.com/supertone-inc/supertonic/blob/main/java/Helper.java) with `duration[i] /= speed;`
- **Swift:** [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) mirroring the same logic

This uniformity ensures that a `speed` value of 1.3 produces identical audio timing and latency characteristics whether calling the Web API, Python bindings, or Go batch processor.

## Practical Configuration Examples

### Web/JavaScript API Usage

Configure speed during the `textToSpeech` call to balance responsiveness and clarity:

```javascript
import { loadTextToSpeech, loadVoiceStyle } from './helper.js';

const { textToSpeech } = await loadTextToSpeech('onnx_models');
const style = await loadVoiceStyle(['style1.json']);

// 1.3x speed for faster synthesis with acceptable quality
const result = await textToSpeech.call(
    "Hello, Supertonic!",
    "en",
    style,
    12,      // denoising steps
    1.3,     // speed factor
    0.3      // silence between chunks
);

```

*The duration reduction is applied internally at line 84 of [`helper.js`](https://github.com/supertone-inc/supertonic/blob/main/helper.js).*

### Python CLI Synthesis

For higher fidelity at the cost of latency, use values below 1.0:

```bash
python example_onnx.py \
    --text "Supertonic demo" \
    --lang en \
    --style style1.json \
    --total_step 12 \
    --speed 0.85

```

*This slower setting increases latency approximately 15% but preserves clearer intonation.*

### Go Batch Processing

For maximum throughput processing multiple utterances:

```go
tts, _ := LoadTextToSpeech("onnx_models")
style, _ := LoadVoiceStyle([]string{"style1.json"})

speed := float32(1.7) // Aggressive speed for lowest latency
wav, durations, _ := tts.Batch(
    []string{"First sentence.", "Second sentence."},
    []string{"en", "en"},
    style,
    12,
    speed,
)

```

*The division occurs in [`helper.go`](https://github.com/supertone-inc/supertonic/blob/main/helper.go) at line 704, reducing the duration tensor before latent generation.*

## Summary

- **Mechanism:** The `speed` parameter divides predicted phoneme durations by the scalar value, directly controlling the size of the latent representation fed to the denoising model.
- **Latency:** Values above 1.0 reduce frame count and inference time up to 50%; values below 1.0 increase latency by 10–30%.
- **Quality:** The 0.9–1.5 range preserves natural speech characteristics; outside this range, compression (>1.5) creates choppy artifacts while stretching (<0.9) produces muddy audio.
- **Consistency:** All Supertonic bindings (JS, Python, Go, Rust, Java, Swift) implement identical duration scaling logic, ensuring cross-platform behavioral consistency.

## Frequently Asked Questions

### What happens if I set the speed parameter to 2.0?

Setting `speed` to 2.0 maximizes inference speed by halving all predicted durations, but typically produces noticeable robotic artifacts and choppy articulation. The developers recommend staying below 1.5 for production applications where audio quality matters.

### Does the speed parameter affect the denoising step count?

No, `speed` only modifies the duration tensor size; the `total_step` or equivalent parameter controls denoising iterations independently. However, fewer duration frames mean fewer total operations per denoising step, which indirectly reduces wall-clock time.

### Why does lower speed increase latency?

Lower `speed` values (0.7–0.9) increase the number of latent frames proportionally. Since the diffusion model processes each frame through its denoising loop, more frames require more computation, extending inference time while potentially improving prosodic naturalness.

### Is the speed parameter applied before or after the acoustic model prediction?

The scaling applies **after** the duration predictor but **before** the vocoder and denoising stages. As implemented in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) and equivalent files, the duration tensor is modified immediately after the duration model inference, ensuring the acoustic model receives properly timed latent representations.