# How the Supertonic Speed Parameter (0.7–2.0) Affects Synthesis Pitch and Pace

> Discover how the Supertonic speed parameter (0.7-2.0) controls synthesis tempo without affecting pitch. Learn to adjust speech pace with this powerful tool. Explore the Supertonic repository for more.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: performance
- Published: 2026-06-14

---

**The speed parameter rescales the predicted duration of speech segments to control tempo without altering pitch or voice characteristics.**

Supertonic, developed by **supertone-inc/supertonic**, generates speech through a cascade of ONNX models that process text into audio. The **speed parameter** (accepting values between 0.7 and 2.0) operates exclusively on the timing layer of this pipeline, stretching or compressing the duration of utterances while preserving the acoustic properties generated by the vocoder.

## How the Speed Parameter Works in Supertonic

Supertonic's inference pipeline runs through a sequence of neural models: duration predictor → text encoder → latent-noise estimator → vocoder. The speed factor intervenes at the very beginning of this chain by modifying the output of the duration predictor.

### Duration Rescaling in the Inference Pipeline

When the pipeline processes input text, the duration predictor generates a raw duration tensor (`dur_onnx`) representing the time allocated to each phoneme or speech segment. Before this tensor proceeds to downstream models, Supertonic applies the speed parameter as a simple division operation:

```python

# py/helper.py – duration rescaling

dur_onnx = dur_onnx / speed    # https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L193

```

This operation appears consistently across all language bindings. By dividing the duration values by the speed factor, the system effectively compresses the timeline when `speed > 1.0` or expands it when `speed < 1.0`. The downstream vocoder then synthesizes the same phonetic content within this modified timeframe, resulting in faster or slower speech delivery.

## Why Pitch Remains Constant Despite Speed Changes

The **pitch** and **timbre** of synthesized speech remain unchanged because these acoustic characteristics are generated by the vocoder from latent representations that are independent of the duration scaling.

The vocoder receives latent features derived from the text encoder and latent-noise estimator, which encode phonetic identity and prosody but not timing. Since the speed parameter only modifies the duration tensor before it influences the audio generation phase, the fundamental frequency (F0) and spectral envelope produced by the vocoder stay consistent with the original voice style. Only the **tempo**—the rate at which words are delivered—changes perceptibly.

## Recommended Speed Range and Quality Considerations

While Supertonic accepts values between 0.7 and 2.0, the recommended operational range is **0.9 to 1.5** for optimal intelligibility.

- **Values below 0.9**: Speech becomes unnaturally drawn out and may exhibit artifacts from excessive time-stretching.
- **Values above 1.5**: Speech sounds rushed and can suffer from reduced clarity and intelligibility.
- **Default value (1.05)**: Provides a subtle speed increase that most listeners perceive as natural and efficient.

## Cross-Platform Implementation Examples

The duration rescaling logic is implemented identically across Supertonic's language bindings. Here are practical examples for common platforms.

### Python CLI and API

In the Python implementation, the speed parameter is exposed both via command-line interface and programmatic API:

```bash

# Faster speech (speed = 1.2)

uv run example_onnx.py \
  --voice-style ../assets/voice_styles/F2.json \
  --text "Quickly synthesized speech." \
  --speed 1.2

# Slower speech (speed = 0.9)

uv run example_onnx.py \
  --voice-style ../assets/voice_styles/M2.json \
  --text "Deliberately spoken text." \
  --speed 0.9

```

```python
from supertonic import load_text_to_speech, load_voice_style, Style

tts = load_text_to_speech(onnx_dir="../assets/onnx")
style = load_voice_style(["../assets/voice_styles/M1.json"])[0]

# speed = 1.3 → 30% faster

wav, dur = tts("Hello, world!", "en", style, total_step=8, speed=1.3)

```

### Node.js Integration

The Node.js binding applies the same division operation within its `_infer` function:

```js
const { TextToSpeech, loadVoiceStyle } = require("./helper");
const tts = await TextToSpeech.load("../assets/onnx");
const [style] = await loadVoiceStyle(["../assets/voice_styles/M1.json"]);

const { wav, duration } = await tts.call(
  "Hello from Node!", "en", style, 8, 0.85   // slower speech
);

```

### Swift Implementation

In Swift, the speed factor is passed as a named parameter to the inference call:

```swift
let tts = try TextToSpeech(onnxDir: "../assets/onnx")
let style = try tts.loadVoiceStyle(paths: ["../assets/voice_styles/F1.json"])
let (wav, duration) = try tts.call(
    "Swift speed demo", "en", style, 8, speed: 1.4, silenceDuration: 0.3
)   // faster speech

```

## Source Code References

The consistent behavior across platforms is ensured by identical implementations in the following files:

- **Python** ([`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)): Core inference; `dur_onnx = dur_onnx / speed`
- **Python** ([`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py)): CLI example showing `--speed` flag
- **JavaScript (Web)** ([`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js)): UI control for speed (slider)
- **Swift** ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)): Speed factor applied to duration
- **Rust** ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)): `*dur /= speed;`
- **Node.js** ([`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js)): Speed handling in `_infer`
- **Java** ([`java/Helper.java`](https://github.com/supertone-inc/supertonic/blob/main/java/Helper.java)): Duration scaling (`duration[i] /= speed;`)
- **Go** ([`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go)): `durOnnx[i] /= speed`

Each implementation divides the duration tensor by the speed parameter before passing it to the vocoder, ensuring that the **tempo adjustment** is uniform across all deployment targets.

## Summary

- The **speed parameter** (0.7–2.0) controls only the tempo of synthesized speech, not pitch or timbre.
- Supertonic rescales duration by dividing the predicted duration tensor (`dur_onnx`) by the speed factor in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py).
- **Pitch preservation** occurs because the vocoder generates acoustic features from latent representations independent of the duration scaling.
- The **recommended range** is 0.9–1.5, with a default of 1.05 for natural-sounding results.
- The implementation is consistent across Python, Rust, Swift, Node.js, Java, and Go bindings.

## Frequently Asked Questions

### Does the speed parameter change the pitch of the voice?

No, the speed parameter does not affect pitch. It only rescales the duration tensor (`dur_onnx`) before synthesis, leaving the fundamental frequency and spectral characteristics generated by the vocoder unchanged. The resulting audio simply plays faster or slower while maintaining the original voice's pitch profile.

### Why is the default speed set to 1.05 instead of 1.0?

The default value of 1.05 provides a slight acceleration that most listeners perceive as more natural and efficient than exactly real-time speech, without introducing the artifacts that occur at higher speeds. This subtle increase improves delivery cadence while staying well within the recommended range of 0.9–1.5.

### What happens if I set the speed below 0.7 or above 2.0?

While the system accepts values between 0.7 and 2.0, values outside this range are not recommended. Below 0.7, speech becomes extremely slow and unnatural. Above 2.0, the compression is so severe that intelligibility degrades significantly. Staying within the 0.9–1.5 range ensures optimal audio quality.

### Is the speed parameter applied differently across programming languages?

No, the implementation is identical across all language bindings. Whether using Python ([`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)), Rust ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)), Swift ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)), or other supported languages, the operation remains `duration = duration / speed`, ensuring consistent behavior regardless of the deployment platform.