# How the Speed Parameter in Supertonic Controls Speech Rate

> Control speech rate with Supertonic's speed parameter. Accelerate or decelerate audio playback while preserving pitch. Learn how this floating-point divisor works.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-12

---

**The `speed` parameter in Supertonic is a floating-point divisor that scales predicted speech durations to accelerate or decelerate audio playback while preserving the original voice pitch.**

The `speed` parameter is a core configuration option in the supertone-inc/supertonic text-to-speech engine that directly manipulates the temporal characteristics of generated audio. As a dimensionless factor applied to the duration predictor's output, this parameter allows developers to fine-tune speech tempo across Python, Node.js, Swift, Java, Go, and Web UI implementations. Understanding how the speed parameter affects audio output is essential for achieving natural-sounding speech synthesis at varying playback rates.

## What Is the Speed Parameter in Supertonic?

Supertonic defines `speed` as a floating-point scaling factor that modulates the tempo of synthesized speech. The system employs a neural duration predictor model to estimate the raw temporal length of each phoneme or utterance. Before this duration converts into actual audio samples, the value is divided by the `speed` parameter to produce the final playback timing.

The default value is **1.05**, which generates slightly faster-than-natural speech while maintaining clarity. Most implementations recommend operating within the **0.9 to 1.5** range to avoid artifacts or unnatural cadence.

## How the Speed Parameter Affects Audio Output

### The Mathematical Mechanism

Internally, Supertonic calculates speech duration through a dedicated duration predictor (DP) ONNX model. The `speed` parameter modifies this output through simple division:

```python
dur_onnx, *_ = self.dp_ort.run(
    None, {"text_ids": text_ids, "style_dp": style.dp, "text_mask": text_mask}
)
dur_onnx = dur_onnx / speed  # Speed applied here

```

This operation occurs in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) (lines 189-194) and equivalent locations across all language bindings. Because the waveform undergoes resampling after duration adjustment, the perceived **pitch remains constant** regardless of speed changes.

### Speed Values Above 1.0

When `speed` exceeds **1.0**, the division operation shortens the predicted duration. This compression creates faster speech output—the higher the value, the quicker the playback. For example, setting `speed = 1.2` yields 20% faster speech than the baseline prediction.

### Speed Values Below 1.0

Conversely, values below **1.0** lengthen the duration, producing slower, more deliberate speech. A setting of `speed = 0.9` extends the duration by approximately 11%, resulting in distinctly slower articulation.

## Implementation Across Language Bindings

The speed division logic is consistent across Supertonic's multi-language SDK. Each binding implements the same mathematical operation in its respective syntax:

**Python** ([`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), lines 189-194):

```python
dur_onnx = dur_onnx / speed

```

**Node.js** ([`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js), lines 217-219):

```javascript
durOnnx[i] /= speed;

```

**Swift** ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift), lines 605-607):

```swift
duration[i] /= speed

```

**Java** ([`java/Helper.java`](https://github.com/supertone-inc/supertonic/blob/main/java/Helper.java), lines 316-318):

```java
duration[i] /= speed;

```

**Go** ([`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go), lines 704-706):

```go
durOnnx[i] /= speed

```

**Web UI** ([`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js), lines 191-194):
The frontend transmits the user-specified speed value to the backend, where it enters the same division pipeline as the native implementations.

## Practical Code Examples

### Python CLI Usage

Execute the ONNX example with accelerated speech using the `--speed` flag:

```bash
python -m supertonic.example_onnx \
    --text "Hello, world!" \
    --lang en \
    --speed 1.2

```

Argument parsing for this interface resides in [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) (lines 30-33).

### Node.js Implementation

Pass the speed parameter as the fifth argument to the `call` method:

```javascript
const { TextToSpeech } = require('./nodejs/helper');

const tts = new TextToSpeech(...);
await tts.call(
    "Hello, world!",
    "en",
    style,
    12,    // totalStep
    0.9    // speed (slower speech)
);

```

This pattern appears in [`nodejs/example_onnx.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/example_onnx.js) (lines 64-70).

### Swift Integration

Configure speed through the `Args` struct before batch processing:

```swift
var args = Args()
args.speed = 1.3  // Faster speech

try TextToSpeech().batch(
    args.text,
    args.lang,
    style,
    args.totalStep,
    speed: args.speed
)

```

The argument definitions exist in [`swift/Sources/ExampleONNX.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/ExampleONNX.swift) (lines 38-40).

### Web UI Configuration

The HTML input element constrains values to the recommended range:

```html
<input type="number" id="speed" value="1.05" min="0.9" max="1.5" step="0.05">

```

When modified, this value propagates through [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js) to the synthesis backend.

## Summary

- The **speed parameter** is a floating-point divisor applied to the duration predictor's output in Supertonic.
- Values **greater than 1.0** increase speech rate; values **less than 1.0** decrease it.
- The **default value is 1.05**, with a recommended operating range of **0.9 to 1.5**.
- Pitch preservation occurs because the system resamples the waveform after adjusting temporal duration.
- All language bindings (Python, Node.js, Swift, Java, Go) implement identical division logic in their respective helper files.

## Frequently Asked Questions

### What is the default speed value in Supertonic?

The default speed value is **1.05**, which produces speech slightly faster than the neural model's natural prediction. This default strikes a balance between efficiency and naturalness for most use cases.

### Does adjusting the speed change the pitch of the voice?

No. The speed parameter only affects temporal duration. Because Supertonic resamples the generated waveform after adjusting the duration, the fundamental frequency and perceived pitch remain constant regardless of speed settings.

### What happens if I set the speed parameter outside the recommended range?

While the system accepts any positive floating-point value, operating outside the **0.9 to 1.5** range may produce artifacts, distorted consonants, or unnatural rhythm. Extreme values could potentially cause the synthesis model to fail or generate unintelligible audio.

### Is the speed parameter available in all Supertonic language bindings?

Yes. The speed parameter is fully implemented across the Python, Node.js, Swift, Java, Go, and Web UI bindings. Each implementation performs the same division operation on the duration predictor output, ensuring consistent behavior regardless of the programming language used.