# How Supertonic Handles Different Audio Lengths in Text-to-Speech Synthesis

> Discover how Supertonic handles diverse audio lengths using advanced techniques like length-to-mask conversion and automatic chunking for seamless text-to-speech synthesis. Explore our SDKs.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-13

---

**Supertonic handles variable audio lengths through length-to-mask conversion, latent-mask creation based on predicted durations, and automatic chunking with concatenation for long texts, implemented consistently across Python, Node.js, and Swift SDKs.**

The `supertone-inc/supertonic` repository provides a flexible text-to-speech pipeline that automatically adapts to inputs of any duration, from sub-second utterances to multi-minute passages. Unlike fixed-length TTS systems, Supertonic dynamically adjusts its internal representations using duration prediction and binary masking mechanisms. This article examines the three core architectural patterns that enable Supertonic to handle different audio lengths across all language bindings.

## Three Mechanisms for Variable-Length Audio

Supertonic’s ability to process arbitrary durations rests on three coordinated mechanisms that appear in every language binding: Python, Node.js, Swift, and the Web demo.

### Length-to-Mask Conversion

The foundation of Supertonic’s flexibility is the conversion of sequence lengths into binary masks that tell the diffusion model which time-steps are valid data versus padding.

In **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** (lines 57‑71), the `length_to_mask()` function takes a list of sequence lengths and produces a boolean tensor where `True` values indicate valid positions. This mechanism is replicated across platforms:

- **Node.js**: `lengthToMask()` in **[`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js)** (lines 306‑311)
- **Swift**: `lengthToMask(_:maxLen:)` in **[`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)** (lines 97‑108)

These masks ensure that regardless of input length, the model only processes valid audio frames while maintaining consistent tensor shapes for batch processing.

### Latent-Mask Creation

After the duration predictor yields a target speech length in seconds, Supertonic converts this duration into a *latent* count—the internal representation used by the diffusion model. This latent mask is applied to randomly-generated noise tensors so only the required time-steps undergo refinement.

In **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** (lines 74‑80), the `get_latent_mask()` function orchestrates this conversion: predicted seconds become wav-sample counts, then latent steps, then binary masks. The cross-platform implementations follow the same logic:

- **Node.js**: `getLatentMask()` in **[`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js)** (lines 166‑172), which calls `lengthToMask()`
- **Swift**: `getLatentMask(_:_)` in **[`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)** (lines 66‑78)

This mask is element-wise multiplied against the Gaussian noise tensor, zero-padding any extra steps while preserving the random latent initialization for valid positions.

### Chunking and Concatenation

For texts exceeding the model’s context window, Supertonic automatically splits input into manageable chunks—approximately 300 tokens for most languages, or 120 tokens for Korean and Japanese. Each chunk is synthesized independently, then stitched together with a 0.3-second silence gap to produce seamless final waveforms.

The chunking logic resides in **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** (lines 28‑44), where `chunk_text()` splits the input and the `TextToSpeech.__call__` method handles concatenation. Equivalent implementations exist in:

- **Node.js**: `_infer` and `batch` methods in **[`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js)** (lines 30‑44), which create silence buffers between chunks
- **Swift**: `chunkText(_:maxLen:)` in **[`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)** (lines 34‑42), followed by post-processing that appends silence

## The Synthesis Pipeline Step-by-Step

When a user invokes the top-level `TextToSpeech` object (Python) or calls `textToSpeech` (Node.js/Swift), the library executes an eight-step pipeline that handles different audio lengths automatically:

1. **Text preprocessing**: Unicode normalization, emoji removal, and language-specific tokenization occur in `UnicodeProcessor._preprocessText` across all platforms.

2. **Duration prediction**: The duration predictor ONNX model (`dp_onnx`) returns a float array of seconds for each input chunk via `TextToSpeech._infer` → `dp_ort.run` (Python).

3. **Latent mask generation**: `get_latent_mask()` converts predicted seconds to wav samples, then to latent steps, then to a binary mask.

4. **Random latent creation**: A Gaussian noise tensor is created with shape **[batch, latent_dim, latent_len]** via `sample_noisy_latent` (Python).

5. **Mask application**: The latent tensor is element-wise multiplied by the latent mask, zero-padding extra steps while preserving valid positions.

6. **Diffusion inference**: The masked latent is refined for `total_step` iterations in `TextToSpeech._infer`.

7. **Vocoder conversion**: The final latent representation converts to waveform audio via `vocoder_onnx`.

8. **Chunk stitching**: Individual chunk waveforms are concatenated with a fixed-duration silence buffer (0.3 s default) in `TextToSpeech.__call__` (Python).

Because the mask derives from the predicted duration rather than a fixed constant, the model adapts to any audio length dynamically.

## Cross-Platform Implementation Consistency

Supertonic maintains identical behavior across Python, Node.js, and Swift by replicating the same mathematical operations in each language. The `length_to_mask`, `get_latent_mask`, and `chunk_text` functions serve as the foundation in all three SDKs, ensuring that a 12-second utterance processes identically whether generated via Python scripts, Node.js servers, or iOS applications.

## Practical Code Examples

The following examples demonstrate how Supertonic automatically handles different audio lengths without manual configuration:

```python

# Python – synthesize two sentences of different lengths

from supertonic import load_text_to_speech, load_voice_style, chunk_text

tts = load_text_to_speech("path/to/onnx_dir")
style = load_voice_style(["style_M1.json"])

# Short sentence (≈1 s)

wav1, dur1 = tts("Hello world!", "en", style, total_step=8)

# Long paragraph (≈12 s)

long_text = "Lorem ipsum dolor sit amet, " * 100
wav2, dur2 = tts(long_text, "en", style, total_step=8)

print(dur1, dur2)          # Different durations are handled automatically

```

```javascript
// Node.js – generate speech for a variable‑length text
import { loadTextToSpeech, loadVoiceStyle } from "./helper.js";

const tts = await loadTextToSpeech("onnx_dir");
const style = await loadVoiceStyle(["style_M1.json"]);

const text = "Short text.";
const [wav, dur] = await tts(text, "en", style, 8);
console.log(`Generated ${dur}s audio`);

```

```swift
// Swift – produce audio for a long paragraph
let tts = try loadTextToSpeech(onnxDir: "onnx_dir")
let style = try loadVoiceStyle(paths: ["style_M1.json"])

let longParagraph = String(repeating: "Supertonic makes TTS easy. ", count: 200)
let (audio, duration) = try tts(text: longParagraph, lang: "en", style: style, totalStep: 8)

print("Audio length: \(duration) s")

```

## Summary

- **Length-to-mask conversion** creates binary validity masks from sequence lengths in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js), and [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift).
- **Latent-mask creation** translates predicted durations (seconds) into latent-space masks that control which diffusion steps are processed.
- **Automatic chunking** splits long texts (>300 tokens) into independent segments with 0.3-second silence gaps, then concatenates results.
- The pipeline derives masks from the duration predictor’s output, enabling dynamic adaptation to any audio length from sub-second to multi-minute.

## Frequently Asked Questions

### How does Supertonic handle very long texts?

Supertonic automatically splits input texts exceeding the context window (300 tokens for most languages, 120 for Korean/Japanese) using the `chunk_text()` function. Each segment is synthesized independently, then concatenated with a 0.3-second silence buffer. This occurs in `TextToSpeech.__call__` (Python), the `_infer` method (Node.js), and the equivalent Swift implementation.

### What is the maximum audio length Supertonic can generate?

There is no hardcoded maximum duration. The system is constrained only by available memory and the chunking mechanism. Because Supertonic processes long inputs as separate chunks that are later concatenated, it can theoretically generate hours of audio by stitching together hundreds of individual segments.

### Does the chunking process affect audio quality?

The chunking mechanism preserves audio quality by maintaining consistent acoustic conditions across segments. Each chunk uses the same voice style embedding and random seed initialization where applicable. The 0.3-second silence gap between chunks ensures natural prosodic breaks without jarring transitions.

### Which platforms support variable-length audio generation?

All official Supertonic SDKs support variable-length audio generation using identical logic: **Python** ([`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)), **Node.js** ([`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js)), and **Swift** ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)). The Web demo also implements these mechanisms, as documented in the repository’s [`web/README.md`](https://github.com/supertone-inc/supertonic/blob/main/web/README.md).