How Supertonic Handles Different Audio Lengths in Text-to-Speech Synthesis
Supertonic handles variable audio lengths through length-to-mask conversion, latent-mask creation based on predicted durations, and automatic chunking with concatenation for long texts, implemented consistently across Python, Node.js, and Swift SDKs.
The supertone-inc/supertonic repository provides a flexible text-to-speech pipeline that automatically adapts to inputs of any duration, from sub-second utterances to multi-minute passages. Unlike fixed-length TTS systems, Supertonic dynamically adjusts its internal representations using duration prediction and binary masking mechanisms. This article examines the three core architectural patterns that enable Supertonic to handle different audio lengths across all language bindings.
Three Mechanisms for Variable-Length Audio
Supertonic’s ability to process arbitrary durations rests on three coordinated mechanisms that appear in every language binding: Python, Node.js, Swift, and the Web demo.
Length-to-Mask Conversion
The foundation of Supertonic’s flexibility is the conversion of sequence lengths into binary masks that tell the diffusion model which time-steps are valid data versus padding.
In py/helper.py (lines 57‑71), the length_to_mask() function takes a list of sequence lengths and produces a boolean tensor where True values indicate valid positions. This mechanism is replicated across platforms:
- Node.js:
lengthToMask()innodejs/helper.js(lines 306‑311) - Swift:
lengthToMask(_:maxLen:)inswift/Sources/Helper.swift(lines 97‑108)
These masks ensure that regardless of input length, the model only processes valid audio frames while maintaining consistent tensor shapes for batch processing.
Latent-Mask Creation
After the duration predictor yields a target speech length in seconds, Supertonic converts this duration into a latent count—the internal representation used by the diffusion model. This latent mask is applied to randomly-generated noise tensors so only the required time-steps undergo refinement.
In py/helper.py (lines 74‑80), the get_latent_mask() function orchestrates this conversion: predicted seconds become wav-sample counts, then latent steps, then binary masks. The cross-platform implementations follow the same logic:
- Node.js:
getLatentMask()innodejs/helper.js(lines 166‑172), which callslengthToMask() - Swift:
getLatentMask(_:_)inswift/Sources/Helper.swift(lines 66‑78)
This mask is element-wise multiplied against the Gaussian noise tensor, zero-padding any extra steps while preserving the random latent initialization for valid positions.
Chunking and Concatenation
For texts exceeding the model’s context window, Supertonic automatically splits input into manageable chunks—approximately 300 tokens for most languages, or 120 tokens for Korean and Japanese. Each chunk is synthesized independently, then stitched together with a 0.3-second silence gap to produce seamless final waveforms.
The chunking logic resides in py/helper.py (lines 28‑44), where chunk_text() splits the input and the TextToSpeech.__call__ method handles concatenation. Equivalent implementations exist in:
- Node.js:
_inferandbatchmethods innodejs/helper.js(lines 30‑44), which create silence buffers between chunks - Swift:
chunkText(_:maxLen:)inswift/Sources/Helper.swift(lines 34‑42), followed by post-processing that appends silence
The Synthesis Pipeline Step-by-Step
When a user invokes the top-level TextToSpeech object (Python) or calls textToSpeech (Node.js/Swift), the library executes an eight-step pipeline that handles different audio lengths automatically:
-
Text preprocessing: Unicode normalization, emoji removal, and language-specific tokenization occur in
UnicodeProcessor._preprocessTextacross all platforms. -
Duration prediction: The duration predictor ONNX model (
dp_onnx) returns a float array of seconds for each input chunk viaTextToSpeech._infer→dp_ort.run(Python). -
Latent mask generation:
get_latent_mask()converts predicted seconds to wav samples, then to latent steps, then to a binary mask. -
Random latent creation: A Gaussian noise tensor is created with shape [batch, latent_dim, latent_len] via
sample_noisy_latent(Python). -
Mask application: The latent tensor is element-wise multiplied by the latent mask, zero-padding extra steps while preserving valid positions.
-
Diffusion inference: The masked latent is refined for
total_stepiterations inTextToSpeech._infer. -
Vocoder conversion: The final latent representation converts to waveform audio via
vocoder_onnx. -
Chunk stitching: Individual chunk waveforms are concatenated with a fixed-duration silence buffer (0.3 s default) in
TextToSpeech.__call__(Python).
Because the mask derives from the predicted duration rather than a fixed constant, the model adapts to any audio length dynamically.
Cross-Platform Implementation Consistency
Supertonic maintains identical behavior across Python, Node.js, and Swift by replicating the same mathematical operations in each language. The length_to_mask, get_latent_mask, and chunk_text functions serve as the foundation in all three SDKs, ensuring that a 12-second utterance processes identically whether generated via Python scripts, Node.js servers, or iOS applications.
Practical Code Examples
The following examples demonstrate how Supertonic automatically handles different audio lengths without manual configuration:
# Python – synthesize two sentences of different lengths
from supertonic import load_text_to_speech, load_voice_style, chunk_text
tts = load_text_to_speech("path/to/onnx_dir")
style = load_voice_style(["style_M1.json"])
# Short sentence (≈1 s)
wav1, dur1 = tts("Hello world!", "en", style, total_step=8)
# Long paragraph (≈12 s)
long_text = "Lorem ipsum dolor sit amet, " * 100
wav2, dur2 = tts(long_text, "en", style, total_step=8)
print(dur1, dur2) # Different durations are handled automatically
// Node.js – generate speech for a variable‑length text
import { loadTextToSpeech, loadVoiceStyle } from "./helper.js";
const tts = await loadTextToSpeech("onnx_dir");
const style = await loadVoiceStyle(["style_M1.json"]);
const text = "Short text.";
const [wav, dur] = await tts(text, "en", style, 8);
console.log(`Generated ${dur}s audio`);
// Swift – produce audio for a long paragraph
let tts = try loadTextToSpeech(onnxDir: "onnx_dir")
let style = try loadVoiceStyle(paths: ["style_M1.json"])
let longParagraph = String(repeating: "Supertonic makes TTS easy. ", count: 200)
let (audio, duration) = try tts(text: longParagraph, lang: "en", style: style, totalStep: 8)
print("Audio length: \(duration) s")
Summary
- Length-to-mask conversion creates binary validity masks from sequence lengths in
py/helper.py,nodejs/helper.js, andswift/Sources/Helper.swift. - Latent-mask creation translates predicted durations (seconds) into latent-space masks that control which diffusion steps are processed.
- Automatic chunking splits long texts (>300 tokens) into independent segments with 0.3-second silence gaps, then concatenates results.
- The pipeline derives masks from the duration predictor’s output, enabling dynamic adaptation to any audio length from sub-second to multi-minute.
Frequently Asked Questions
How does Supertonic handle very long texts?
Supertonic automatically splits input texts exceeding the context window (300 tokens for most languages, 120 for Korean/Japanese) using the chunk_text() function. Each segment is synthesized independently, then concatenated with a 0.3-second silence buffer. This occurs in TextToSpeech.__call__ (Python), the _infer method (Node.js), and the equivalent Swift implementation.
What is the maximum audio length Supertonic can generate?
There is no hardcoded maximum duration. The system is constrained only by available memory and the chunking mechanism. Because Supertonic processes long inputs as separate chunks that are later concatenated, it can theoretically generate hours of audio by stitching together hundreds of individual segments.
Does the chunking process affect audio quality?
The chunking mechanism preserves audio quality by maintaining consistent acoustic conditions across segments. Each chunk uses the same voice style embedding and random seed initialization where applicable. The 0.3-second silence gap between chunks ensures natural prosodic breaks without jarring transitions.
Which platforms support variable-length audio generation?
All official Supertonic SDKs support variable-length audio generation using identical logic: Python (py/helper.py), Node.js (nodejs/helper.js), and Swift (swift/Sources/Helper.swift). The Web demo also implements these mechanisms, as documented in the repository’s web/README.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →