Supertonic Expressive Tags: Adding Laughter, Breathing, and Sighs to Synthesized Speech
Supertonic supports expressive tags like <laugh>, <breath>, and <sigh> that you insert directly into text strings to add natural vocal expressions to synthesized speech without requiring special API calls.
The supertone-inc/supertonic repository provides a lightweight text-to-speech (TTS) engine that uses ONNX Runtime to generate expressive, human-like speech. Unlike traditional TTS systems that require complex prosody controls, Supertonic expressive tags allow developers to trigger specific vocal effects—such as laughter, breathing pauses, and sighs—using simple XML-like markers embedded in the input text.
How Supertonic Expressive Tags Work
Supertonic's pipeline treats expressive tags as ordinary tokens rather than special metadata, allowing them to flow naturally through the text processing and inference stages.
Unicode Pre-processing Pipeline
The UnicodeProcessor class in [py/helper.py](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) handles text normalization before inference. While it strips emojis and adds language delimiters (such as <en>…</en>), the processor does not remove or transform custom angle-bracket tags. This means markers like <laugh>, <breath>, and <sigh> are preserved verbatim and passed to the model unchanged.
ONNX Model Inference
During the TextToSpeech._infer method (also in py/helper.py), the processed token IDs—including any expressive tags—are fed directly into the text encoder, duration predictor, vector estimator, and vocoder. The ONNX model has been specifically trained to recognize these tags and inject the corresponding prosodic cues: a chuckle for <laugh>, a short pause for <breath>, or a soft vocalization for <sigh>.
Using Expressive Tags in Python
The Python SDK processes expressive tags automatically when you include them in your input string. No additional configuration or API parameters are required.
from supertonic import TTS, load_voice_style, Style
# Load the pre-trained ONNX assets (auto-download on first run)
tts = TTS(auto_download=True)
# Load a voice style (replace with your own style file path)
style = load_voice_style(["assets/styles/M1.json"])
text = (
"Hey there, <laugh> thanks for joining us! "
"Now, let's take a quick <breath> before we continue. "
"Finally, we end with a soft <sigh>."
)
# Synthesize speech – expressive tags are interpreted by the model
wav, duration = tts.synthesize(text, voice_style=style, lang="en")
tts.save_audio(wav, "expressive_demo.wav")
print(f"Generated {duration:.2f}s of audio → expressive_demo.wav")
Key implementation details:
- Tags are inserted directly into the raw text string as
<tagname>format. - The
TTS.synthesize()method accepts the tagged text through its standardtextparameter. - The high-level API is exposed through [
py/__init__.py](https://github.com/supertone-inc/supertonic/blob/main/py/__init__.py).
Using Expressive Tags in Node.js
The Node.js SDK mirrors the Python API, maintaining identical tag semantics across JavaScript environments.
const { TTS, loadVoiceStyle } = require('supertonic');
// Initialise TTS (downloads ONNX assets on first run)
const tts = new TTS({ autoDownload: true });
async function run() {
const style = await loadVoiceStyle(['assets/styles/M1.json']);
const text = "Welcome! <laugh> This is a demo. <breath> Let's continue. <sigh>";
const { wav, duration } = await tts.synthesize(text, { voiceStyle: style, lang: 'en' });
await tts.saveAudio(wav, 'expressive_demo.wav');
console.log(`Generated ${duration.toFixed(2)} s → expressive_demo.wav`);
}
run();
The Node.js implementation, documented in [nodejs/README.md](https://github.com/supertone-inc/supertonic/blob/main/nodejs/README.md), processes the tagged text through the same ONNX pipeline as the Python version.
Using Expressive Tags in the Browser
Supertonic's Web SDK supports on-device inference with identical expressive tag functionality, enabling client-side TTS with vocal expressions.
<script type="module">
import { TTS } from "./web/supertonic.js";
async function demo() {
const tts = await TTS.create({ autoDownload: true });
const text = "Hello world! <laugh> This runs in the browser. <breath> Enjoy!";
const { wav } = await tts.synthesize(text, { lang: "en" });
const blob = new Blob([wav.buffer], { type: "audio/wav" });
const url = URL.createObjectURL(blob);
const audio = new Audio(url);
audio.play();
}
demo();
</script>
As documented in [web/README.md](https://github.com/supertone-inc/supertonic/blob/main/web/README.md), the WebAssembly-based runtime preserves expressive tags through the same Unicode preprocessing stage before ONNX inference.
Key Source Files and Architecture
Understanding the source code helps clarify why expressive tags work consistently across all platforms:
- [
py/helper.py](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) – Contains theUnicodeProcessorclass that normalizes input while preserving expressive tags, and theTextToSpeechclass that manages the ONNX inference pipeline via the_infermethod. - [
README.md](https://github.com/supertone-inc/supertonic/blob/main/README.md#L232) – Documents the expressive tag feature at line 232, confirming that Supertonic "supports simple expression tags such as<laugh>,<breath>, and<sigh>." - [
py/example_onnx.py](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) – Provides end-to-end usage examples that can be adapted to demonstrate expressive tags. - Node.js and Web SDKs – Located in
nodejs/andweb/directories respectively, implementing the same tag-preserving logic for JavaScript environments.
Summary
- Supertonic expressive tags (
<laugh>,<breath>,<sigh>) are preserved by theUnicodeProcessorinpy/helper.pyand interpreted directly by the ONNX model. - Tags work across all language bindings—Python, Node.js, Java, C++, Go, and Swift—without requiring platform-specific API calls.
- Insert tags directly into text strings; the
TTS.synthesize()methods handle them as ordinary tokens during inference. - The feature is documented in the repository README and demonstrated in the example scripts.
Frequently Asked Questions
What expressive tags does Supertonic support?
Supertonic officially supports <laugh> for vocalized chuckles, <breath> for breathing pauses, and <sigh> for soft exhalations. According to the source code in py/helper.py, any angle-bracket tag that passes through the Unicode processor unchanged will be treated as a token by the ONNX model, though these three tags are specifically trained and documented in the README at line 232.
Do expressive tags work in all Supertonic language bindings?
Yes. Because the UnicodeProcessor preserves all angle-bracket tags before they reach the ONNX inference engine, Supertonic expressive tags function identically across Python, Node.js, Java, C++, Go, Swift, and browser-based WebAssembly implementations. The tags are processed at the model level rather than the SDK level.
How does Supertonic process expressive tags during inference?
The TextToSpeech._infer method in py/helper.py feeds tokenized input—including expressive tags—directly into the text encoder and duration predictor components of the ONNX model. The model has been trained to associate specific tag tokens with acoustic features that produce laughter, breathing sounds, or sighs in the generated waveform.
Can I use multiple expressive tags in a single text input?
Yes. You can combine multiple tags within a single string, as demonstrated in the code examples. The preprocessor preserves all tags in sequence, and the model renders them with appropriate timing. For example, "Hello how are you today " will synthesize speech with a chuckle, followed later by a breath pause and a final sigh.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →