How Supertonic Handles Natural Text: Financial Expressions, Phone Numbers, and Technical Units
Supertonic processes natural text through a syntax-aware, language-agnostic Unicode normalization pipeline that preserves numeric patterns and abbreviations while stripping emojis and standardizing punctuation, enabling accurate TTS synthesis of financial figures, phone numbers, and technical units without manual preprocessing.
The supertone-inc/supertonic repository provides an ONNX-based text-to-speech engine that treats arbitrary user-written sentences as natural text. Its preprocessing pipeline prepares input for the acoustic model through deterministic normalization steps implemented identically across Python, Swift, Go, Java, and C++ bindings.
The Six-Stage Preprocessing Pipeline
Supertonic normalizes every input string through a consistent six-step sequence defined in py/helper.py and mirrored across all language bindings.
Unicode Normalization with NFKD
Every input string undergoes Compatibility Decomposition (NFKD) to split accented characters and Hangul Jamo into their base code points. This guarantees deterministic Unicode representation regardless of the source encoding.
In py/helper.py (lines 21-30), Python uses unicodedata.normalize('NFKD', text), while swift/Sources/Helper.swift (lines 59-67) implements the same decomposition using native Swift string normalization. This step ensures that characters like "é" become "e" + combining accent, preventing tokenization mismatches.
Emoji and Symbol Stripping
A compiled regular expression removes the full Unicode emoji range (U+1F600–U+1F64F, U+2600–U+26FF, and pictographic extensions) along with miscellaneous symbols. As implemented in py/helper.py (lines 25-41), this prevents non-speech tokens from reaching the ONNX acoustic model.
Standard-Character Replacements
The pipeline replaces typographic punctuation with ASCII equivalents. En-dashes (–) become hyphens (-), smart quotes (“”) become straight quotes ("), and stray brackets or pipes are removed. This standardization occurs in py/helper.py (lines 43-61) to ensure consistent token boundaries.
Whitespace and Punctuation Normalization
Duplicate spaces collapse into single spaces, and spacing around commas, periods, exclamation marks, and other terminal punctuation is standardized. The regex patterns in py/helper.py (lines 78-85) ensure each sentence ends with explicit punctuation, creating clean token boundaries for the downstream tokenizer.
Language Tagging
After cleaning, text is wrapped in XML-style language tags (<en> ... </en>, <ko> ... </ko>, etc.) to signal language-specific phoneme table selection. This tagging occurs in py/helper.py (lines 99-105) before chunking begins.
Smart Text Chunking
Long utterances split hierarchically: paragraph → sentence → chunk (≤300 tokens). The chunking logic respects common abbreviations like "Mr.", "Dr.", and "e.g." to avoid cutting within numbers or mid-abbreviation. Implemented in py/helper.py (lines 88-115) and mirrored in Swift (chunk_text), Go (chunkText), and C++, this ensures coherent prosody across chunk boundaries.
Processing Financial Expressions, Phone Numbers, and Technical Units
Because the pipeline is syntax-aware but semantics-agnostic, it preserves the structural patterns that the ONNX TTS model expects, without requiring hand-crafted phonetic annotations.
Financial expressions such as $5.2M and $450K retain their dollar signs and magnitude suffixes. The numeric tokenizer emits sequences like "five point two million" and "four hundred fifty thousand" because the $, digits, and M/K symbols remain attached through normalization.
Phone numbers including formats like (212) 555-0142 ext. 402 preserve parentheses, hyphens, and the literal "ext." string. Digits are spoken individually while "extension" renders naturally from the preserved "ext" abbreviation.
Technical units such as 2.3h and 30kph retain decimal points and unit abbreviations. The model's training on multilingual corpora containing these patterns allows it to map 2.3h to "two point three hours" and 30kph to "thirty kilometers per hour" without explicit substitution rules.
Implementation Across Language Bindings
The preprocessing logic executes identically across all supported languages before feeding data to the ONNX session.
Python
from supertonic import load_text_to_speech, load_voice_style, Style
tts = load_text_to_speech("onnx_dir")
style = load_voice_style(["voice_styles/F1.json"])[0]
wav, dur = tts(
text="The startup secured $5.2M in venture capital.",
lang="en",
style=style,
total_step=30)
Source: py/example_onnx.py (lines 1-13)
Swift
let tts = try TextToSpeech(onnxDir: "onnx_dir")
let style = try loadVoiceStyle(paths: ["voice_styles/F1.json"]).first!
let (wav, duration) = try tts(
text: "(212) 555-0142 ext. 402",
lang: "en",
style: style,
totalStep: 30)
Source: swift/ExampleONNX.swift (lines 1-12)
Go
processor := NewTextProcessor("onnx_dir/unicode_indexer.json")
textIds, textMask := processor.Process([]string{
"Our drone battery lasts 2.3h at 30kph."}, []string{"en"})
Source: go/helper.go (lines 340-359)
Key Source Files
py/helper.py: Canonical implementation of Unicode normalization, emoji stripping, punctuation fixing, language tagging, and chunking logic.swift/Sources/Helper.swift: macOS/iOS port using native Swift string operations and regular expressions.go/helper.go: Go implementation usingnorm.NFKDand regex utilities for the same preprocessing steps.cpp/helper.cpp: C++ version usingstd::regexand manual Unicode handling for performance-critical applications.
Summary
- NFKD Unicode decomposition ensures consistent character representation across all language bindings before tokenization.
- Emoji and symbol stripping removes non-speech content that could confuse the acoustic model.
- Syntax-aware chunking respects abbreviations and numeric patterns to maintain natural prosody boundaries.
- Semantics-agnostic preservation allows financial expressions, phone numbers, and technical units to pass through unchanged, relying on the TTS model's trained numeric tokenizers for correct pronunciation.
- Cross-platform consistency guarantees identical preprocessing whether using Python, Swift, Go, Java, or C++ bindings.
Frequently Asked Questions
Does Supertonic require manual formatting for currency symbols like $ or €?
No. The preprocessing pipeline preserves currency symbols and magnitude abbreviations (M, K, B) as literal characters. The ONNX model's tokenizer, trained on multilingual corpora containing these patterns, maps $5.2M to the phoneme sequence for "five point two million dollars" without explicit substitution rules.
How does Supertonic handle phone number extensions like "ext. 402"?
The pipeline treats "ext." as regular punctuation attached to digits. Because the chunking logic preserves abbreviations and the tokenizer recognizes "ext" as "extension," the acoustic model renders "(212) 555-0142 ext. 402" with appropriate pauses and the full word "extension" followed by the digits "four zero two."
Can Supertonic process technical abbreviations in languages other than English?
Yes. After Unicode normalization and language tagging (e.g., <ko> for Korean), unit abbreviations like kph or h pass through unchanged regardless of the language tag. The model selects phoneme tables based on the wrapper tags while retaining the literal abbreviations, enabling correct multilingual pronunciation of technical measurements.
What happens to emojis and special Unicode characters in the input text?
Emojis and pictographs are removed entirely by a compiled regex covering the Unicode emoji ranges (U+1F600–U+1F64F and extensions) as defined in py/helper.py (lines 25-41). This prevents silent tokens or unexpected phonemes from reaching the acoustic model, ensuring clean synthesis of the remaining text content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →