How Supertonic Preprocesses Text: Unicode Normalization and Emoji Removal
Supertonic's UnicodeProcessor class in py/helper.py executes a 10-step normalization pipeline that converts raw Unicode text into clean, tokenized input by applying NFKD decomposition, stripping emojis and symbols via regex, standardizing punctuation, and wrapping results in XML language tags.
The supertone-inc/supertonic repository implements a neural text-to-speech engine that requires rigorous preprocessing to handle real-world input. Before reaching the ONNX encoder, every string passes through the UnicodeProcessor._preprocess_text method, which systematically eliminates noise and enforces linguistic consistency.
The UnicodeProcessor Implementation
The preprocessing logic resides in the UnicodeProcessor class defined in py/helper.py. When invoked via the __call__ method, the processor iterates over input batches and delegates string cleaning to the private _preprocess_text method. This method accepts a text string and a lang code, executing a fixed sequence of normalization steps before returning the formatted result.
Step-by-Step Preprocessing Pipeline
1. Unicode Normalization (NFKD)
The pipeline begins by decomposing characters into their canonical base forms using Python's unicodedata module. Specifically, normalize("NFKD", text) separates diacritics from base characters, ensuring that downstream tokenizers handle accented characters predictably. This operation occurs at lines 21-23 of py/helper.py.
2. Emoji and Wide-Unicode Symbol Removal
Next, a compiled regex pattern removes emojis, emoticons, transport symbols, and pictographs that occupy Unicode blocks without linguistic meaning. The pattern covers U+1F600-U+1FAFF (emoticons), U+2600-U+26FF (miscellaneous symbols), U+2700-U+27BF (dingbats), and flag sequences, applied via emoji_pattern.sub("", text) at lines 25-42.
3. Punctuation and Character Normalization
The processor replaces "noisy" punctuation variants with simpler equivalents using a dictionary-based replacement loop. Curly quotes become straight quotes, en-dashes and em-dashes convert to hyphens, and various brackets, pipes, and arrows normalize to standard ASCII or space characters (lines 44-65).
4. Miscellaneous Symbol Removal
Additional informal symbols—hearts (♥, ♡), stars (☆), and copyright signs (©)—are stripped via a secondary regex re.sub(r"[♥☆♡©\\]", "", text) at lines 66-68, removing visual clutter that lacks phonetic value.
5. Expression Normalization and Spacing Fixes
Common shorthand tokens expand to explicit words: @ becomes "at", while e.g., and i.e., convert to their textual equivalents. Simultaneously, the pipeline corrects spacing errors by collapsing spaces that appear before commas, periods, exclamation marks, question marks, semicolons, colons, and apostrophes using targeted re.sub calls (lines 70-84).
6. Whitespace and Quote Deduplication
Repeated quotation marks ("", '', and backticks) reduce to single instances through iterative replacement loops (lines 87-93), preventing malformed string artifacts. The method then collapses any whitespace sequence to a single space and trims leading or trailing spaces using re.sub(r"\s+", " ", text).strip() (lines 95-96).
7. Final Formatting and Language Tagging
The pipeline guarantees that the final character is sentence-ending punctuation (or a closing quote/bracket), adding a period if necessary (lines 98-100). Finally, it validates the supplied lang parameter against an AVAILABLE_LANGS whitelist and wraps the cleaned text in XML-style tags: f"<{lang}>{text}</{lang}>" (lines 102-105).
Practical Usage Examples
Preprocessing a Single String
Access the _preprocess_text method directly to inspect how individual strings transform:
from supertonic.py.helper import UnicodeProcessor
# Path to the Unicode indexer generated during model export
indexer_path = "path/to/unicode_indexer.json"
# Initialise the processor
processor = UnicodeProcessor(indexer_path)
# Raw user input (contains emojis, fancy quotes, and an en‑dash)
raw = "Hey 👋! Let’s meet at 10 am – don’t be late “please”."
# Internally the processor will call _preprocess_text for each element
cleaned = processor._preprocess_text(raw, lang="en")
print(cleaned)
# Expected output (formatted with language tags):
# <en>Hey! Let's meet at 10 am - don't be late "please".</en>
Batch Processing via Callable Interface
For inference, call the processor object directly to handle lists efficiently:
texts = [
"Good morning 🌅! How are you?",
"안녕하세요? 오늘도 화이팅! 💪"
]
langs = ["en", "ko"]
# UnicodeProcessor implements __call__ → it processes a list in one go
ids, mask = processor(texts, langs) # returns integer IDs and attention mask
Integration with the TTS Pipeline
In production workflows, preprocessing happens automatically before the ONNX encoder:
from supertonic.py.helper import load_text_to_speech, load_voice_style
onnx_dir = "models/tts_onnx"
tts = load_text_to_speech(onnx_dir, use_gpu=False)
style_paths = ["style_1.json", "style_2.json"]
style = load_voice_style(style_paths)
wav, duration = tts("I love pizza 🍕!", "en", style, total_step=30)
Summary
- Location: The
UnicodeProcessorclass inpy/helper.pyorchestrates all text cleaning. - Unicode Handling: NFKD normalization (lines 21-23) decomposes characters for reliable token mapping.
- Emoji Removal: Regex patterns covering U+1F600-U+1FAFF and related blocks strip non-linguistic symbols (lines 25-42, 66-68).
- Punctuation Standardization: Dictionary replacements normalize dashes, quotes, and special characters (lines 44-65).
- Formatting: The pipeline enforces sentence-ending punctuation and wraps output in XML language tags (lines 98-105).
Frequently Asked Questions
Does Supertonic support all Unicode languages?
The processor validates language codes against an internal AVAILABLE_LANGS list before wrapping text in XML tags (lines 102-105). Only codes present in this whitelist proceed to the encoder, ensuring model compatibility.
What happens to emojis that aren't removed by the initial regex?
After NFKD decomposition separates combined characters, the primary emoji regex covers U+1F600-U+1FAFF, miscellaneous symbols, and flags (lines 25-42). Any remaining symbols like hearts or stars are caught by a secondary regex at lines 66-68, ensuring comprehensive visual noise removal.
Can I customize the punctuation replacements?
The _preprocess_text method uses hardcoded replacement dictionaries defined at lines 44-65 and 70-76. To modify behavior, you must edit py/helper.py directly or subclass UnicodeProcessor to override the replacement logic before model export.
Why does the pipeline enforce sentence-ending punctuation?
Downstream TTS models require properly delimited input for attention mechanisms. Lines 98-100 check whether the final character is punctuation, a closing quote, or bracket; if not, the method appends a period to guarantee a complete sentence boundary.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →