Supertonic Text Preprocessing and Unicode Normalization: Complete Pipeline Guide
Supertonic converts raw input strings into clean, language-tagged Unicode representations using a ten-stage normalization pipeline implemented in the UnicodeProcessor class before feeding them to ONNX-based text-to-speech models.
The Supertonic open-source text-to-speech library sanitizes real-world text through rigorous preprocessing defined in py/helper.py. This pipeline handles arbitrary Unicode input—including emojis, special symbols, and mixed punctuation—and transforms it into deterministic token sequences suitable for neural TTS inference.
The UnicodeProcessor Normalization Pipeline
The preprocessing logic resides in UnicodeProcessor._preprocess_text within [py/helper.py](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py). The method executes ten distinct transformation stages to ensure model-ready output.
1. Unicode Compatibility Decomposition
The processor first applies NFKD normalization using unicodedata.normalize("NFKD", text) (lines 21-24). This compatibility decomposition separates combined characters into base glyphs and combining marks, standardizing code points for consistent downstream processing.
2. Emoji Stripping
A compiled regular expression defined in lines 26-41 removes emoji characters across wide Unicode ranges. This prevents unsupported symbols from entering the acoustic model while preserving valid linguistic content.
3. Symbol and Dash Normalization
The pipeline replaces various dash characters, underscores, quotation marks, and miscellaneous punctuation with simpler equivalents or spaces using a replacements dictionary (lines 44-65).
4. Special Symbol Removal
Characters such as hearts (♥, ♡), stars (☆), and copyright symbols (©) are stripped via re.sub(r"[♥☆♡©\\]", "", text) (lines 66-68) to ensure clean phonetic processing.
5. Expression Expansion
Abbreviations and shortcuts including "@", "e.g.", and "i.e." expand into full words through the expr_replacements dictionary (lines 70-77), improving pronunciation accuracy.
6. Punctuation Spacing Normalization
The processor fixes spacing around commas, periods, exclamation points, question marks, semicolons, colons, and apostrophes using a series of re.sub calls (lines 78-86).
7. Quote Deduplication
Consecutive double-quotes, single-quotes, and back-ticks collapse into single characters via while loops (lines 87-94), eliminating formatting artifacts from user input.
8. Whitespace Normalization
Multiple spaces collapse into single spaces with re.sub(r"\s+", " ", text).strip() (lines 95-96), and leading or trailing whitespace is trimmed.
9. Sentence Termination
The system guarantees proper sentence boundaries by checking for terminal punctuation. If none exists, it appends a period using if not re.search(...): text += "." (lines 98-101).
10. Language Tagging
Finally, the cleaned text wraps in XML-style language tags (<en>…</en>, <ja>…</ja>) after validating the language code (lines 102-105), producing the final string format consumed by the unicode indexer.
Token Conversion and Model Integration
The UnicodeProcessor.__call__ method orchestrates the full workflow. After preprocessing, it converts the Unicode string into integer IDs using the pre-computed unicode indexer JSON file (models/onnx/unicode_indexer.json), returning a padded ID matrix and binary attention mask (text_mask) ready for ONNX inference.
Implementation Examples
Basic Python Usage
The following example demonstrates initializing the processor and preparing mixed-language batches:
from helper import UnicodeProcessor, load_text_processor
# Path to the json file that maps Unicode code points to model-specific IDs
unicode_indexer_path = "models/onnx/unicode_indexer.json"
# Initialise the processor
processor = UnicodeProcessor(unicode_indexer_path)
texts = [
"Hello world! 😊 This is a test – with dashes, emojis, and “quotes”.",
"こんにちは、世界!"
]
langs = ["en", "ja"]
# Get token IDs and mask ready for the ONNX model
text_ids, text_mask = processor(texts, langs)
print("IDs shape:", text_ids.shape) # (batch, max_len)
print("Mask shape:", text_mask.shape) # (batch, 1, max_len)
End-to-End TTS Inference
The _preprocess_text logic executes automatically within the high-level TTS interface:
from helper import (
load_text_to_speech,
load_voice_style,
chunk_text,
)
# Load the whole TTS stack (config + ONNX sessions)
tts = load_text_to_speech(onnx_dir="models/onnx", use_gpu=False)
# Load a single speaker style
style = load_voice_style(["styles/english_female.json"])
# Long input text – the system will chunk automatically
raw_text = """Supertonic provides state‑of‑the‑art TTS.
It handles Unicode, removes emojis, and normalises punctuation."""
wav, duration = tts(
text=raw_text,
lang="en",
style=style,
total_step=60,
speed=1.05,
silence_duration=0.3,
)
# `wav` is a NumPy array (1, samples) ready to be written to a WAV file
Swift Implementation
Supertonic mirrors this logic in Swift within [swift/Sources/Helper.swift](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift):
let processor = try UnicodeProcessor(unicodeIndexerPath: "unicode_indexer.json")
let (ids, mask) = try processor.process(texts: ["Hello 🌟"], langs: ["en"])
Summary
- Supertonic text preprocessing employs a ten-stage pipeline in
UnicodeProcessor._preprocess_textto normalize raw Unicode input before TTS inference. - The system applies NFKD normalization to decompose characters, strips emojis via compiled regex, and replaces special symbols with standardized equivalents.
- Punctuation spacing fixes, quote deduplication, and forced sentence termination ensure consistent formatting for the acoustic model.
- The
UnicodeProcessor.__call__method converts cleaned text to integer IDs usingunicode_indexer.json, outputting padded tensors and binary masks for ONNX inference. - Both Python (
py/helper.py) and Swift (swift/Sources/Helper.swift) implementations provide identical preprocessing guarantees across platforms.
Frequently Asked Questions
What Unicode normalization form does Supertonic use?
Supertonic applies NFKD (Compatibility Decomposition) via unicodedata.normalize("NFKD", text) in the _preprocess_text method (lines 21-24). This separates base characters from combining marks, ensuring consistent code point representation before emoji removal and symbol replacement occur.
Why does the pipeline strip emojis instead of converting them to words?
The preprocessor removes emojis using a compiled regex pattern covering wide Unicode ranges (lines 26-41 in py/helper.py). This prevents unsupported symbols from reaching the ONNX-based acoustic models, as the current implementation focuses on spoken language content rather than descriptive text-to-speech conversion for pictographic symbols.
How does Supertonic ensure proper sentence boundaries?
The pipeline guarantees proper termination by checking if the text ends with punctuation using re.search. If no terminal punctuation exists, it appends a period (lines 98-101 in py/helper.py). This ensures the TTS model receives properly bounded phonetic units for natural prosody generation.
Can I extend Supertonic preprocessing to languages beyond English and Japanese?
Yes. The UnicodeProcessor validates language codes and wraps output in XML-style tags like <en> or <ja> (lines 102-105). You can extend support to additional languages by ensuring models/onnx/unicode_indexer.json contains the necessary code point mappings for your target language's character set, as the NFKD normalization and symbol cleaning handle Unicode generically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →