# How Supertonic Preprocesses Text: Unicode Normalization and Emoji Removal

> Discover how Supertonic preprocesses text with its UnicodeProcessor. Learn about Unicode normalization, emoji removal, and a 10-step pipeline for clean, tokenized input.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-15

---

**Supertonic's `UnicodeProcessor` class in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) executes a 10-step normalization pipeline that converts raw Unicode text into clean, tokenized input by applying NFKD decomposition, stripping emojis and symbols via regex, standardizing punctuation, and wrapping results in XML language tags.**

The supertone-inc/supertonic repository implements a neural text-to-speech engine that requires rigorous preprocessing to handle real-world input. Before reaching the ONNX encoder, every string passes through the `UnicodeProcessor._preprocess_text` method, which systematically eliminates noise and enforces linguistic consistency.

## The UnicodeProcessor Implementation

The preprocessing logic resides in the `UnicodeProcessor` class defined in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py). When invoked via the `__call__` method, the processor iterates over input batches and delegates string cleaning to the private `_preprocess_text` method. This method accepts a `text` string and a `lang` code, executing a fixed sequence of normalization steps before returning the formatted result.

## Step-by-Step Preprocessing Pipeline

### 1. Unicode Normalization (NFKD)

The pipeline begins by decomposing characters into their canonical base forms using Python's `unicodedata` module. Specifically, `normalize("NFKD", text)` separates diacritics from base characters, ensuring that downstream tokenizers handle accented characters predictably. This operation occurs at lines 21-23 of [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py).

### 2. Emoji and Wide-Unicode Symbol Removal

Next, a compiled regex pattern removes emojis, emoticons, transport symbols, and pictographs that occupy Unicode blocks without linguistic meaning. The pattern covers U+1F600-U+1FAFF (emoticons), U+2600-U+26FF (miscellaneous symbols), U+2700-U+27BF (dingbats), and flag sequences, applied via `emoji_pattern.sub("", text)` at lines 25-42.

### 3. Punctuation and Character Normalization

The processor replaces "noisy" punctuation variants with simpler equivalents using a dictionary-based replacement loop. Curly quotes become straight quotes, en-dashes and em-dashes convert to hyphens, and various brackets, pipes, and arrows normalize to standard ASCII or space characters (lines 44-65).

### 4. Miscellaneous Symbol Removal

Additional informal symbols—hearts (♥, ♡), stars (☆), and copyright signs (©)—are stripped via a secondary regex `re.sub(r"[♥☆♡©\\]", "", text)` at lines 66-68, removing visual clutter that lacks phonetic value.

### 5. Expression Normalization and Spacing Fixes

Common shorthand tokens expand to explicit words: `@` becomes "at", while `e.g.,` and `i.e.,` convert to their textual equivalents. Simultaneously, the pipeline corrects spacing errors by collapsing spaces that appear before commas, periods, exclamation marks, question marks, semicolons, colons, and apostrophes using targeted `re.sub` calls (lines 70-84).

### 6. Whitespace and Quote Deduplication

Repeated quotation marks (`""`, `''`, and backticks) reduce to single instances through iterative replacement loops (lines 87-93), preventing malformed string artifacts. The method then collapses any whitespace sequence to a single space and trims leading or trailing spaces using `re.sub(r"\s+", " ", text).strip()` (lines 95-96).

### 7. Final Formatting and Language Tagging

The pipeline guarantees that the final character is sentence-ending punctuation (or a closing quote/bracket), adding a period if necessary (lines 98-100). Finally, it validates the supplied `lang` parameter against an `AVAILABLE_LANGS` whitelist and wraps the cleaned text in XML-style tags: `f"<{lang}>{text}</{lang}>"` (lines 102-105).

## Practical Usage Examples

### Preprocessing a Single String

Access the `_preprocess_text` method directly to inspect how individual strings transform:

```python
from supertonic.py.helper import UnicodeProcessor

# Path to the Unicode indexer generated during model export

indexer_path = "path/to/unicode_indexer.json"

# Initialise the processor

processor = UnicodeProcessor(indexer_path)

# Raw user input (contains emojis, fancy quotes, and an en‑dash)

raw = "Hey 👋! Let’s meet at 10 am – don’t be late “please”."

# Internally the processor will call _preprocess_text for each element

cleaned = processor._preprocess_text(raw, lang="en")
print(cleaned)

# Expected output (formatted with language tags):

# <en>Hey! Let's meet at 10 am - don't be late "please".</en>

```

### Batch Processing via Callable Interface

For inference, call the processor object directly to handle lists efficiently:

```python
texts = [
    "Good morning 🌅! How are you?",
    "안녕하세요? 오늘도 화이팅! 💪"
]
langs = ["en", "ko"]

# UnicodeProcessor implements __call__ → it processes a list in one go

ids, mask = processor(texts, langs)   # returns integer IDs and attention mask

```

### Integration with the TTS Pipeline

In production workflows, preprocessing happens automatically before the ONNX encoder:

```python
from supertonic.py.helper import load_text_to_speech, load_voice_style

onnx_dir = "models/tts_onnx"
tts = load_text_to_speech(onnx_dir, use_gpu=False)

style_paths = ["style_1.json", "style_2.json"]
style = load_voice_style(style_paths)

wav, duration = tts("I love pizza 🍕!", "en", style, total_step=30)

```

## Summary

- **Location**: The `UnicodeProcessor` class in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) orchestrates all text cleaning.
- **Unicode Handling**: NFKD normalization (lines 21-23) decomposes characters for reliable token mapping.
- **Emoji Removal**: Regex patterns covering U+1F600-U+1FAFF and related blocks strip non-linguistic symbols (lines 25-42, 66-68).
- **Punctuation Standardization**: Dictionary replacements normalize dashes, quotes, and special characters (lines 44-65).
- **Formatting**: The pipeline enforces sentence-ending punctuation and wraps output in XML language tags (lines 98-105).

## Frequently Asked Questions

### Does Supertonic support all Unicode languages?

The processor validates language codes against an internal `AVAILABLE_LANGS` list before wrapping text in XML tags (lines 102-105). Only codes present in this whitelist proceed to the encoder, ensuring model compatibility.

### What happens to emojis that aren't removed by the initial regex?

After NFKD decomposition separates combined characters, the primary emoji regex covers U+1F600-U+1FAFF, miscellaneous symbols, and flags (lines 25-42). Any remaining symbols like hearts or stars are caught by a secondary regex at lines 66-68, ensuring comprehensive visual noise removal.

### Can I customize the punctuation replacements?

The `_preprocess_text` method uses hardcoded replacement dictionaries defined at lines 44-65 and 70-76. To modify behavior, you must edit [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) directly or subclass `UnicodeProcessor` to override the replacement logic before model export.

### Why does the pipeline enforce sentence-ending punctuation?

Downstream TTS models require properly delimited input for attention mechanisms. Lines 98-100 check whether the final character is punctuation, a closing quote, or bracket; if not, the method appends a period to guarantee a complete sentence boundary.