# Supertonic Text Preprocessing and Unicode Normalization: Complete Pipeline Guide

> Learn the Supertonic text preprocessing and Unicode normalization pipeline. Convert raw text to clean Unicode for ONNX TTS models with this complete guide.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-05-14

---

**Supertonic converts raw input strings into clean, language-tagged Unicode representations using a ten-stage normalization pipeline implemented in the `UnicodeProcessor` class before feeding them to ONNX-based text-to-speech models.**

The **Supertonic** open-source text-to-speech library sanitizes real-world text through rigorous preprocessing defined in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py). This pipeline handles arbitrary Unicode input—including emojis, special symbols, and mixed punctuation—and transforms it into deterministic token sequences suitable for neural TTS inference.

## The UnicodeProcessor Normalization Pipeline

The preprocessing logic resides in `UnicodeProcessor._preprocess_text` within [[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py). The method executes ten distinct transformation stages to ensure model-ready output.

### 1. Unicode Compatibility Decomposition

The processor first applies **NFKD normalization** using `unicodedata.normalize("NFKD", text)` (lines 21-24). This compatibility decomposition separates combined characters into base glyphs and combining marks, standardizing code points for consistent downstream processing.

### 2. Emoji Stripping

A compiled regular expression defined in lines 26-41 removes emoji characters across wide Unicode ranges. This prevents unsupported symbols from entering the acoustic model while preserving valid linguistic content.

### 3. Symbol and Dash Normalization

The pipeline replaces various dash characters, underscores, quotation marks, and miscellaneous punctuation with simpler equivalents or spaces using a `replacements` dictionary (lines 44-65).

### 4. Special Symbol Removal

Characters such as hearts (♥, ♡), stars (☆), and copyright symbols (©) are stripped via `re.sub(r"[♥☆♡©\\]", "", text)` (lines 66-68) to ensure clean phonetic processing.

### 5. Expression Expansion

Abbreviations and shortcuts including "@", "e.g.", and "i.e." expand into full words through the `expr_replacements` dictionary (lines 70-77), improving pronunciation accuracy.

### 6. Punctuation Spacing Normalization

The processor fixes spacing around commas, periods, exclamation points, question marks, semicolons, colons, and apostrophes using a series of `re.sub` calls (lines 78-86).

### 7. Quote Deduplication

Consecutive double-quotes, single-quotes, and back-ticks collapse into single characters via `while` loops (lines 87-94), eliminating formatting artifacts from user input.

### 8. Whitespace Normalization

Multiple spaces collapse into single spaces with `re.sub(r"\s+", " ", text).strip()` (lines 95-96), and leading or trailing whitespace is trimmed.

### 9. Sentence Termination

The system guarantees proper sentence boundaries by checking for terminal punctuation. If none exists, it appends a period using `if not re.search(...): text += "."` (lines 98-101).

### 10. Language Tagging

Finally, the cleaned text wraps in XML-style language tags (`<en>…</en>`, `<ja>…</ja>`) after validating the language code (lines 102-105), producing the final string format consumed by the unicode indexer.

## Token Conversion and Model Integration

The **`UnicodeProcessor.__call__`** method orchestrates the full workflow. After preprocessing, it converts the Unicode string into integer IDs using the pre-computed **unicode indexer** JSON file ([`models/onnx/unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/models/onnx/unicode_indexer.json)), returning a padded ID matrix and binary attention mask (`text_mask`) ready for ONNX inference.

## Implementation Examples

### Basic Python Usage

The following example demonstrates initializing the processor and preparing mixed-language batches:

```python
from helper import UnicodeProcessor, load_text_processor

# Path to the json file that maps Unicode code points to model-specific IDs

unicode_indexer_path = "models/onnx/unicode_indexer.json"

# Initialise the processor

processor = UnicodeProcessor(unicode_indexer_path)

texts = [
    "Hello world! 😊 This is a test – with dashes, emojis, and “quotes”.",
    "こんにちは、世界！"
]
langs = ["en", "ja"]

# Get token IDs and mask ready for the ONNX model

text_ids, text_mask = processor(texts, langs)

print("IDs shape:", text_ids.shape)      # (batch, max_len)

print("Mask shape:", text_mask.shape)    # (batch, 1, max_len)

```

### End-to-End TTS Inference

The `_preprocess_text` logic executes automatically within the high-level TTS interface:

```python
from helper import (
    load_text_to_speech,
    load_voice_style,
    chunk_text,
)

# Load the whole TTS stack (config + ONNX sessions)

tts = load_text_to_speech(onnx_dir="models/onnx", use_gpu=False)

# Load a single speaker style

style = load_voice_style(["styles/english_female.json"])

# Long input text – the system will chunk automatically

raw_text = """Supertonic provides state‑of‑the‑art TTS. 
It handles Unicode, removes emojis, and normalises punctuation."""

wav, duration = tts(
    text=raw_text,
    lang="en",
    style=style,
    total_step=60,
    speed=1.05,
    silence_duration=0.3,
)

# `wav` is a NumPy array (1, samples) ready to be written to a WAV file

```

### Swift Implementation

Supertonic mirrors this logic in Swift within [[`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift):

```swift
let processor = try UnicodeProcessor(unicodeIndexerPath: "unicode_indexer.json")
let (ids, mask) = try processor.process(texts: ["Hello 🌟"], langs: ["en"])

```

## Summary

- **Supertonic text preprocessing** employs a ten-stage pipeline in `UnicodeProcessor._preprocess_text` to normalize raw Unicode input before TTS inference.
- The system applies **NFKD normalization** to decompose characters, strips emojis via compiled regex, and replaces special symbols with standardized equivalents.
- **Punctuation spacing fixes**, quote deduplication, and forced sentence termination ensure consistent formatting for the acoustic model.
- The `UnicodeProcessor.__call__` method converts cleaned text to integer IDs using [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json), outputting padded tensors and binary masks for ONNX inference.
- Both **Python** ([`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)) and **Swift** ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)) implementations provide identical preprocessing guarantees across platforms.

## Frequently Asked Questions

### What Unicode normalization form does Supertonic use?

Supertonic applies **NFKD (Compatibility Decomposition)** via `unicodedata.normalize("NFKD", text)` in the `_preprocess_text` method (lines 21-24). This separates base characters from combining marks, ensuring consistent code point representation before emoji removal and symbol replacement occur.

### Why does the pipeline strip emojis instead of converting them to words?

The preprocessor removes emojis using a compiled regex pattern covering wide Unicode ranges (lines 26-41 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)). This prevents unsupported symbols from reaching the ONNX-based acoustic models, as the current implementation focuses on spoken language content rather than descriptive text-to-speech conversion for pictographic symbols.

### How does Supertonic ensure proper sentence boundaries?

The pipeline guarantees proper termination by checking if the text ends with punctuation using `re.search`. If no terminal punctuation exists, it appends a period (lines 98-101 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)). This ensures the TTS model receives properly bounded phonetic units for natural prosody generation.

### Can I extend Supertonic preprocessing to languages beyond English and Japanese?

Yes. The `UnicodeProcessor` validates language codes and wraps output in XML-style tags like `<en>` or `<ja>` (lines 102-105). You can extend support to additional languages by ensuring [`models/onnx/unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/models/onnx/unicode_indexer.json) contains the necessary code point mappings for your target language's character set, as the NFKD normalization and symbol cleaning handle Unicode generically.