# Supertonic Text Preprocessing Pipeline: How UnicodeProcessor Prepares Text for TTS

> Explore the Supertonic text preprocessing pipeline and understand how UnicodeProcessor prepares text for TTS. Learn about normalization, decomposition, and index mapping for model-ready tokens.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: deep-dive
- Published: 2026-06-12

---

**The `UnicodeProcessor` class in Supertonic converts raw input strings into model-ready token IDs through a seven-step pipeline that includes text normalization, Unicode decomposition, and index mapping, implemented in [`cpp/helper.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/helper.cpp).**

The text preprocessing pipeline in the [supertone-inc/supertonic](https://github.com/supertone-inc/supertonic) repository transforms arbitrary Unicode strings into the tensor format required by the neural TTS models. This pipeline is orchestrated by the `UnicodeProcessor` class and operates entirely in C++ for performance, ensuring deterministic tokenization across 31 supported languages before any audio generation begins.

## Overview of the Text Preprocessing Pipeline

The pipeline executes sequentially within the `UnicodeProcessor::call` method (lines 560–571 in [`cpp/helper.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/helper.cpp)). Each raw input string undergoes the following transformation stages:

1. **Entry Point**: The `TextToSpeech::_infer` method receives batches of text and language codes (lines 878–886).
2. **Text Normalization**: `preprocessText` applies symbol replacement, emoji stripping, and punctuation cleanup (lines 52–98, 104–161).
3. **Unicode Decomposition**: Complex characters like Hangul syllables and Latin accented letters are decomposed into base code-points via `textToUnicodeValues` (lines 302–346).
4. **Index Mapping**: Each Unicode value is mapped to an integer ID using the [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) lookup table (lines 668–686).
5. **Mask Generation**: A binary mask is created via `lengthToMask` to mark valid token positions versus padding (lines 440–557).
6. **Output**: The pipeline returns `text_ids` (2-D tensor) and `text_mask` (3-D binary tensor) for the ONNX models (lines 562–571).

## The Text Normalization Stage

The `preprocessText` function (lines 168–182) performs language-agnostic string cleaning before Unicode conversion. According to the Supertonic source code, this stage executes ten specific operations:

- **Symbol replacement**: Dashes, underscores, arrows, and quotes are converted to ASCII equivalents.
- **Emoji removal**: A 4-byte UTF-8 regex (`[\xF0][\x9F][\x80-\xBF][\x80-\xBF]`) strips emoji sequences.
- **Special-symbol purge**: Characters like “♥”, “☆”, and “©” are removed entirely.
- **Expression rewrites**: Shorthands expand to full words (“@” → “ at ”, “e.g.,” → “for example,”, “i.e.,” → “that is,”).
- **Punctuation spacing**: Spaces before commas, periods, and exclamation marks are normalized.
- **Duplicate-quote collapse**: Sequences like `""` become `"`, `''` becomes `'`, and ```` becomes ```.
- **Whitespace normalization**: Consecutive spaces collapse to single spaces, and leading/trailing whitespace is trimmed.
- **Sentence finalization**: If the text lacks terminal punctuation (including multibyte symbols like “…” or “。”), a period is appended.
- **Language validation**: The code must exist in `AVAILABLE_LANGS`; otherwise, a runtime error is thrown.
- **Tag wrapping**: The cleaned text is wrapped in `<lang>…</lang>` tags to signal the downstream phoneme set.

## Unicode Decomposition and Tokenization

After normalization, the `textToUnicodeValues` function (lines 302–346) converts the string into a vector of `uint16_t` code-points. This step uses the `decomposeCharacter` helper (lines 71–89) to handle complex scripts:

**Hangul Decomposition**: Korean syllables (U+AC00–U+D7A3) are split into constituent Jamo (leading consonant, vowel, and optional trailing consonant) using Unicode Standard Annex #15 algorithms.

**Latin Character Expansion**: The hard-coded `LATIN_DECOMPOSITIONS` map applies NFKD-style normalization, expanding characters like “Á” into “A” + “́” and “ç” into “c” + “̧”.

**Pass-through**: All other Unicode code-points remain unchanged, creating a canonical representation that the model's indexer expects.

## From Unicode Values to Model Inputs

The final stages convert decomposed Unicode values into the integer tensors consumed by the ONNX models.

**Index Mapping**: Each `uint16_t` value is looked up in the [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) file (loaded at construction time) to produce the integer IDs required by the neural network (lines 668–686).

**Mask Creation**: The `lengthToMask` utility (lines 440–557) generates a binary `text_mask` from the sequence lengths, marking valid token positions as `1.0` and padding as `0.0`. This allows the text encoder to ignore padded positions during inference (lines 488–494).

## Implementation Examples

### Using UnicodeProcessor Directly in C++

The following example demonstrates direct instantiation of the processor with the Unicode indexer:

```cpp
#include "helper.h"

int main() {
    // Load the Unicode-to-ID indexer produced by the training pipeline
    auto processor = std::make_unique<UnicodeProcessor>("path/to/unicode_indexer.json");

    std::vector<std::string> texts = {"Hello 👋! こんにちは。", "¡Hola, mundo!"};
    std::vector<std::string> langs = {"en", "es"};

    std::vector<std::vector<int64_t>> text_ids;
    std::vector<std::vector<std::vector<float>>> text_mask;

    // Execute the full preprocessing pipeline
    processor->call(texts, langs, text_ids, text_mask);

    // text_ids now contains integer tokens for the ONNX models
    // text_mask marks valid positions versus padding
}

```

### Python Wrapper Usage

The Python API mirrors the C++ implementation through a compiled extension:

```python
from supertonic.py.helper import UnicodeProcessor

processor = UnicodeProcessor("unicode_indexer.json")

texts = ["Hello 👋! こんにちは。", "¡Hola, mundo!"]
langs = ["en", "es"]

text_ids, text_mask = processor.process(texts, langs)

```

### Integration with TextToSpeech

For end-to-end inference, the `TextToSpeech` class handles the UnicodeProcessor invocation internally:

```cpp
auto tts = loadTextToSpeech(env, onnx_dir, false);
auto result = tts->call(
    memory_info,           // Ort::MemoryInfo
    "Hello world! 🎉",     // Raw text input
    "en",                  // Language code
    style,                 // Pre-loaded speaker style
    50,                    // Total diffusion steps
    1.0f,                  // Speed factor
    0.2f);                 // Silence between chunks

```

The `call` method automatically triggers the `UnicodeProcessor` pipeline before feeding tensors to the duration predictor and text encoder.

## Summary

- **The text preprocessing pipeline** in Supertonic is implemented in [`cpp/helper.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/helper.cpp) and orchestrated by the `UnicodeProcessor` class.
- **Normalization occurs first** via `preprocessText`, handling emoji removal, symbol replacement, and language validation before adding `<lang>` tags.
- **Unicode decomposition** splits Hangul syllables into Jamo and expands Latin accented characters using NFKD-style tables in `textToUnicodeValues`.
- **Tokenization** maps decomposed code-points to integer IDs using [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json), then generates binary masks for padding handling.
- **Cross-language support** is built-in, with the pipeline handling 31 languages through the `AVAILABLE_LANGS` validation and language-specific tagging.

## Frequently Asked Questions

### What is the UnicodeProcessor in Supertonic?

The `UnicodeProcessor` is a C++ class defined in [`cpp/helper.h`](https://github.com/supertone-inc/supertonic/blob/main/cpp/helper.h) and implemented in [`cpp/helper.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/helper.cpp) that manages the entire text preprocessing pipeline for Supertonic's TTS models. It converts raw Unicode strings into integer token IDs and binary masks through normalization, decomposition, and index mapping stages, ensuring consistent input formatting for the neural networks.

### How does Supertonic handle Unicode normalization?

Supertonic applies custom Unicode decomposition rather than standard library normalization. The `decomposeCharacter` function (lines 71–89) splits Hangul syllables (U+AC00–U+D7A3) into Jamo components and uses a hard-coded `LATIN_DECOMPOSITIONS` map to separate Latin accented characters into base letters and combining diacritics, producing a canonical form suitable for the model's vocabulary.

### What languages does the Supertonic text preprocessing pipeline support?

The pipeline supports 31 languages as defined in the `AVAILABLE_LANGS` constant. During the `preprocessText` stage, the processor validates the provided language code against this list and wraps the normalized text in `<lang>…</lang>` tags, allowing the downstream phoneme converters to apply language-specific rules.

### How does the pipeline handle emojis and special characters?

The `preprocessText` function removes emojis using a specific 4-byte UTF-8 regex pattern (`[\xF0][\x9F][\x80-\xBF][\x80-\xBF]`) and strips special symbols like “♥” and “©”. It also replaces typographic symbols (arrows, dashes, quotes) with ASCII equivalents, ensuring the final input contains only characters representable in the model's unicode indexer.