Supertonic Text Preprocessing Pipeline: How UnicodeProcessor Prepares Text for TTS

The UnicodeProcessor class in Supertonic converts raw input strings into model-ready token IDs through a seven-step pipeline that includes text normalization, Unicode decomposition, and index mapping, implemented in cpp/helper.cpp.

The text preprocessing pipeline in the supertone-inc/supertonic repository transforms arbitrary Unicode strings into the tensor format required by the neural TTS models. This pipeline is orchestrated by the UnicodeProcessor class and operates entirely in C++ for performance, ensuring deterministic tokenization across 31 supported languages before any audio generation begins.

Overview of the Text Preprocessing Pipeline

The pipeline executes sequentially within the UnicodeProcessor::call method (lines 560–571 in cpp/helper.cpp). Each raw input string undergoes the following transformation stages:

  1. Entry Point: The TextToSpeech::_infer method receives batches of text and language codes (lines 878–886).
  2. Text Normalization: preprocessText applies symbol replacement, emoji stripping, and punctuation cleanup (lines 52–98, 104–161).
  3. Unicode Decomposition: Complex characters like Hangul syllables and Latin accented letters are decomposed into base code-points via textToUnicodeValues (lines 302–346).
  4. Index Mapping: Each Unicode value is mapped to an integer ID using the unicode_indexer.json lookup table (lines 668–686).
  5. Mask Generation: A binary mask is created via lengthToMask to mark valid token positions versus padding (lines 440–557).
  6. Output: The pipeline returns text_ids (2-D tensor) and text_mask (3-D binary tensor) for the ONNX models (lines 562–571).

The Text Normalization Stage

The preprocessText function (lines 168–182) performs language-agnostic string cleaning before Unicode conversion. According to the Supertonic source code, this stage executes ten specific operations:

  • Symbol replacement: Dashes, underscores, arrows, and quotes are converted to ASCII equivalents.
  • Emoji removal: A 4-byte UTF-8 regex ([\xF0][\x9F][\x80-\xBF][\x80-\xBF]) strips emoji sequences.
  • Special-symbol purge: Characters like “♥”, “☆”, and “©” are removed entirely.
  • Expression rewrites: Shorthands expand to full words (“@” → “ at ”, “e.g.,” → “for example,”, “i.e.,” → “that is,”).
  • Punctuation spacing: Spaces before commas, periods, and exclamation marks are normalized.
  • Duplicate-quote collapse: Sequences like "" become ", '' becomes ', and ```` becomes ```.
  • Whitespace normalization: Consecutive spaces collapse to single spaces, and leading/trailing whitespace is trimmed.
  • Sentence finalization: If the text lacks terminal punctuation (including multibyte symbols like “…” or “。”), a period is appended.
  • Language validation: The code must exist in AVAILABLE_LANGS; otherwise, a runtime error is thrown.
  • Tag wrapping: The cleaned text is wrapped in <lang>…</lang> tags to signal the downstream phoneme set.

Unicode Decomposition and Tokenization

After normalization, the textToUnicodeValues function (lines 302–346) converts the string into a vector of uint16_t code-points. This step uses the decomposeCharacter helper (lines 71–89) to handle complex scripts:

Hangul Decomposition: Korean syllables (U+AC00–U+D7A3) are split into constituent Jamo (leading consonant, vowel, and optional trailing consonant) using Unicode Standard Annex #15 algorithms.

Latin Character Expansion: The hard-coded LATIN_DECOMPOSITIONS map applies NFKD-style normalization, expanding characters like “Á” into “A” + “́” and “ç” into “c” + “̧”.

Pass-through: All other Unicode code-points remain unchanged, creating a canonical representation that the model's indexer expects.

From Unicode Values to Model Inputs

The final stages convert decomposed Unicode values into the integer tensors consumed by the ONNX models.

Index Mapping: Each uint16_t value is looked up in the unicode_indexer.json file (loaded at construction time) to produce the integer IDs required by the neural network (lines 668–686).

Mask Creation: The lengthToMask utility (lines 440–557) generates a binary text_mask from the sequence lengths, marking valid token positions as 1.0 and padding as 0.0. This allows the text encoder to ignore padded positions during inference (lines 488–494).

Implementation Examples

Using UnicodeProcessor Directly in C++

The following example demonstrates direct instantiation of the processor with the Unicode indexer:

#include "helper.h"

int main() {
    // Load the Unicode-to-ID indexer produced by the training pipeline
    auto processor = std::make_unique<UnicodeProcessor>("path/to/unicode_indexer.json");

    std::vector<std::string> texts = {"Hello 👋! こんにちは。", "¡Hola, mundo!"};
    std::vector<std::string> langs = {"en", "es"};

    std::vector<std::vector<int64_t>> text_ids;
    std::vector<std::vector<std::vector<float>>> text_mask;

    // Execute the full preprocessing pipeline
    processor->call(texts, langs, text_ids, text_mask);

    // text_ids now contains integer tokens for the ONNX models
    // text_mask marks valid positions versus padding
}

Python Wrapper Usage

The Python API mirrors the C++ implementation through a compiled extension:

from supertonic.py.helper import UnicodeProcessor

processor = UnicodeProcessor("unicode_indexer.json")

texts = ["Hello 👋! こんにちは。", "¡Hola, mundo!"]
langs = ["en", "es"]

text_ids, text_mask = processor.process(texts, langs)

Integration with TextToSpeech

For end-to-end inference, the TextToSpeech class handles the UnicodeProcessor invocation internally:

auto tts = loadTextToSpeech(env, onnx_dir, false);
auto result = tts->call(
    memory_info,           // Ort::MemoryInfo
    "Hello world! 🎉",     // Raw text input
    "en",                  // Language code
    style,                 // Pre-loaded speaker style
    50,                    // Total diffusion steps
    1.0f,                  // Speed factor
    0.2f);                 // Silence between chunks

The call method automatically triggers the UnicodeProcessor pipeline before feeding tensors to the duration predictor and text encoder.

Summary

  • The text preprocessing pipeline in Supertonic is implemented in cpp/helper.cpp and orchestrated by the UnicodeProcessor class.
  • Normalization occurs first via preprocessText, handling emoji removal, symbol replacement, and language validation before adding <lang> tags.
  • Unicode decomposition splits Hangul syllables into Jamo and expands Latin accented characters using NFKD-style tables in textToUnicodeValues.
  • Tokenization maps decomposed code-points to integer IDs using unicode_indexer.json, then generates binary masks for padding handling.
  • Cross-language support is built-in, with the pipeline handling 31 languages through the AVAILABLE_LANGS validation and language-specific tagging.

Frequently Asked Questions

What is the UnicodeProcessor in Supertonic?

The UnicodeProcessor is a C++ class defined in cpp/helper.h and implemented in cpp/helper.cpp that manages the entire text preprocessing pipeline for Supertonic's TTS models. It converts raw Unicode strings into integer token IDs and binary masks through normalization, decomposition, and index mapping stages, ensuring consistent input formatting for the neural networks.

How does Supertonic handle Unicode normalization?

Supertonic applies custom Unicode decomposition rather than standard library normalization. The decomposeCharacter function (lines 71–89) splits Hangul syllables (U+AC00–U+D7A3) into Jamo components and uses a hard-coded LATIN_DECOMPOSITIONS map to separate Latin accented characters into base letters and combining diacritics, producing a canonical form suitable for the model's vocabulary.

What languages does the Supertonic text preprocessing pipeline support?

The pipeline supports 31 languages as defined in the AVAILABLE_LANGS constant. During the preprocessText stage, the processor validates the provided language code against this list and wraps the normalized text in <lang>…</lang> tags, allowing the downstream phoneme converters to apply language-specific rules.

How does the pipeline handle emojis and special characters?

The preprocessText function removes emojis using a specific 4-byte UTF-8 regex pattern ([\xF0][\x9F][\x80-\xBF][\x80-\xBF]) and strips special symbols like “♥” and “©”. It also replaces typographic symbols (arrows, dashes, quotes) with ASCII equivalents, ensuring the final input contains only characters representable in the model's unicode indexer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →