How the UnicodeProcessor Class Preprocesses and Indexes Text for Supertonic's TTS Model

The UnicodeProcessor class in the supertone-inc/supertonic repository converts raw text into language-tagged integer sequences through Unicode normalization, regex-based cleaning, and code-point-to-ID mapping using a JSON indexer.

The UnicodeProcessor serves as the initial text processing stage in Supertonic's open-source text-to-speech pipeline. Located in the supertone-inc/supertonic repository, this class transforms raw input strings into clean, language-annotated sequences of integer IDs that feed directly into the ONNX text encoder. Understanding how the UnicodeProcessor class preprocesses and indexes text is essential for customizing the pipeline or debugging tokenization issues.

UnicodeProcessor Initialization and Indexer Loading

The processor begins by loading a deterministic mapping from Unicode code points to numeric token IDs. In py/helper.py, lines 16-20, the __init__ method loads the JSON indexer:

def __init__(self, unicode_indexer_path: str):
    with open(unicode_indexer_path, "r") as f:
        self.indexer = json.load(f)          # ← loads a dict  {code_point → int_id}

The unicode_indexer.json file is produced offline and contains every Unicode code point that the model can handle. This mapping ensures that the same character always resolves to the same integer ID, regardless of input source.

Text Preprocessing Pipeline

The _preprocess_text method in py/helper.py (lines 21-105) executes a ten-step normalization pipeline. Each step operates on Unicode code points to ensure language-agnostic handling:

  • Unicode normalization: Applies normalize("NFKD", text) to decompose characters (e.g., "é" → "e" + "´") into canonical forms.
  • Emoji removal: Strips emojis using regex covering \U0001f600-\U0001f6ff and symbols \u2600-\u26ff that fall outside the speech vocabulary.
  • Symbol replacement: Converts "–", "‑", "—", "_", quotes, brackets, and arrows to ASCII equivalents or spaces.
  • Decorative symbol stripping: Removes non-phonetic characters like ♥☆♡© via re.sub(r"[♥☆♡©\\]", "", text).
  • Expression expansion: Expands @ to " at ", e.g., to "for example, ", and i.e., to "that is, " for clearer lexical cues.
  • Punctuation spacing: Collapses spaces before commas and periods to guarantee tidy token boundaries.
  • Quote normalization: Collapses duplicate quotes ("", '', ``` ```) into single instances to prevent empty tokens.
  • Whitespace normalization: Uses re.sub(r"\s+", " ", text).strip() to collapse whitespace runs into single spaces.
  • Sentence finalization: Appends a period if the string lacks terminating punctuation, ensuring a closing token for the model.
  • Language tagging: Wraps text with <lang> tags (e.g., <en>Hello</en>) and validates supported language codes with ValueError for unsupported inputs.

All operations are purely text-based and execute before any numeric conversion occurs.

Converting Characters to Unicode Values

After cleaning, the _text_to_unicode_values method converts characters to their integer code points. Found in py/helper.py, lines 11-15:

def _text_to_unicode_values(self, text: str) -> np.ndarray:
    unicode_values = np.array([ord(char) for char in text], dtype=np.uint16)
    return unicode_values

The ord() function yields the integer code point for each character (e.g., "a" → 97). The resulting NumPy array uses 16-bit unsigned integers to match the JSON indexer keys.

Indexing the Processed Text

In the __call__ method (lines 24-30 in py/helper.py), the processor executes the final mapping:

  1. Calls _preprocess_text on each input string.
  2. Obtains the Unicode code point array via _text_to_unicode_values.
  3. Looks up each code point in self.indexer to obtain model-specific integer tokens.
for i, text in enumerate(text_list):
    unicode_vals = self._text_to_unicode_values(text)
    text_ids[i, : len(unicode_vals)] = np.array(
        [self.indexer[val] for val in unicode_vals], dtype=np.int64
    )

The resulting 2-D text_ids array has shape (batch, max_seq_len), padded with zeros where sequences are shorter than the longest input. A complementary binary mask (text_mask) is generated from length vectors via length_to_mask, allowing downstream ONNX modules to ignore padding during inference.

Cross-Platform Implementation Examples

The UnicodeProcessor architecture remains consistent across language bindings. Here are implementations in Python, Node.js, Java, and Go.

Python (Reference Implementation)

from py.helper import UnicodeProcessor

# Load the processor (the JSON indexer ships with the model bundle)

processor = UnicodeProcessor("models/unicode_indexer.json")

texts = [
    "Supertonic makes TTS easy 😊!",
    "여보세요, 이는 테스트입니다."
]
langs = ["en", "ko"]

ids, mask = processor(texts, langs)

print("Token IDs:\n", ids)
print("Mask:\n", mask)

Node.js

const { UnicodeProcessor } = require("./nodejs/helper.js");
const processor = new UnicodeProcessor("models/unicode_indexer.json");

const texts = ["Supertonic makes TTS easy 😊!", "여보세요, 이는 테스트입니다."];
const langs = ["en", "ko"];

const { textIds, textMask } = processor.call(texts, langs);
console.log(textIds);
console.log(textMask);

Source: nodejs/helper.js, lines 13-38

Java

UnicodeProcessor processor = new UnicodeProcessor("models/unicode_indexer.json");
List<String> texts = List.of("Supertonic makes TTS easy 😊!", "여보세요, 이는 테스트입니다.");
List<String> langs = List.of("en", "ko");
TextProcessResult result = processor.call(texts, langs);
int[][] ids = result.textIds;      // [batch][seqLen]
float[][][] mask = result.textMask; // [batch][1][seqLen]

Source: java/Helper.java, lines 67-82

Go

up, _ := NewUnicodeProcessor("models/unicode_indexer.json")
ids, mask := up.Call([]string{
    "Supertonic makes TTS easy 😊!",
    "여보세요, 이는 테스트입니다.",
}, []string{"en", "ko"})
fmt.Println(ids)  // [][]int64
fmt.Println(mask) // [][][]float64

Source: go/helper.go, lines 102-116

Summary

  • The UnicodeProcessor initializes by loading a JSON indexer from unicode_indexer.json that maps Unicode code points to integer IDs.
  • The _preprocess_text method normalizes input via NFKD decomposition, removes emojis and decorative symbols, expands abbreviations, and wraps text in language tags.
  • The _text_to_unicode_values method converts cleaned strings into 16-bit integer arrays using Python's ord() function.
  • During __call__, code points are mapped to model-specific token IDs, producing padded batches of shape (batch, max_seq_len) with accompanying binary masks.
  • Consistent implementations exist across Python, Node.js, Java, Go, C++, Swift, C#, and Dart to ensure platform-agnostic text processing.

Frequently Asked Questions

What is the purpose of the UnicodeProcessor in Supertonic's pipeline?

The UnicodeProcessor prepares raw strings for the ONNX text encoder by normalizing Unicode characters, removing unsupported symbols like emojis, and converting text into sequences of integer token IDs. It serves as the bridge between arbitrary user input and the fixed vocabulary expected by the neural network.

How does the UnicodeProcessor handle emojis and special characters?

The processor removes emojis using regex patterns covering the Unicode ranges \U0001f600-\U0001f6ff and \u2600-\u26ff, and strips decorative symbols like ♥☆♡© via character class substitution. This prevents out-of-vocabulary errors during inference, as these symbols have no phonetic representation in the model.

Why does the UnicodeProcessor use NFKD normalization?

NFKD (Normalization Form KD) decomposes characters into their canonical constituents, converting characters like "é" into "e" plus a combining acute accent. This ensures the model processes a consistent, decomposed representation of text regardless of how the input was composed, reducing vocabulary fragmentation.

What is the output format of the UnicodeProcessor?

The processor returns two arrays: text_ids, a 2-D integer array of shape (batch, max_seq_len) containing padded token IDs, and text_mask, a binary mask indicating which positions contain valid tokens versus padding. These feed directly into the duration predictor and text encoder ONNX models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →