# How the UnicodeProcessor Class Preprocesses and Indexes Text for Supertonic's TTS Model

> Learn how the UnicodeProcessor class preprocesses and indexes text for Supertonic's TTS model. Discover Unicode normalization, regex cleaning, and code-point mapping for efficient text processing.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: internals
- Published: 2026-06-14

---

**The UnicodeProcessor class in the supertone-inc/supertonic repository converts raw text into language-tagged integer sequences through Unicode normalization, regex-based cleaning, and code-point-to-ID mapping using a JSON indexer.**

The UnicodeProcessor serves as the initial text processing stage in Supertonic's open-source text-to-speech pipeline. Located in the `supertone-inc/supertonic` repository, this class transforms raw input strings into clean, language-annotated sequences of integer IDs that feed directly into the ONNX text encoder. Understanding how the UnicodeProcessor class preprocesses and indexes text is essential for customizing the pipeline or debugging tokenization issues.

## UnicodeProcessor Initialization and Indexer Loading

The processor begins by loading a deterministic mapping from Unicode code points to numeric token IDs. In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), lines 16-20, the `__init__` method loads the JSON indexer:

```python
def __init__(self, unicode_indexer_path: str):
    with open(unicode_indexer_path, "r") as f:
        self.indexer = json.load(f)          # ← loads a dict  {code_point → int_id}

```

The [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) file is produced offline and contains every Unicode code point that the model can handle. This mapping ensures that the same character always resolves to the same integer ID, regardless of input source.

## Text Preprocessing Pipeline

The `_preprocess_text` method in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) (lines 21-105) executes a ten-step normalization pipeline. Each step operates on Unicode code points to ensure language-agnostic handling:

- **Unicode normalization**: Applies `normalize("NFKD", text)` to decompose characters (e.g., "é" → "e" + "´") into canonical forms.
- **Emoji removal**: Strips emojis using regex covering `\U0001f600-\U0001f6ff` and symbols `\u2600-\u26ff` that fall outside the speech vocabulary.
- **Symbol replacement**: Converts "–", "‑", "—", "_", quotes, brackets, and arrows to ASCII equivalents or spaces.
- **Decorative symbol stripping**: Removes non-phonetic characters like `♥☆♡©` via `re.sub(r"[♥☆♡©\\]", "", text)`.
- **Expression expansion**: Expands `@` to " at ", `e.g.,` to "for example, ", and `i.e.,` to "that is, " for clearer lexical cues.
- **Punctuation spacing**: Collapses spaces before commas and periods to guarantee tidy token boundaries.
- **Quote normalization**: Collapses duplicate quotes (`""`, `''`, ` ``` ``` `) into single instances to prevent empty tokens.
- **Whitespace normalization**: Uses `re.sub(r"\s+", " ", text).strip()` to collapse whitespace runs into single spaces.
- **Sentence finalization**: Appends a period if the string lacks terminating punctuation, ensuring a closing token for the model.
- **Language tagging**: Wraps text with `<lang>` tags (e.g., `<en>Hello</en>`) and validates supported language codes with `ValueError` for unsupported inputs.

All operations are purely text-based and execute before any numeric conversion occurs.

## Converting Characters to Unicode Values

After cleaning, the `_text_to_unicode_values` method converts characters to their integer code points. Found in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), lines 11-15:

```python
def _text_to_unicode_values(self, text: str) -> np.ndarray:
    unicode_values = np.array([ord(char) for char in text], dtype=np.uint16)
    return unicode_values

```

The `ord()` function yields the integer code point for each character (e.g., "a" → 97). The resulting NumPy array uses 16-bit unsigned integers to match the JSON indexer keys.

## Indexing the Processed Text

In the `__call__` method (lines 24-30 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)), the processor executes the final mapping:

1. Calls `_preprocess_text` on each input string.
2. Obtains the Unicode code point array via `_text_to_unicode_values`.
3. Looks up each code point in `self.indexer` to obtain model-specific integer tokens.

```python
for i, text in enumerate(text_list):
    unicode_vals = self._text_to_unicode_values(text)
    text_ids[i, : len(unicode_vals)] = np.array(
        [self.indexer[val] for val in unicode_vals], dtype=np.int64
    )

```

The resulting 2-D `text_ids` array has shape **(batch, max_seq_len)**, padded with zeros where sequences are shorter than the longest input. A complementary binary mask (`text_mask`) is generated from length vectors via `length_to_mask`, allowing downstream ONNX modules to ignore padding during inference.

## Cross-Platform Implementation Examples

The UnicodeProcessor architecture remains consistent across language bindings. Here are implementations in Python, Node.js, Java, and Go.

### Python (Reference Implementation)

```python
from py.helper import UnicodeProcessor

# Load the processor (the JSON indexer ships with the model bundle)

processor = UnicodeProcessor("models/unicode_indexer.json")

texts = [
    "Supertonic makes TTS easy 😊!",
    "여보세요, 이는 테스트입니다."
]
langs = ["en", "ko"]

ids, mask = processor(texts, langs)

print("Token IDs:\n", ids)
print("Mask:\n", mask)

```

### Node.js

```javascript
const { UnicodeProcessor } = require("./nodejs/helper.js");
const processor = new UnicodeProcessor("models/unicode_indexer.json");

const texts = ["Supertonic makes TTS easy 😊!", "여보세요, 이는 테스트입니다."];
const langs = ["en", "ko"];

const { textIds, textMask } = processor.call(texts, langs);
console.log(textIds);
console.log(textMask);

```

*Source:* [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js), lines 13-38

### Java

```java
UnicodeProcessor processor = new UnicodeProcessor("models/unicode_indexer.json");
List<String> texts = List.of("Supertonic makes TTS easy 😊!", "여보세요, 이는 테스트입니다.");
List<String> langs = List.of("en", "ko");
TextProcessResult result = processor.call(texts, langs);
int[][] ids = result.textIds;      // [batch][seqLen]
float[][][] mask = result.textMask; // [batch][1][seqLen]

```

*Source:* [`java/Helper.java`](https://github.com/supertone-inc/supertonic/blob/main/java/Helper.java), lines 67-82

### Go

```go
up, _ := NewUnicodeProcessor("models/unicode_indexer.json")
ids, mask := up.Call([]string{
    "Supertonic makes TTS easy 😊!",
    "여보세요, 이는 테스트입니다.",
}, []string{"en", "ko"})
fmt.Println(ids)  // [][]int64
fmt.Println(mask) // [][][]float64

```

*Source:* [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go), lines 102-116

## Summary

- The UnicodeProcessor initializes by loading a JSON indexer from [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) that maps Unicode code points to integer IDs.
- The `_preprocess_text` method normalizes input via NFKD decomposition, removes emojis and decorative symbols, expands abbreviations, and wraps text in language tags.
- The `_text_to_unicode_values` method converts cleaned strings into 16-bit integer arrays using Python's `ord()` function.
- During `__call__`, code points are mapped to model-specific token IDs, producing padded batches of shape **(batch, max_seq_len)** with accompanying binary masks.
- Consistent implementations exist across Python, Node.js, Java, Go, C++, Swift, C#, and Dart to ensure platform-agnostic text processing.

## Frequently Asked Questions

### What is the purpose of the UnicodeProcessor in Supertonic's pipeline?

The UnicodeProcessor prepares raw strings for the ONNX text encoder by normalizing Unicode characters, removing unsupported symbols like emojis, and converting text into sequences of integer token IDs. It serves as the bridge between arbitrary user input and the fixed vocabulary expected by the neural network.

### How does the UnicodeProcessor handle emojis and special characters?

The processor removes emojis using regex patterns covering the Unicode ranges `\U0001f600-\U0001f6ff` and `\u2600-\u26ff`, and strips decorative symbols like `♥☆♡©` via character class substitution. This prevents out-of-vocabulary errors during inference, as these symbols have no phonetic representation in the model.

### Why does the UnicodeProcessor use NFKD normalization?

NFKD (Normalization Form KD) decomposes characters into their canonical constituents, converting characters like "é" into "e" plus a combining acute accent. This ensures the model processes a consistent, decomposed representation of text regardless of how the input was composed, reducing vocabulary fragmentation.

### What is the output format of the UnicodeProcessor?

The processor returns two arrays: `text_ids`, a 2-D integer array of shape **(batch, max_seq_len)** containing padded token IDs, and `text_mask`, a binary mask indicating which positions contain valid tokens versus padding. These feed directly into the duration predictor and text encoder ONNX models.