How Supertonic's Text Normalization Handles Special Characters and Unicode

Supertonic normalizes incoming text across all supported language bindings by applying Unicode NFKD (Normalization Form Compatibility Decomposition) to decompose characters into base glyphs and combining marks, then strips the diacritical marks to produce clean ASCII output for the speech synthesis engine.

The supertone-inc/supertonic repository implements Supertonic's text normalization pipeline across Python, JavaScript, Java, C++, and Dart to preprocess TTS input. By applying Unicode NFKD decomposition according to the source files in the repository, the library ensures that special characters and Unicode glyphs never trigger out-of-vocabulary errors in the neural model.

The NFKD Decomposition Strategy

Supertonic's text normalization relies on Unicode NFKD (Normalization Form Compatibility Decomposition). This process splits precomposed characters—such as "é" or "ñ"—into a base ASCII letter followed by separate combining diacritical marks.

The pipeline executes two distinct steps:

  1. Decomposition: Convert characters to their canonical base form plus combining marks using NFKD.
  2. Stripping: Remove all combining diacritical marks (Unicode category Mn), leaving only the base ASCII characters.

This approach transforms strings like "Café" into "Cafe" and "naïve" into "naive", ensuring the downstream TTS model receives predictable, ASCII-only input.

Language-Specific Implementation Details

Each language binding in the Supertonic repository implements this normalization strategy using native libraries or optimized lookup tables.

Python Implementation

In py/helper.py at line 6, the Python binding calls the standard library's unicodedata.normalize function:

text = normalize("NFKD", text)

After decomposition, the code filters out combining marks to generate the final ASCII string.

JavaScript and Node.js

The Node.js implementation in nodejs/helper.js at line 20 leverages ECMAScript 6's built-in String.prototype.normalize method:

text = text.normalize('NFKD');

This provides native NFKD support without external dependencies.

Java

Java developers will find the normalization logic in java/Helper.java at line 143, utilizing java.text.Normalizer:

text = Normalizer.normalize(text, Normalizer.Form.NFKD);

The implementation follows the same decomposition-then-strip pattern as the other bindings.

C++

The C++ implementation in cpp/helper.cpp starting at line 53 takes a different approach for performance. Rather than linking the full ICU library, it uses a manual lookup table that maps pre-composed characters to their base-plus-combining sequences. The surrounding code then strips the combining marks, effectively mirroring NFKD behavior with minimal overhead.

Dart (Flutter)

For mobile platforms, flutter/lib/helper.dart at line 55 contains a static table of code-point replacements (e.g., mapping 0x00C1 to [0x0041, 0x0301]). This table decomposes accented characters into base letters and diacritics, which are subsequently removed to ensure platform-independent normalization.

Why Supertonic Uses NFKD

The choice of NFKD over other normalization forms provides three critical advantages for speech synthesis:

  • Compatibility: NFKD maps legacy characters, ligatures, and superscripts to their canonical equivalents, preventing vocabulary mismatches.
  • Predictability: The TTS model was trained on normalized, ASCII-only transcripts; NFKD ensures input consistency with the training data.
  • Speed: Normalization operates as a single-pass transformation in standard libraries, or via pre-computed tables in C++ and Dart, adding negligible latency to the inference pipeline.

Practical Usage Examples

Here is how Supertonic's text normalization handles complex input across different languages:

Python:

from supertonic import Supertonic
engine = Supertonic()
text = "Café 🎉 — résumé"
clean = engine.normalize_text(text)   # → "Cafe  - resume"

Node.js:

const { Supertonic } = require('supertonic');
const st = new Supertonic();
const text = "¡Hola, Señor!";
const clean = st.normalizeText(text); // "Hola, Senor!"

Java:

Supertonic st = new Supertonic();
String text = "naïve façade";
String clean = st.normalizeText(text); // "naive facade"

C++:

Supertonic st;
std::string text = u8"Łódź – piękny";
std::string clean = st.normalizeText(text); // "Lodz - piekny"

Dart (Flutter):

final st = Supertonic();
final clean = st.normalizeText('Résumé 🍀'); // "Resume "

Summary

  • Supertonic applies Unicode NFKD normalization across all five language bindings to decompose special characters into base ASCII letters and combining marks.
  • After decomposition, the pipeline strips combining diacritical marks (Unicode category Mn) to produce clean ASCII output.
  • Implementation varies by language: Python uses unicodedata.normalize in py/helper.py, Node.js uses String.normalize in nodejs/helper.js, Java uses java.text.Normalizer in java/Helper.java, while C++ (cpp/helper.cpp) and Dart (flutter/lib/helper.dart) employ optimized lookup tables for performance.
  • This normalization prevents out-of-vocabulary errors by ensuring the TTS engine receives consistent, ASCII-compatible text regardless of input complexity.

Frequently Asked Questions

What Unicode normalization form does Supertonic use?

Supertonic uses NFKD (Normalization Form Compatibility Decomposition). This form decomposes characters into their base code points plus combining diacritical marks, which are subsequently removed to generate ASCII-compatible text for the speech synthesis model.

How does Supertonic handle emojis and special symbols?

Emojis and special symbols that do not decompose into ASCII base characters are typically removed or replaced with whitespace during the stripping phase. The normalization process focuses on producing ASCII-only output, so non-Latin characters and symbols are filtered out after NFKD decomposition to prevent vocabulary mismatches in the TTS engine.

Is the normalization behavior consistent across all Supertonic language bindings?

Yes. While implementation details differ—Python and Node.js use built-in normalization libraries, Java uses java.text.Normalizer, and C++ and Dart use manual lookup tables—the behavioral output is identical. All implementations apply NFKD decomposition followed by removal of combining marks (Unicode category Mn) to ensure cross-platform consistency.

Why does Supertonic convert text to ASCII instead of keeping Unicode characters?

The TTS model was trained on normalized, ASCII-only transcripts. Converting input to ASCII guarantees that out-of-vocabulary tokens never reach the model, eliminating pronunciation errors that could occur when the engine encounters unfamiliar Unicode glyphs. This approach also simplifies vocabulary management and reduces model complexity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →