How Supertonic Handles Financial Expressions, Phone Numbers, and Technical Units

Supertonic's on-device TTS pipeline includes a built-in text normalizer that automatically expands monetary amounts, telephone numbers, and engineering units into spoken form without requiring pre-processing or phonetic annotations.

The supertone-inc/supertonic repository provides an on-device text-to-speech engine that processes complex real-world strings through a built-in text normalizer. This normalizer handles financial expressions, phone numbers, and technical units using regex-based rules embedded in language-specific helper modules, ensuring consistent pronunciation across all supported runtimes.

Built-in Text Normalization Architecture

The normalizer operates as the first stage in the inference pipeline. According to the source code, raw input strings pass through a series of regex-based rewrite rules before tokenization. These rules detect patterns for currency symbols, telephone number formats, and unit abbreviations, expanding them into full spoken words that the 99M-parameter ONNX model can pronounce naturally.

Financial Expression Handling

The normalizer detects currency symbols ($), decimal points, and magnitude abbreviations (K, M). In README.md (lines 346-360), the documentation shows that $5.2M expands to "five point two million dollars" and $450K becomes "four hundred fifty thousand dollars". The implementation uses regular-expression patterns embedded in the core inference code to capture these financial expressions and convert them to their spoken equivalents.

Phone Number Handling

For telephone numbers, the system recognizes area-code parentheses, hyphen separators, and the "ext." abbreviation. As documented in README.md (lines 375-388), the text (212) 555-0142 ext. 402 normalizes to "two one two, five five five, zero one four two, extension four zero two". The phone-number rule set tokenizes each digit individually while preserving the extension marker.

Technical Unit Handling

Technical units like 2.3h or 30kph are parsed by recognizing decimal numbers paired with unit abbreviations. The unit-recognition rule, found in README.md (lines 401-414), maps 2.3h to "two point three hours" and 30kph to "thirty kilometres per hour". This parsing handles the decimal point and unit suffix separately before combining them into natural speech.

Cross-Platform Implementation

The normalizer lives in each SDK's language-specific helper module, ensuring identical behavior across all supported runtimes. The Python implementation resides in py/helper.py, utilizing unicodedata.normalize combined with regex rules. The JavaScript version in nodejs/helper.js uses text.normalize('NFKD'), while java/Helper.java employs Normalizer.normalize. Similar stubs exist in go/helper.go, cpp/helper.cpp, and rust/src/helper.rs, maintaining the same text-normalization logic across Python, Node.js, Java, C++, C#, Go, and Swift.

Practical Usage Examples

The following Python SDK demonstrates how the normalizer processes each category automatically:

from supertonic import TTS

tts = TTS(auto_download=True)

# Financial expression

financial_text = "The startup secured $5.2M in venture capital, a huge leap from their initial $450K seed round."
wav, _ = tts.synthesize(text=financial_text, lang="en")
tts.save_audio(wav, "financial.wav")

# Phone number

phone_text = "You can reach the hotel front desk at (212) 555-0142 ext. 402 anytime."
wav, _ = tts.synthesize(text=phone_text, lang="en")
tts.save_audio(wav, "phone.wav")

# Technical unit

unit_text = "Our drone battery lasts 2.3h when flying at 30kph with full camera payload."
wav, _ = tts.synthesize(text=unit_text, lang="en")
tts.save_audio(wav, "unit.wav")

Running this script produces three WAV files containing correctly pronounced expansions of the financial, telephone, and technical unit expressions.

Summary

  • Supertonic's built-in text normalizer automatically handles financial expressions, phone numbers, and technical units without external preprocessing.
  • The system uses regex-based rules embedded in language-specific helper files like py/helper.py and nodejs/helper.js.
  • Financial expressions such as $5.2M expand to "five point two million dollars" using magnitude abbreviation detection.
  • Phone numbers including extensions like (212) 555-0142 ext. 402 are tokenized into individual digits with proper pauses.
  • Technical units like 30kph normalize to "thirty kilometres per hour" through combined number and suffix parsing.
  • All processing occurs on-device across CPU, GPU, WebGPU, or mobile runtimes without requiring network services.

Frequently Asked Questions

Does Supertonic require phonetic annotations for financial expressions?

No. According to the repository's README (lines 442-449), the system handles text normalization automatically, eliminating the need for pre-processing or phonetic annotations. You can input raw strings like $5.2M directly into the synthesize() method.

Which source files contain the text normalization logic?

The normalization logic resides in the helper modules for each language binding. Key files include py/helper.py for Python, nodejs/helper.js for JavaScript, and java/Helper.java for Java. These files implement the regex patterns that detect and expand financial expressions, phone numbers, and technical units.

Are phone number formats consistent across all programming languages?

Yes. Because the normalizer is implemented in each SDK's helper module with identical regex rules, phone number handling produces the same spoken output whether you use Python, Node.js, Java, C++, C#, Go, or Swift. The (212) 555-0142 ext. 402 example will always expand to "two one two, five five five, zero one four two, extension four zero two".

Can the normalizer handle combined expressions in a single sentence?

Yes. The normalizer processes the entire input text sequentially, applying rules for financial expressions, phone numbers, and technical units within the same string. For example, "Call (212) 555-0142 for $5.2M deals" correctly normalizes both the phone number and monetary amount before tokenization for the 99M-parameter ONNX model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →