How Supertonic Handles Language-Agnostic Synthesis with `lang="na"`

Supertonic processes language-agnostic synthesis requests by validating lang="na" against an allow-list, wrapping input text in <na>...</na> XML tags, and routing the encoded tokens through the same multilingual ONNX inference graph used for all 31 supported languages.

The supertone-inc/supertonic repository provides a text-to-speech pipeline designed for 31 distinct languages alongside a special language-agnostic mode. When developers specify lang="na" (short for not-applicable), the system leverages its unified multilingual architecture to synthesize speech without requiring language-specific adapters or separate model weights.

How Language-Agnostic Synthesis Works

When lang="na" is supplied to the TextToSpeech.__call__ method, Supertonic executes a four-stage pipeline that treats the input as generic multilingual text rather than targeting a specific language.

Language Validation in py/helper.py

The system first validates the requested language code against the AVAILABLE_LANGS tuple defined in py/helper.py at line 13. This list contains all 31 supported ISO codes plus the string "na", ensuring that language-agnostic requests are treated as first-class citizens alongside specific languages like "en" or "ko".

XML Tagging via UnicodeProcessor._preprocess_text

Once validated, the input text undergoes preprocessing in the UnicodeProcessor._preprocess_text method (lines 102-105 in py/helper.py). This method wraps the cleaned text in XML-like language tags. For lang="na", the text becomes:

<na>Your input text here</na>

This tagging scheme allows the downstream neural models to recognize that the content should be processed using the model's internal multilingual representations rather than language-specific normalization rules.

Unified ONNX Inference Pipeline

The tagged text is tokenized into Unicode IDs and fed through the text encoder ONNX model, followed by the duration predictor and vector estimator. Crucially, the inference code in py/example_onnx.py (lines 100-108) does not branch to a different sub-graph when processing "na"; it uses the exact same multilingual ONNX sub-graphs employed for all other languages. The models were trained on a corpus that explicitly included the <na> tag, enabling the system to infer pronunciation and prosody internally without external language adapters.

Implementation Examples

Python SDK Usage

When using the high-level Python SDK, pass lang="na" to the synthesize method:

from supertonic import TTS

# Load the model (auto-download on first run)

tts = TTS(auto_download=True)

# Get a voice style (e.g., the default "M1" style)

style = tts.get_voice_style(voice_name="M1")

# Language-agnostic synthesis – note `lang="na"`

wav, duration = tts.synthesize(
    text="¡Hola! This is a mixed-language test.",
    lang="na",               # ⬅️ language-agnostic

    voice_style=style,
    total_steps=8,
    speed=1.05,
)

# Save the audio

tts.save_audio(wav, "na_demo.wav")

The lang="na" argument propagates from TTS.synthesize through TextToSpeech.__call__ and ultimately reaches UnicodeProcessor for preprocessing.

CLI Execution with example_onnx.py

For command-line inference, use the provided example script:

python py/example_onnx.py \
  --onnx-dir ../assets/onnx \
  --text "Bonjour, 世界! This sentence mixes languages." \
  --lang na \
  --voice-style ../assets/voice_styles/M1.json \
  --save-dir na_results

The script parses the --lang argument (lines 60-64) and invokes text_to_speech(text, lang, style, ...) using the same preprocessing pipeline described above.

Summary

  • lang="na" validation: Permitted via the AVAILABLE_LANGS tuple in py/helper.py (line 13), treating "not-applicable" as a valid language code.
  • Tagging mechanism: UnicodeProcessor._preprocess_text wraps input in <na>...</na> tags (lines 102-105) to signal multilingual processing.
  • Unified inference: The same ONNX text encoder, duration predictor, and vector estimator process "na" requests without language-specific adapters.
  • Training data: The models learned multilingual representations from a corpus containing <na> tagged examples, enabling internal language detection.
  • API consistency: Both the Python SDK and CLI handle lang="na" identically to specific language codes, requiring no special configuration flags.

Frequently Asked Questions

What does lang="na" stand for in Supertonic?

lang="na" stands for not-applicable. It is included in the AVAILABLE_LANGS list within py/helper.py as a special identifier indicating that the input text contains mixed languages or unknown language content that should be handled by the model's multilingual capabilities rather than a specific language adapter.

How does Supertonic handle text preprocessing for language-agnostic mode?

The UnicodeProcessor._preprocess_text method in py/helper.py (lines 102-105) validates the language code and wraps the input text in XML-style tags. For lang="na", this produces <na>input text</na>, which signals the ONNX inference pipeline to use generalized multilingual representations instead of language-specific normalization.

Does using lang="na" require loading different ONNX models?

No. According to the README and implementation in py/example_onnx.py, Supertonic uses the same ONNX sub-graphs for every language including "na". The text encoder, duration predictor, and vector estimator remain constant; the model relies on its training with <na> tags to infer language properties internally, eliminating the need for separate language-specific adapters.

Where is the language validation logic implemented in the source code?

Language validation occurs in py/helper.py at line 13, where the AVAILABLE_LANGS tuple defines all permissible codes. When the TextToSpeech class or CLI script processes a request, it checks against this list before calling UnicodeProcessor._preprocess_text, ensuring only supported languages (including "na") reach the inference stage.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →