Debugging Synthesis Errors and Common Issues in Supertonic

Supertonic synthesis errors typically stem from mismatched language codes, incorrect voice-style-to-text ratios, missing ONNX model assets, or disabled text chunking in batch mode, all of which can be diagnosed by tracing the pipeline from py/example_onnx.py through py/helper.py.

Supertonic is a multilingual on-device Text-to-Speech (TTS) system developed by supertone-inc that runs inference entirely through ONNX models. Understanding the synthesis pipeline architecture—spanning from the UnicodeProcessor text normalization to the four-model ONNX inference chain—is essential for rapidly isolating failure points. This guide walks through the most common runtime errors, their root causes in the source code, and practical debugging techniques.

Understanding the Supertonic Architecture

The synthesis flow is orchestrated across two primary Python files. In py/helper.py, the UnicodeProcessor class (lines 16-105) handles text normalization, stripping emojis, replacing symbols, and wrapping text in language tags like <en>…</en>. The TextToSpeech class (lines 40-155) manages the four ONNX sub-models: the duration predictor, text encoder, latent-vector estimator, and vocoder. It also handles chunking of long texts and speed scaling.

The command-line interface in py/example_onnx.py serves as the front-end, parsing arguments and invoking load_text_to_speech() to initialize the ONNX Runtime sessions. The inference chain flows as follows: argument parsing → model loading (load_onnx_all) → voice style loading (load_voice_style) → text preprocessing (UnicodeProcessor.__call__) → sub-model execution (dp_ort, text_enc_ort, vector_est_ort, vocoder_ort) → post-processing and WAV writing.

Common Synthesis Errors and Solutions

Invalid Language Codes

If you encounter ValueError: Invalid language: xx, the language code supplied via the --lang argument is not present in the AVAILABLE_LANGS list defined in py/helper.py. Verify your code against the supported ISO-639-1 codes in the source file, or update the list if you are adding support for a new language.

GPU Mode Limitations

Passing --use-gpu or setting use_gpu=True in load_text_to_speech() triggers a NotImplementedError: GPU mode is not fully tested. According to the source code in py/helper.py, GPU inference is deliberately guarded because it is not yet stable. To resolve this, run inference on the CPU by omitting the flag. If GPU support is mandatory, you must build a custom ONNX Runtime with GPU providers and remove the guard clause in the repository.

Voice-Style and Text Count Mismatches

The error AssertionError: Number of voice styles (…) must match number of texts (…) occurs when the --voice-style and --text arguments have differing lengths. Ensure a 1:1 correspondence between style JSONs and input strings:

--voice-style style1.json style2.json --text "Hello" "Hola"

Chunking Disabled in Batch Mode

When using --batch, Supertonic disables automatic text chunking, which can cause failures with long inputs. The chunk_text function in py/helper.py (lines 88-130) splits text into ≤300-character chunks, but this is bypassed in batch mode. For long-form synthesis, omit the --batch flag. If you require both batch processing and chunking, manually split texts using chunk_text() and invoke the TTS object iteratively.

Truncated or Silent Audio

Audio truncation usually indicates inaccurate duration predictions from the duration predictor model. Inspect the duration values returned by dp_ort and compare them against expected waveform lengths. You can mitigate this by adjusting the --speed factor or increasing --total-step to improve model stability during inference.

Missing ONNX Model Files

The error onnxruntime.capi.onnxruntime_pybind11_state.NotImplementedError: CPUExecutionProvider often masks a missing model file. Verify that --onnx-dir points to a directory containing all four required files: duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, and vocoder.onnx. The default value (../assets/onnx) assumes execution from within the py/ folder.

JSON and Unicode Decoding Errors

UnicodeDecodeError or json.JSONDecodeError typically indicate corrupted or missing asset files. Ensure unicode_indexer.json and voice-style JSONs are present in the assets directory. If damaged, re-download the assets from the Hugging Face release linked in the top-level README.

Debugging Techniques for ONNX Inference

To isolate failures, wrap the TextToSpeech.__call__ invocation in a try/except block and log the full traceback before the exception propagates.

Use the timer context manager available in the codebase to profile which sub-model introduces latency. For shape mismatches, print intermediate tensor dimensions immediately before feeding them to the ONNX sessions:

print(text_ids.shape, text_mask.shape)

Inspect the duration predictor outputs (print(duration)) to verify that the inferred lengths align with the actual phoneme sequences generated by the text encoder.

Code Examples for Troubleshooting

Basic Single-Text Synthesis

This command runs with default chunking enabled (recommended for long texts):

uv run py/example_onnx.py \
  --voice-style ../assets/voice_styles/M1.json \
  --text "The quick brown fox jumps over the lazy dog." \
  --lang en

Reproducing Language Code Errors

Test validation logic programmatically:

from helper import UnicodeProcessor

proc = UnicodeProcessor("../assets/unicode_indexer.json")
try:
    proc._preprocess_text("Hello world", "xx")
except ValueError as e:
    print("Caught error:", e)

Manual Chunking Implementation

When batch mode is required but inputs exceed length limits:

from helper import chunk_text, sanitize_filename

chunks = chunk_text(open("long_paragraph.txt").read())
for i, chunk in enumerate(chunks):
    # Process each chunk individually and concatenate results

    print(f"Processing chunk {i}: {chunk[:50]}...")

Verifying Model Loading

Confirm that ONNX sessions are initialized correctly:

from helper import load_text_to_speech

tts = load_text_to_speech("../assets/onnx", use_gpu=False)
print("Duration predictor inputs:", tts.dp_ort.get_inputs())
print("Vocoder outputs:", tts.vocoder_ort.get_outputs())

Summary

  • Language validation is strict; ensure codes exist in AVAILABLE_LANGS inside py/helper.py.
  • GPU inference is explicitly disabled in the current release due to stability concerns.
  • Batch mode disables automatic chunking, which may cause failures with texts longer than 300 characters.
  • Audio truncation can be mitigated by inspecting duration predictor outputs and adjusting speed parameters.
  • ONNX model paths must resolve to the four specific .onnx files in the assets directory.
  • Debug using try/except blocks around TextToSpeech.__call__, timer contexts, and intermediate tensor shape printing.

Frequently Asked Questions

Why does Supertonic throw "Invalid language" errors?

Supertonic validates language codes against an internal AVAILABLE_LANGS list in py/helper.py. If your ISO-639-1 code is not present, the UnicodeProcessor._preprocess_text() method raises a ValueError. Add your language to the list or use a supported code like en, ko, or ja.

How do I fix truncated or silent audio output?

Truncation occurs when the duration predictor underestimates phoneme lengths. Trace the duration variable output from the dp_ort session in py/helper.py. If values are too low, increase the --total-step parameter or reduce the --speed factor to allow the model more steps for generation.

Can I use GPU acceleration with Supertonic?

No. The load_text_to_speech() function explicitly raises a NotImplementedError when use_gpu=True. The repository maintainers have not fully tested GPU execution paths, so the code aborts to prevent undefined behavior. CPU execution is the stable path.

Why does batch processing disable automatic chunking?

The TextToSpeech class in py/helper.py disables chunking when the internal batch processing logic is active. This is a documented limitation because batch mode assumes you have pre-segmented your inputs. To process long texts, either remove --batch to enable automatic chunking, or manually split inputs using chunk_text() before passing them to the TTS engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →