# Debugging Synthesis Errors and Common Issues in Supertonic

> Troubleshoot Supertonic synthesis errors by diagnosing common issues like mismatched language codes, incorrect voice ratios, and missing ONNX assets. Trace the pipeline to resolve problems effectively.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-05-14

---

**Supertonic synthesis errors typically stem from mismatched language codes, incorrect voice-style-to-text ratios, missing ONNX model assets, or disabled text chunking in batch mode, all of which can be diagnosed by tracing the pipeline from** [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) **through** [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)**.**

Supertonic is a multilingual on-device Text-to-Speech (TTS) system developed by supertone-inc that runs inference entirely through ONNX models. Understanding the synthesis pipeline architecture—spanning from the UnicodeProcessor text normalization to the four-model ONNX inference chain—is essential for rapidly isolating failure points. This guide walks through the most common runtime errors, their root causes in the source code, and practical debugging techniques.

## Understanding the Supertonic Architecture

The synthesis flow is orchestrated across two primary Python files. In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), the **UnicodeProcessor** class (lines 16-105) handles text normalization, stripping emojis, replacing symbols, and wrapping text in language tags like `<en>…</en>`. The **TextToSpeech** class (lines 40-155) manages the four ONNX sub-models: the duration predictor, text encoder, latent-vector estimator, and vocoder. It also handles chunking of long texts and speed scaling.

The command-line interface in [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) serves as the front-end, parsing arguments and invoking `load_text_to_speech()` to initialize the ONNX Runtime sessions. The inference chain flows as follows: argument parsing → model loading (`load_onnx_all`) → voice style loading (`load_voice_style`) → text preprocessing (`UnicodeProcessor.__call__`) → sub-model execution (`dp_ort`, `text_enc_ort`, `vector_est_ort`, `vocoder_ort`) → post-processing and WAV writing.

## Common Synthesis Errors and Solutions

### Invalid Language Codes

If you encounter `ValueError: Invalid language: xx`, the language code supplied via the `--lang` argument is not present in the `AVAILABLE_LANGS` list defined in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py). Verify your code against the supported ISO-639-1 codes in the source file, or update the list if you are adding support for a new language.

### GPU Mode Limitations

Passing `--use-gpu` or setting `use_gpu=True` in `load_text_to_speech()` triggers a `NotImplementedError: GPU mode is not fully tested`. According to the source code in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), GPU inference is deliberately guarded because it is not yet stable. To resolve this, run inference on the CPU by omitting the flag. If GPU support is mandatory, you must build a custom ONNX Runtime with GPU providers and remove the guard clause in the repository.

### Voice-Style and Text Count Mismatches

The error `AssertionError: Number of voice styles (…) must match number of texts (…)` occurs when the `--voice-style` and `--text` arguments have differing lengths. Ensure a 1:1 correspondence between style JSONs and input strings:

```bash
--voice-style style1.json style2.json --text "Hello" "Hola"

```

### Chunking Disabled in Batch Mode

When using `--batch`, Supertonic disables automatic text chunking, which can cause failures with long inputs. The `chunk_text` function in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) (lines 88-130) splits text into ≤300-character chunks, but this is bypassed in batch mode. For long-form synthesis, **omit** the `--batch` flag. If you require both batch processing and chunking, manually split texts using `chunk_text()` and invoke the TTS object iteratively.

### Truncated or Silent Audio

Audio truncation usually indicates inaccurate duration predictions from the duration predictor model. Inspect the `duration` values returned by `dp_ort` and compare them against expected waveform lengths. You can mitigate this by adjusting the `--speed` factor or increasing `--total-step` to improve model stability during inference.

### Missing ONNX Model Files

The error `onnxruntime.capi.onnxruntime_pybind11_state.NotImplementedError: CPUExecutionProvider` often masks a missing model file. Verify that `--onnx-dir` points to a directory containing all four required files: `duration_predictor.onnx`, `text_encoder.onnx`, `vector_estimator.onnx`, and `vocoder.onnx`. The default value (`../assets/onnx`) assumes execution from within the `py/` folder.

### JSON and Unicode Decoding Errors

`UnicodeDecodeError` or `json.JSONDecodeError` typically indicate corrupted or missing asset files. Ensure [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) and voice-style JSONs are present in the assets directory. If damaged, re-download the assets from the Hugging Face release linked in the top-level README.

## Debugging Techniques for ONNX Inference

To isolate failures, wrap the `TextToSpeech.__call__` invocation in a try/except block and log the full traceback before the exception propagates. 

Use the `timer` context manager available in the codebase to profile which sub-model introduces latency. For shape mismatches, print intermediate tensor dimensions immediately before feeding them to the ONNX sessions:

```python
print(text_ids.shape, text_mask.shape)

```

Inspect the duration predictor outputs (`print(duration)`) to verify that the inferred lengths align with the actual phoneme sequences generated by the text encoder.

## Code Examples for Troubleshooting

### Basic Single-Text Synthesis

This command runs with default chunking enabled (recommended for long texts):

```bash
uv run py/example_onnx.py \
  --voice-style ../assets/voice_styles/M1.json \
  --text "The quick brown fox jumps over the lazy dog." \
  --lang en

```

### Reproducing Language Code Errors

Test validation logic programmatically:

```python
from helper import UnicodeProcessor

proc = UnicodeProcessor("../assets/unicode_indexer.json")
try:
    proc._preprocess_text("Hello world", "xx")
except ValueError as e:
    print("Caught error:", e)

```

### Manual Chunking Implementation

When batch mode is required but inputs exceed length limits:

```python
from helper import chunk_text, sanitize_filename

chunks = chunk_text(open("long_paragraph.txt").read())
for i, chunk in enumerate(chunks):
    # Process each chunk individually and concatenate results

    print(f"Processing chunk {i}: {chunk[:50]}...")

```

### Verifying Model Loading

Confirm that ONNX sessions are initialized correctly:

```python
from helper import load_text_to_speech

tts = load_text_to_speech("../assets/onnx", use_gpu=False)
print("Duration predictor inputs:", tts.dp_ort.get_inputs())
print("Vocoder outputs:", tts.vocoder_ort.get_outputs())

```

## Summary

- **Language validation** is strict; ensure codes exist in `AVAILABLE_LANGS` inside [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py).
- **GPU inference** is explicitly disabled in the current release due to stability concerns.
- **Batch mode** disables automatic chunking, which may cause failures with texts longer than 300 characters.
- **Audio truncation** can be mitigated by inspecting duration predictor outputs and adjusting speed parameters.
- **ONNX model paths** must resolve to the four specific `.onnx` files in the assets directory.
- **Debug using** try/except blocks around `TextToSpeech.__call__`, timer contexts, and intermediate tensor shape printing.

## Frequently Asked Questions

### Why does Supertonic throw "Invalid language" errors?

Supertonic validates language codes against an internal `AVAILABLE_LANGS` list in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py). If your ISO-639-1 code is not present, the `UnicodeProcessor._preprocess_text()` method raises a `ValueError`. Add your language to the list or use a supported code like `en`, `ko`, or `ja`.

### How do I fix truncated or silent audio output?

Truncation occurs when the duration predictor underestimates phoneme lengths. Trace the `duration` variable output from the `dp_ort` session in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py). If values are too low, increase the `--total-step` parameter or reduce the `--speed` factor to allow the model more steps for generation.

### Can I use GPU acceleration with Supertonic?

No. The `load_text_to_speech()` function explicitly raises a `NotImplementedError` when `use_gpu=True`. The repository maintainers have not fully tested GPU execution paths, so the code aborts to prevent undefined behavior. CPU execution is the stable path.

### Why does batch processing disable automatic chunking?

The `TextToSpeech` class in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) disables chunking when the internal batch processing logic is active. This is a documented limitation because batch mode assumes you have pre-segmented your inputs. To process long texts, either remove `--batch` to enable automatic chunking, or manually split inputs using `chunk_text()` before passing them to the TTS engine.