Required ONNX Model Files and How load_onnx_all() Loads Them in Supertonic
Supertonic requires four specific ONNX model files—duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, and vocoder.onnx—which the load_onnx_all() function in py/helper.py loads into onnxruntime.InferenceSession objects and returns as a tuple in that exact order.
The supertone-inc/supertonic repository implements a neural text-to-speech pipeline that relies on pre-trained ONNX model files for inference. Understanding which ONNX model files are required and how the load_onnx_all() utility orchestrates their loading is essential for initializing the synthesis pipeline correctly.
Required ONNX Model Files
According to the supertone-inc/supertonic source code, the TTS system expects four distinct ONNX models in a single directory (commonly ../assets/onnx). Each model serves a specific role in the synthesis pipeline:
Duration Predictor (duration_predictor.onnx)
The duration predictor estimates the length of each phoneme or utterance in frames. During the loading process, load_onnx_all() constructs the path using os.path.join(onnx_dir, "duration_predictor.onnx") at line 297 in py/helper.py.
Text Encoder (text_encoder.onnx)
The text encoder converts tokenized input text into a latent representation suitable for the diffusion model. The function builds this path at line 298 in py/helper.py with os.path.join(onnx_dir, "text_encoder.onnx").
Vector Estimator (vector_estimator.onnx)
This model generates style-conditioned vectors that guide the diffusion process. The corresponding file path is constructed at line 299 in py/helper.py via os.path.join(onnx_dir, "vector_estimator.onnx").
Vocoder (vocoder.onnx)
The vocoder synthesizes the final output waveform from the diffusion latent representation. The path is assembled at line 300 in py/helper.py using os.path.join(onnx_dir, "vocoder.onnx").
How load_onnx_all() Loads the Models
The load_onnx_all() function in py/helper.py orchestrates the loading sequence by constructing absolute paths for each model, creating InferenceSession instances, and returning them in a fixed order.
Path Construction and Session Creation
First, the function builds the absolute paths for all four model files:
dp_onnx_path = os.path.join(onnx_dir, "duration_predictor.onnx") # py/helper.py#L297
text_enc_onnx_path = os.path.join(onnx_dir, "text_encoder.onnx") # py/helper.py#L298
vector_est_onnx_path = os.path.join(onnx_dir, "vector_estimator.onnx") # py/helper.py#L299
vocoder_onnx_path = os.path.join(onnx_dir, "vocoder.onnx") # py/helper.py#L300
Then, each path is passed to the internal load_onnx() helper along with session options and execution providers:
dp_ort = load_onnx(dp_onnx_path, opts, providers) # py/helper.py#L302
text_enc_ort = load_onnx(text_enc_onnx_path, opts, providers) # py/helper.py#L303
vector_est_ort = load_onnx(vector_est_onnx_path, opts, providers) # py/helper.py#L304
vocoder_ort = load_onnx(vocoder_onnx_path, opts, providers) # py/helper.py#L305
Return Value
The function returns a tuple containing the four sessions in the order mentioned above:
return dp_ort, text_enc_ort, vector_est_ort, vocoder_ort # py/helper.py#L306
This return statement appears at line 306 in py/helper.py.
Practical Usage Examples
Direct Loading with load_onnx_all()
You can directly invoke the loader to initialize the four model sessions:
import onnxruntime as ort
from helper import load_onnx_all
onnx_dir = "../assets/onnx"
opts = ort.SessionOptions()
providers = ["CPUExecutionProvider"] # Use ["CUDAExecutionProvider"] for GPU
dp_session, text_enc_session, vec_est_session, vocoder_session = load_onnx_all(
onnx_dir, opts, providers
)
print("Loaded models:")
print(f"• Duration predictor: {dp_session}")
print(f"• Text encoder: {text_enc_session}")
print(f"• Vector estimator: {vec_est_session}")
print(f"• Vocoder: {vocoder_session}")
Integration via High-Level API
Typically, you will use load_text_to_speech() from py/helper.py, which internally calls load_onnx_all():
from helper import load_text_to_speech
tts = load_text_to_speech("../assets/onnx", use_gpu=False)
wav, duration = tts("Hello world!", "en", style_vector, total_step=8, speed=1.05)
Verifying Model Files on Disk
Before loading, verify that all required files exist in the target directory:
import os
def list_onnx_models(dir_path: str) -> list[str]:
return [f for f in os.listdir(dir_path) if f.endswith(".onnx")]
print(list_onnx_models("../assets/onnx"))
# Output: ['duration_predictor.onnx', 'text_encoder.onnx', 'vector_estimator.onnx', 'vocoder.onnx']
Summary
- Four required files:
duration_predictor.onnx,text_encoder.onnx,vector_estimator.onnx, andvocoder.onnxmust reside in the specified ONNX directory. - Loading location: The
load_onnx_all()function inpy/helper.pyhandles the loading process at lines 297-306. - Fixed order: The function returns sessions as a tuple
(duration_predictor, text_encoder, vector_estimator, vocoder). - Internal helper: Each model is loaded via the internal
load_onnx()function with specified session options and execution providers. - High-level wrapper:
load_text_to_speech()provides a convenient wrapper that internally callsload_onnx_all().
Frequently Asked Questions
What directory should contain the ONNX model files?
The ONNX model files should be placed in a directory specified by the onnx_dir parameter passed to load_onnx_all(). By convention, the Supertonic examples use ../assets/onnx relative to the Python script location, but any absolute or relative path works as long as it contains all four required .onnx files.
Can I use GPU acceleration when loading these models?
Yes. Pass ["CUDAExecutionProvider"] as the providers argument to load_onnx_all() instead of ["CPUExecutionProvider"]. Ensure you have the appropriate CUDA drivers and ONNX Runtime GPU packages installed. The load_text_to_speech() function simplifies this by accepting a use_gpu boolean parameter.
What happens if one of the required ONNX files is missing?
The load_onnx() internal helper (called by load_onnx_all()) will raise an error when attempting to create an onnxruntime.InferenceSession for the missing file. Verify all four files exist using os.listdir() or similar checks before calling the loader.
How do I know which model corresponds to which return value from load_onnx_all()?
The return order is fixed and documented in the source code at line 306 of py/helper.py: duration predictor first, text encoder second, vector estimator third, and vocoder fourth. You should unpack the tuple in this exact order: dp_session, text_enc_session, vec_est_session, vocoder_session = load_onnx_all(...).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →