# Required ONNX Model Files and How load_onnx_all() Loads Them in Supertonic

> Discover the four essential ONNX model files required by Supertonic and learn how load_onnx_all() efficiently loads them using onnxruntime for your AI audio projects.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: api-reference
- Published: 2026-06-14

---

**Supertonic requires four specific ONNX model files—`duration_predictor.onnx`, `text_encoder.onnx`, `vector_estimator.onnx`, and `vocoder.onnx`—which the `load_onnx_all()` function in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) loads into `onnxruntime.InferenceSession` objects and returns as a tuple in that exact order.**

The supertone-inc/supertonic repository implements a neural text-to-speech pipeline that relies on pre-trained ONNX model files for inference. Understanding which ONNX model files are required and how the `load_onnx_all()` utility orchestrates their loading is essential for initializing the synthesis pipeline correctly.

## Required ONNX Model Files

According to the supertone-inc/supertonic source code, the TTS system expects four distinct ONNX models in a single directory (commonly `../assets/onnx`). Each model serves a specific role in the synthesis pipeline:

### Duration Predictor (`duration_predictor.onnx`)

The duration predictor estimates the length of each phoneme or utterance in frames. During the loading process, `load_onnx_all()` constructs the path using `os.path.join(onnx_dir, "duration_predictor.onnx")` at line 297 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py).

### Text Encoder (`text_encoder.onnx`)

The text encoder converts tokenized input text into a latent representation suitable for the diffusion model. The function builds this path at line 298 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) with `os.path.join(onnx_dir, "text_encoder.onnx")`.

### Vector Estimator (`vector_estimator.onnx`)

This model generates style-conditioned vectors that guide the diffusion process. The corresponding file path is constructed at line 299 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) via `os.path.join(onnx_dir, "vector_estimator.onnx")`.

### Vocoder (`vocoder.onnx`)

The vocoder synthesizes the final output waveform from the diffusion latent representation. The path is assembled at line 300 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) using `os.path.join(onnx_dir, "vocoder.onnx")`.

## How load_onnx_all() Loads the Models

The `load_onnx_all()` function in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) orchestrates the loading sequence by constructing absolute paths for each model, creating `InferenceSession` instances, and returning them in a fixed order.

### Path Construction and Session Creation

First, the function builds the absolute paths for all four model files:

```python
dp_onnx_path = os.path.join(onnx_dir, "duration_predictor.onnx")       # py/helper.py#L297

text_enc_onnx_path = os.path.join(onnx_dir, "text_encoder.onnx")      # py/helper.py#L298

vector_est_onnx_path = os.path.join(onnx_dir, "vector_estimator.onnx") # py/helper.py#L299

vocoder_onnx_path = os.path.join(onnx_dir, "vocoder.onnx")            # py/helper.py#L300

```

Then, each path is passed to the internal `load_onnx()` helper along with session options and execution providers:

```python
dp_ort = load_onnx(dp_onnx_path, opts, providers)          # py/helper.py#L302

text_enc_ort = load_onnx(text_enc_onnx_path, opts, providers)   # py/helper.py#L303

vector_est_ort = load_onnx(vector_est_onnx_path, opts, providers) # py/helper.py#L304

vocoder_ort = load_onnx(vocoder_onnx_path, opts, providers)   # py/helper.py#L305

```

### Return Value

The function returns a tuple containing the four sessions in the order mentioned above:

```python
return dp_ort, text_enc_ort, vector_est_ort, vocoder_ort  # py/helper.py#L306

```

This return statement appears at line 306 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py).

## Practical Usage Examples

### Direct Loading with load_onnx_all()

You can directly invoke the loader to initialize the four model sessions:

```python
import onnxruntime as ort
from helper import load_onnx_all

onnx_dir = "../assets/onnx"
opts = ort.SessionOptions()
providers = ["CPUExecutionProvider"]  # Use ["CUDAExecutionProvider"] for GPU

dp_session, text_enc_session, vec_est_session, vocoder_session = load_onnx_all(
    onnx_dir, opts, providers
)

print("Loaded models:")
print(f"• Duration predictor: {dp_session}")
print(f"• Text encoder: {text_enc_session}")
print(f"• Vector estimator: {vec_est_session}")
print(f"• Vocoder: {vocoder_session}")

```

### Integration via High-Level API

Typically, you will use `load_text_to_speech()` from [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), which internally calls `load_onnx_all()`:

```python
from helper import load_text_to_speech

tts = load_text_to_speech("../assets/onnx", use_gpu=False)
wav, duration = tts("Hello world!", "en", style_vector, total_step=8, speed=1.05)

```

### Verifying Model Files on Disk

Before loading, verify that all required files exist in the target directory:

```python
import os

def list_onnx_models(dir_path: str) -> list[str]:
    return [f for f in os.listdir(dir_path) if f.endswith(".onnx")]

print(list_onnx_models("../assets/onnx"))

# Output: ['duration_predictor.onnx', 'text_encoder.onnx', 'vector_estimator.onnx', 'vocoder.onnx']

```

## Summary

- **Four required files**: `duration_predictor.onnx`, `text_encoder.onnx`, `vector_estimator.onnx`, and `vocoder.onnx` must reside in the specified ONNX directory.
- **Loading location**: The `load_onnx_all()` function in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) handles the loading process at lines 297-306.
- **Fixed order**: The function returns sessions as a tuple `(duration_predictor, text_encoder, vector_estimator, vocoder)`.
- **Internal helper**: Each model is loaded via the internal `load_onnx()` function with specified session options and execution providers.
- **High-level wrapper**: `load_text_to_speech()` provides a convenient wrapper that internally calls `load_onnx_all()`.

## Frequently Asked Questions

### What directory should contain the ONNX model files?

The ONNX model files should be placed in a directory specified by the `onnx_dir` parameter passed to `load_onnx_all()`. By convention, the Supertonic examples use `../assets/onnx` relative to the Python script location, but any absolute or relative path works as long as it contains all four required `.onnx` files.

### Can I use GPU acceleration when loading these models?

Yes. Pass `["CUDAExecutionProvider"]` as the `providers` argument to `load_onnx_all()` instead of `["CPUExecutionProvider"]`. Ensure you have the appropriate CUDA drivers and ONNX Runtime GPU packages installed. The `load_text_to_speech()` function simplifies this by accepting a `use_gpu` boolean parameter.

### What happens if one of the required ONNX files is missing?

The `load_onnx()` internal helper (called by `load_onnx_all()`) will raise an error when attempting to create an `onnxruntime.InferenceSession` for the missing file. Verify all four files exist using `os.listdir()` or similar checks before calling the loader.

### How do I know which model corresponds to which return value from load_onnx_all()?

The return order is fixed and documented in the source code at line 306 of [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py): duration predictor first, text encoder second, vector estimator third, and vocoder fourth. You should unpack the tuple in this exact order: `dp_session, text_enc_session, vec_est_session, vocoder_session = load_onnx_all(...)`.