How to Install Supertonic: Complete Setup Guide for Python, Node.js, and More

Install Supertonic by cloning the repository with Git LFS enabled, downloading the ONNX model assets from Hugging Face, and running the language-specific SDK commands found in the py/, nodejs/, web/, or other platform directories.

Supertonic is a cross-platform, on-device Text-to-Speech (TTS) system built on ONNX Runtime. The supertone-inc/supertonic repository provides a 99‑M‑parameter open‑weight model and language-specific SDKs for Python, Node.js, Java, C++, C#, Go, Swift, iOS, Rust, and Flutter. This guide covers the complete installation workflow from prerequisites to running your first synthesis.

Prerequisites: Git LFS and Model Assets

Before installing any SDK, you must prepare the model files and supporting metadata.

Install Git LFS

The model assets are hosted on Hugging Face and tracked via Git LFS. Initialize Git LFS before cloning:

git lfs install

Download Model Assets

Clone the specific model version into the assets/ directory. For Supertonic‑3, run:

git clone https://huggingface.co/Supertone/supertonic-3 assets

This downloads the ONNX files, preset voice JSONs, and metadata required for inference. Once downloaded, all SDKs run entirely offline with zero network traffic during synthesis.

Language-Specific Installation Steps

Each supported language includes a helper module (e.g., helper.py, helper.js, helper.cpp) that abstracts model loading, text preprocessing (including ten expression tags), and audio post‑processing into 44.1 kHz 16‑bit WAV output.

Python

Install the PyPI package and run the example which auto‑downloads the model on first use:

pip install supertonic

Run the example in py/example_onnx.py:

from supertonic import TTS

# First run downloads the model assets from Hugging Face automatically.

tts = TTS(auto_download=True)

# Choose a voice style (e.g. "M1")

style = tts.get_voice_style(voice_name="M1")

# Synthesize speech

wav, duration = tts.synthesize(
    text="Supertonic is a lightning‑fast, on‑device TTS system.",
    lang="en",
    voice_style=style,
    total_steps=8,
    speed=1.05,
)

# Save the audio file

tts.save_audio(wav, "output.wav")
print(f"Generated {duration[0]:.2f}s of audio")

Core implementation logic resides in py/helper.py.

Node.js

Navigate to the Node.js directory and install dependencies:

cd nodejs/
npm install
npm start

The entry point nodejs/example_onnx.js demonstrates the async API:

const { TTS } = require("./helper.js");

// Initialise the TTS object; model files are expected under ./assets
const tts = new TTS({ autoDownload: true });

async function synth() {
  const style = await tts.getVoiceStyle("M1");
  const { wav, duration } = await tts.synthesize({
    text: "Supertonic runs locally with no network calls.",
    lang: "en",
    voiceStyle: style,
    totalSteps: 8,
    speed: 1.0,
  });
  await tts.saveAudio(wav, "output.wav");
  console.log(`Generated ${duration}s of audio`);
}
synth();

Wrapper functions are defined in nodejs/helper.js.

Browser (WebGPU/WASM)

For browser-based inference using WebGPU:

cd web/
npm install
npm run dev

Open http://localhost:5173 to access the UI. The Web implementation loads ONNX Runtime WebGPU via web/main.js and runs inference locally in the browser.

Java

Build and run the Java SDK using Maven:

cd java/
mvn clean install

C++

Compile the native C++ wrapper with CMake:

cd cpp/
cmake .. && cmake --build .

The high-performance wrapper is implemented in cpp/helper.cpp, with a CLI demo available in cpp/example_onnx.cpp.

C#

Restore dependencies and run the .NET application:

cd csharp/
dotnet restore && dotnet run

The interop layer is defined in csharp/Helper.cs, demonstrated in csharp/ExampleONNX.cs.

Go

Download modules and execute the example:

cd go/
go mod download && go run example_onnx.go helper.go

Go bindings for ONNX Runtime are provided in go/helper.go, with usage demonstrated in go/example_onnx.go.

Swift

Build the release version for macOS:

cd swift/
swift build -c release

The Swift wrapper is located at swift/Sources/ExampleONNX.swift.

iOS

Generate the Xcode project and build:

cd ios/ExampleiOSApp
xcodegen generate

Integration examples are found in ios/ExampleiOSApp/App.swift.

Rust

Build the Rust implementation:

cd rust/
cargo build --release

The safe Rust wrapper is implemented in rust/src/example_onnx.rs.

Flutter

Install Flutter dependencies:

cd flutter/
flutter pub get

The Flutter SDK configuration is defined in flutter/pubspec.yaml and requires a recent Flutter SDK version.

Verifying Your Installation

After installation, verify that the ONNX Runtime loads the model correctly by checking the console output for successfulasset loading. Each SDK automatically caches the 99‑M‑parameter model locally after the first download, ensuring subsequent syntheses require no network access. Confirm that assets/ contains the .onnx model files and voice preset JSONs before running examples.

Summary

  • Enable Git LFS (git lfs install) before cloning to handle large model files tracked on Hugging Face.

  • Download assets via git clone https://huggingface.co/Supertone/supertonic-3 assets to obtain the ONNX model and voice presets.

  • Choose your SDK: Python (pip install supertonic), Node.js (npm install), C++ (cmake), C# (dotnet), Go (go mod download), Swift (swift build), Rust (cargo build), or Flutter (flutter pub get).

  • Run offline: After initial setup, all inference runs locally using the helper modules (e.g., helper.py, helper.js) with no external API calls.

  • Output format: All SDKs produce 44.1 kHz 16‑bit WAV audio via the shared preprocessing and post‑processing pipeline.

Frequently Asked Questions

Do I need an internet connection to use Supertonic after installation?

No. After the initial download of the ONNX model assets from Hugging Face, Supertonic runs entirely on-device. The TTS class in each SDK (e.g., TTS(auto_download=True) in Python) fetches the model files only on the first run, after which all synthesis occurs offline with zero network traffic.

Where are the model files stored in the repository?

The model files are not stored directly in the Git repository; they are tracked via Git LFS and hosted on Hugging Face. You must run git clone https://huggingface.co/Supertone/supertonic-3 assets to populate the assets/ directory with the ONNX weights and voice preset JSONs required by the helper modules.

What is the purpose of the helper files in each SDK directory?

The helper files (e.g., py/helper.py, nodejs/helper.js, cpp/helper.cpp) provide a unified abstraction layer that handles ONNX Runtime initialization, text preprocessing (including support for ten expression tags), and audio post‑processing. They convert raw model outputs into standardized 44.1 kHz 16‑bit WAV files, simplifying the API for end users.

Can I run Supertonic in a web browser without a backend server?

Yes. The web/ directory contains a WebGPU/WASM implementation that runs entirely in the browser. After running npm install and npm run dev in the web/ folder, the application loads ONNX Runtime WebGPU via web/main.js and performs inference locally using the client's GPU, requiring no server-side processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →