How to Deploy Supertonic in a Production Environment: Complete Setup Guide

Deploy Supertonic in a production environment by containerizing the Python HTTP server, mounting the Git-LFS managed ONNX model assets from the assets directory, and fronting the service with NGINX for TLS termination and load balancing.

Supertonic is a lightweight, on-device text-to-speech (TTS) system developed by Supertone Inc. that runs inference using ONNX Runtime. To deploy Supertonic in a production environment, you must host the ~99M-parameter model checkpoint, expose the inference service via the built-in supertonic serve command, and configure container orchestration for scalability. The architecture supports both cloud-native deployments and edge devices like Raspberry Pi without requiring GPU acceleration.

Prepare the Model Assets

Production deployment starts with acquiring the model files stored in Git-LFS. The assets directory in the supertone-inc/supertonic repository contains model.onnx, voice_styles.json, and other runtime resources required by the inference engine.

Install Git LFS and clone the model assets from Hugging Face:


# Install Git LFS (macOS example)

brew install git-lfs && git lfs install

# Clone the model assets

git clone https://huggingface.co/Supertone/supertonic-3 assets

The SUPERTONIC_ASSETS environment variable can optionally point to this directory if you mount the assets outside the default path.

Run the Built-in HTTP Server

The Python SDK provides a production-ready HTTP server via the supertonic serve command. Install the package with server extras to include the necessary dependencies:

pip install 'supertonic[serve]'
supertonic serve --host 0.0.0.0 --port 8080

The server exposes two endpoints:

  • POST /v1/tts – Native Supertonic JSON format accepting text, lang, voice_style, total_steps, and speed parameters.
  • POST /v1/audio/speech – OpenAI-compatible endpoint matching the OpenAI TTS API specification.

Full configuration options are documented in py/README.md within the repository.

Containerize for Production

Containerization ensures consistent environments across development and production. Below is a minimal Dockerfile that bundles the Python server with the model assets:

FROM python:3.12-slim

# Install system dependencies (ONNX Runtime requires libgomp)

RUN apt-get update && apt-get install -y --no-install-recommends \
    libgomp1 && rm -rf /var/lib/apt/lists/*

# Create non-root user for security

RUN useradd -m appuser
USER appuser
WORKDIR /app

# Install Supertonic with server extras

RUN pip install --no-cache-dir "supertonic[serve]"

# Copy model assets into container

COPY --chown=appuser assets /app/assets

EXPOSE 8080

CMD ["supertonic", "serve", "--host", "0.0.0.0", "--port", "8080"]

Build and push the image to your registry:

docker build -t supertonic:latest .
docker tag supertonic:latest your-registry/supertonic:latest
docker push your-registry/supertonic:latest

Deploy using Docker Compose:

version: "3.8"
services:
  tts:
    image: your-registry/supertonic:latest
    restart: unless-stopped
    ports:
      - "8080:8080"
    environment:
      - SUPERTONIC_ASSETS=/app/assets

Configure Reverse Proxy and Systemd

For production traffic, front the container with a reverse proxy to handle TLS and connection pooling.

NGINX Configuration

server {
    listen 443 ssl;
    server_name tts.example.com;

    ssl_certificate /etc/letsencrypt/live/tts.example.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/tts.example.com/privkey.pem;

    location / {
        proxy_pass http://localhost:8080;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
    }
}

Reload NGINX after configuration changes to apply TLS termination.

Systemd Service

For bare-metal or VM deployments without containers, run Supertonic as a systemd service:

[Unit]
Description=Supertonic TTS Service
After=network.target

[Service]
User=appuser
Group=appuser
WorkingDirectory=/opt/supertonic
ExecStart=/usr/local/bin/supertonic serve --host 0.0.0.0 --port 8080
Restart=on-failure
Environment=SUPERTONIC_ASSETS=/opt/supertonic/assets

[Install]
WantedBy=multi-user.target

Enable the service with:

systemctl daemon-reload
systemctl enable --now supertonic.service

Monitor logs using journalctl -u supertonic.

Deploy to Edge Devices

Supertonic's ONNX architecture enables deployment on resource-constrained devices without GPU acceleration.

Raspberry Pi

Install the ONNX Runtime and copy the assets folder to the device. Run the Python server or compile the Rust example for lower memory overhead:

cd rust && cargo build --release && ./target/release/example_onnx --text "Hello" --lang en

iOS

Use the Swift SDK located in ios/ExampleiOSApp. Embed the assets directory in the app bundle and reference the files directly from the Swift inference wrapper.

Web Browser

Deploy the browser-based demo using onnxruntime-web as documented in web/README.md. This configuration runs entirely client-side without requiring a backend server, suitable for serverless deployments.

Flutter

Add the Flutter SDK as a dependency and bundle the assets per the instructions in flutter/lib/main.dart and the accompanying flutter/README.md.

Monitoring and Scaling

Production deployments require observability and horizontal scaling capabilities.

Metrics Collection: Wrap the Python server with a Prometheus client to scrape latency and memory usage statistics from the ONNX Runtime. Install the client with pip install prometheus_client and expose metrics on a separate port.

Autoscaling: In Kubernetes, configure a Horizontal Pod Autoscaler (HPA) based on CPU utilization or custom Prometheus metrics to handle traffic spikes across multiple replicas.

Caching: For repeated utterances, implement a caching layer using Redis or a shared volume to store generated WAV files, reducing redundant inference calls.

Summary

  • Acquire model assets by cloning the Hugging Face repository with Git LFS to obtain model.onnx and voice_styles.json.
  • Start the service using supertonic serve from the Python SDK with supertonic[serve] extras installed.
  • Containerize using the provided Dockerfile, ensuring libgomp1 is installed and assets are copied to /app/assets.
  • Secure traffic by placing NGINX in front of the container to terminate TLS and manage connections.
  • Run persistently via systemd on Linux VMs or deploy to Kubernetes for orchestration.
  • Scale horizontally using Prometheus metrics and caching strategies to optimize resource utilization.

Frequently Asked Questions

Does Supertonic require a GPU for production deployment?

No. Supertonic runs inference on the CPU using ONNX Runtime, making it suitable for cost-effective cloud instances and edge devices without GPU acceleration. The ~99M-parameter model is optimized for on-device execution, allowing production deployment on standard x86 and ARM processors.

How do I scale Supertonic to handle high traffic?

Deploy multiple container instances behind a load balancer and use Kubernetes Horizontal Pod Autoscaler to scale based on CPU utilization. Because the inference is stateless, you can distribute requests across replicas. Implement caching for repeated text inputs to reduce computational overhead.

What is the difference between the /v1/tts and /v1/audio/speech endpoints?

The /v1/tts endpoint accepts native Supertonic parameters including voice_style, total_steps, and speed for fine-grained control over synthesis. The /v1/audio/speech endpoint provides OpenAI API compatibility, accepting the same payload format as OpenAI's TTS service for easier integration with existing client libraries.

Can I deploy Supertonic without using Docker?

Yes. Install the Python SDK directly on a Linux VM using pip install 'supertonic[serve]', clone the model assets to a directory like /opt/supertonic/assets, and run the service using the provided systemd unit file. This approach is suitable for environments where containerization is restricted or unnecessary.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →