# How to Deploy Supertonic in a Production Environment: Complete Setup Guide

> Learn to deploy Supertonic in a production environment. Containerize the Python HTTP server, manage model assets with Git-LFS, and use NGINX for TLS and load balancing. Get the complete setup guide.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-12

---

**Deploy Supertonic in a production environment by containerizing the Python HTTP server, mounting the Git-LFS managed ONNX model assets from the `assets` directory, and fronting the service with NGINX for TLS termination and load balancing.**

Supertonic is a lightweight, on-device text-to-speech (TTS) system developed by Supertone Inc. that runs inference using ONNX Runtime. To deploy Supertonic in a production environment, you must host the ~99M-parameter model checkpoint, expose the inference service via the built-in `supertonic serve` command, and configure container orchestration for scalability. The architecture supports both cloud-native deployments and edge devices like Raspberry Pi without requiring GPU acceleration.

## Prepare the Model Assets

Production deployment starts with acquiring the model files stored in Git-LFS. The `assets` directory in the supertone-inc/supertonic repository contains `model.onnx`, [`voice_styles.json`](https://github.com/supertone-inc/supertonic/blob/main/voice_styles.json), and other runtime resources required by the inference engine.

Install Git LFS and clone the model assets from Hugging Face:

```bash

# Install Git LFS (macOS example)

brew install git-lfs && git lfs install

# Clone the model assets

git clone https://huggingface.co/Supertone/supertonic-3 assets

```

The `SUPERTONIC_ASSETS` environment variable can optionally point to this directory if you mount the assets outside the default path.

## Run the Built-in HTTP Server

The Python SDK provides a production-ready HTTP server via the `supertonic serve` command. Install the package with server extras to include the necessary dependencies:

```bash
pip install 'supertonic[serve]'
supertonic serve --host 0.0.0.0 --port 8080

```

The server exposes two endpoints:

- **`POST /v1/tts`** – Native Supertonic JSON format accepting `text`, `lang`, `voice_style`, `total_steps`, and `speed` parameters.
- **`POST /v1/audio/speech`** – OpenAI-compatible endpoint matching the OpenAI TTS API specification.

Full configuration options are documented in [`py/README.md`](https://github.com/supertone-inc/supertonic/blob/main/py/README.md) within the repository.

## Containerize for Production

Containerization ensures consistent environments across development and production. Below is a minimal Dockerfile that bundles the Python server with the model assets:

```dockerfile
FROM python:3.12-slim

# Install system dependencies (ONNX Runtime requires libgomp)

RUN apt-get update && apt-get install -y --no-install-recommends \
    libgomp1 && rm -rf /var/lib/apt/lists/*

# Create non-root user for security

RUN useradd -m appuser
USER appuser
WORKDIR /app

# Install Supertonic with server extras

RUN pip install --no-cache-dir "supertonic[serve]"

# Copy model assets into container

COPY --chown=appuser assets /app/assets

EXPOSE 8080

CMD ["supertonic", "serve", "--host", "0.0.0.0", "--port", "8080"]

```

Build and push the image to your registry:

```bash
docker build -t supertonic:latest .
docker tag supertonic:latest your-registry/supertonic:latest
docker push your-registry/supertonic:latest

```

Deploy using Docker Compose:

```yaml
version: "3.8"
services:
  tts:
    image: your-registry/supertonic:latest
    restart: unless-stopped
    ports:
      - "8080:8080"
    environment:
      - SUPERTONIC_ASSETS=/app/assets

```

## Configure Reverse Proxy and Systemd

For production traffic, front the container with a reverse proxy to handle TLS and connection pooling.

### NGINX Configuration

```nginx
server {
    listen 443 ssl;
    server_name tts.example.com;

    ssl_certificate /etc/letsencrypt/live/tts.example.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/tts.example.com/privkey.pem;

    location / {
        proxy_pass http://localhost:8080;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
    }
}

```

Reload NGINX after configuration changes to apply TLS termination.

### Systemd Service

For bare-metal or VM deployments without containers, run Supertonic as a systemd service:

```ini
[Unit]
Description=Supertonic TTS Service
After=network.target

[Service]
User=appuser
Group=appuser
WorkingDirectory=/opt/supertonic
ExecStart=/usr/local/bin/supertonic serve --host 0.0.0.0 --port 8080
Restart=on-failure
Environment=SUPERTONIC_ASSETS=/opt/supertonic/assets

[Install]
WantedBy=multi-user.target

```

Enable the service with:

```bash
systemctl daemon-reload
systemctl enable --now supertonic.service

```

Monitor logs using `journalctl -u supertonic`.

## Deploy to Edge Devices

Supertonic's ONNX architecture enables deployment on resource-constrained devices without GPU acceleration.

### Raspberry Pi

Install the ONNX Runtime and copy the `assets` folder to the device. Run the Python server or compile the Rust example for lower memory overhead:

```bash
cd rust && cargo build --release && ./target/release/example_onnx --text "Hello" --lang en

```

### iOS

Use the Swift SDK located in `ios/ExampleiOSApp`. Embed the `assets` directory in the app bundle and reference the files directly from the Swift inference wrapper.

### Web Browser

Deploy the browser-based demo using `onnxruntime-web` as documented in [`web/README.md`](https://github.com/supertone-inc/supertonic/blob/main/web/README.md). This configuration runs entirely client-side without requiring a backend server, suitable for serverless deployments.

### Flutter

Add the Flutter SDK as a dependency and bundle the assets per the instructions in `flutter/lib/main.dart` and the accompanying [`flutter/README.md`](https://github.com/supertone-inc/supertonic/blob/main/flutter/README.md).

## Monitoring and Scaling

Production deployments require observability and horizontal scaling capabilities.

**Metrics Collection**: Wrap the Python server with a Prometheus client to scrape latency and memory usage statistics from the ONNX Runtime. Install the client with `pip install prometheus_client` and expose metrics on a separate port.

**Autoscaling**: In Kubernetes, configure a Horizontal Pod Autoscaler (HPA) based on CPU utilization or custom Prometheus metrics to handle traffic spikes across multiple replicas.

**Caching**: For repeated utterances, implement a caching layer using Redis or a shared volume to store generated WAV files, reducing redundant inference calls.

## Summary

- **Acquire model assets** by cloning the Hugging Face repository with Git LFS to obtain `model.onnx` and [`voice_styles.json`](https://github.com/supertone-inc/supertonic/blob/main/voice_styles.json).
- **Start the service** using `supertonic serve` from the Python SDK with `supertonic[serve]` extras installed.
- **Containerize** using the provided Dockerfile, ensuring `libgomp1` is installed and assets are copied to `/app/assets`.
- **Secure traffic** by placing NGINX in front of the container to terminate TLS and manage connections.
- **Run persistently** via systemd on Linux VMs or deploy to Kubernetes for orchestration.
- **Scale horizontally** using Prometheus metrics and caching strategies to optimize resource utilization.

## Frequently Asked Questions

### Does Supertonic require a GPU for production deployment?

No. Supertonic runs inference on the CPU using ONNX Runtime, making it suitable for cost-effective cloud instances and edge devices without GPU acceleration. The ~99M-parameter model is optimized for on-device execution, allowing production deployment on standard x86 and ARM processors.

### How do I scale Supertonic to handle high traffic?

Deploy multiple container instances behind a load balancer and use Kubernetes Horizontal Pod Autoscaler to scale based on CPU utilization. Because the inference is stateless, you can distribute requests across replicas. Implement caching for repeated text inputs to reduce computational overhead.

### What is the difference between the `/v1/tts` and `/v1/audio/speech` endpoints?

The `/v1/tts` endpoint accepts native Supertonic parameters including `voice_style`, `total_steps`, and `speed` for fine-grained control over synthesis. The `/v1/audio/speech` endpoint provides OpenAI API compatibility, accepting the same payload format as OpenAI's TTS service for easier integration with existing client libraries.

### Can I deploy Supertonic without using Docker?

Yes. Install the Python SDK directly on a Linux VM using `pip install 'supertonic[serve]'`, clone the model assets to a directory like `/opt/supertonic/assets`, and run the service using the provided systemd unit file. This approach is suitable for environments where containerization is restricted or unnecessary.