# Supertonic Memory Footprint and Latency on Raspberry Pi and E-Readers

> Discover Supertonic memory footprint and latency on Raspberry Pi and e-readers. Achieve real-time factors of 0.5 and 0.3 on edge devices with just 300 MiB RAM.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: performance
- Published: 2026-06-14

---

**Supertonic runs locally on edge devices with approximately 300 MiB of RAM and achieves real-time factors of 0.5 on Raspberry Pi 4 and 0.3 on Onyx Boox Go 6 e-readers.**

The **supertone-inc/supertonic** repository provides a 99-million-parameter ONNX text-to-speech model specifically optimized for CPU-only deployment. Unlike cloud-dependent TTS systems, Supertonic is engineered to operate entirely offline on resource-constrained hardware, making it viable for Raspberry Pi single-board computers and ARM-based e-readers with limited memory budgets.

## Memory Footprint on Edge Devices

Supertonic maintains a minimal memory profile through efficient ONNX Runtime execution and quantized model weights.

### Model Size and Disk Usage

The compressed ONNX model occupies roughly **140 MiB** on disk. After decompression, the 99 M-parameter model expands to approximately **200 MiB**. This compact size allows the model to fit comfortably on devices with limited storage, such as SD cards used in Raspberry Pi deployments.

### Runtime Memory Allocation

During inference, the total RAM usage reaches approximately **300 MiB**. This includes:
- The loaded model weights (~200 MiB)
- ONNX Runtime session buffers
- Tensor allocation for audio generation
- Working memory for tokenization

According to the performance metrics in `img/metrics/runtime_cpu_gpu_latency_memory.png`, this footprint leaves sufficient headroom on 512 MiB to 1 GiB systems, making it suitable for devices like the **Raspberry Pi 4** and **Onyx Boox Go 6**.

## Latency Benchmarks on Raspberry Pi and E-Readers

Supertonic delivers interactive latency on ARM processors without GPU acceleration.

### Raspberry Pi 4 Performance

On a **Raspberry Pi 4 (4 GB RAM, ARM Cortex-A72)**, the observed latency follows the formula:

**150 ms + 20 ms × audio_seconds**

This translates to a **real-time factor (RTF) of approximately 0.5**, meaning the system generates one second of speech in roughly 0.5 seconds of wall-clock time. The latency budget includes model loading, tokenization, and the ONNX inference pass.

### E-Reader Performance (Onyx Boox Go 6)

On an **Onyx Boox Go 6 e-reader (ARM Cortex-A53, 2 GB RAM, no GPU)**, the latency profile shifts to:

**300 ms + 15 ms × audio_seconds**

This yields an **RTF of approximately 0.3**, demonstrating that even low-power e-readers can achieve real-time speech synthesis. The device maintains this performance in airplane mode, processing all data locally without cloud connectivity.

### Real-Time Factor Explained

An **RTF below 1.0** indicates real-time capability. Supertonic achieves:
- **0.5 RTF** on Raspberry Pi 4 (faster than real-time)
- **0.3 RTF** on Cortex-A53 e-readers (faster than real-time)

These measurements appear in the repository's benchmark visualization at `img/metrics/runtime_cpu_gpu_latency_memory.png`, which compares CPU runtime, memory consumption, and latency across hardware platforms.

## Measuring Memory and Latency in Code

You can verify these metrics on your target device using the Python SDK. The implementation in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) handles ONNX session initialization, while [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) demonstrates inference patterns.

```python
from supertonic import TTS
import time, psutil, os

# Initialize the TTS engine (auto-downloads model on first run)

tts = TTS(auto_download=True)

def print_mem():
    """Print current process memory in MiB."""
    proc = psutil.Process(os.getpid())
    print(f"Memory usage: {proc.memory_info().rss / 1024**2:.1f} MiB")

# Warm-up loads ONNX session (see py/helper.py)

tts.get_voice_style("M1")
print_mem()  # ~200 MiB after model load

text = "Supertonic runs fast on tiny devices."
start = time.time()
wav, _ = tts.synthesize(text=text, lang="en", voice_style=tts.get_voice_style("M1"))
latency = time.time() - start

print(f"Latency: {latency*1000:.0f} ms")
print_mem()  # ~300 MiB after inference buffers

```

Typical output on Raspberry Pi 4:

```

Memory usage: 202.3 MiB
Latency: 158 ms
Memory usage: 298.7 MiB

```

## Implementation Details

The C++ implementation in [`cpp/example_onnx.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/example_onnx.cpp) exposes explicit memory management through `Ort::MemoryInfo` objects, mirroring the strategy used to achieve low-memory operation on embedded devices. This approach minimizes heap fragmentation during inference, critical for devices with 512 MiB to 1 GiB of RAM.

Key files referenced:
- [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md) - Contains the edge-device readiness statement and performance chart
- `img/metrics/runtime_cpu_gpu_latency_memory.png` - Visual benchmark of latency and memory across hardware
- [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) - Python SDK implementation for model loading
- [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) - Python inference example
- [`cpp/example_onnx.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/example_onnx.cpp) - C++ implementation showing memory-efficient ONNX Runtime usage

## Summary

- **Memory Footprint**: Approximately 300 MiB RAM during inference (200 MiB model + 100 MiB buffers)
- **Disk Space**: 140 MiB compressed, 200 MiB decompressed
- **Raspberry Pi 4 Latency**: 150 ms + 20 ms × audio_seconds (RTF ~0.5)
- **E-Reader Latency**: 300 ms + 15 ms × audio_seconds (RTF ~0.3)
- **Hardware Requirements**: CPU-only; no GPU necessary
- **Target Devices**: Raspberry Pi 4, Onyx Boox Go 6, and similar ARM-based edge devices

## Frequently Asked Questions

### How much RAM does Supertonic require?

Supertonic requires approximately **300 MiB of RAM** during active inference. This includes the 200 MiB ONNX model and approximately 100 MiB for tensor buffers and runtime overhead. The system can run on devices with as little as 512 MiB total RAM, though 1 GiB is recommended for comfortable operation alongside the operating system.

### Can Supertonic run on Raspberry Pi Zero?

While the repository demonstrates operation on Raspberry Pi 4 and e-readers with Cortex-A53/A72 processors, the **300 MiB memory footprint** may challenge the 512 MiB RAM limit of Raspberry Pi Zero when accounting for OS overhead. Performance would likely degrade due to the single-core CPU and slower memory bus, though the ONNX Runtime's CPU optimizations in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) might still permit offline processing with increased latency.

### What is the RTF (Real-Time Factor) for edge devices?

On **Raspberry Pi 4**, Supertonic achieves an RTF of approximately **0.5**, generating speech twice as fast as real-time. On **Cortex-A53 e-readers** like the Onyx Boox Go 6, the RTF is approximately **0.3**, still maintaining real-time performance. Values below 1.0 indicate the system can synthesize speech faster than it takes to play back, enabling live streaming applications.

### Does Supertonic require GPU acceleration?

No. Supertonic is explicitly designed for **CPU-only inference** using ONNX Runtime. The architecture avoids GPU dependencies, making it suitable for e-readers, Raspberry Pi devices, and other edge hardware lacking discrete graphics or CUDA support. The [`cpp/example_onnx.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/example_onnx.cpp) implementation demonstrates pure CPU memory management strategies optimized for ARM architectures.