# Performance Tradeoffs of the total_steps Parameter (5-12) in Supertonic TTS

> Explore total_steps parameter tradeoffs in Supertonic TTS. Learn how increasing steps enhances audio quality while impacting inference latency, GPU load, and memory. Optimize your performance.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: performance
- Published: 2026-06-14

---

**Increasing `total_steps` from 5 to 12 in Supertonic's S-I2 text-to-speech pipeline improves audio quality through finer denoising iterations, but each additional step linearly increases inference latency, GPU compute load, and memory consumption.**

The `total_steps` parameter in the supertone-inc/supertonic repository controls the diffusion denoising process in the S-I2 neural text-to-speech model. Understanding the specific performance tradeoffs between values 5 and 12 helps developers optimize for real-time mobile applications versus high-fidelity offline rendering.

## How total_steps Controls the Diffusion Process

In Supertonic’s architecture, `total_steps` defines the number of iterations the model runs to transform random noise into coherent speech. The default value is **8**, stored in the private variable `_totalSteps` in `flutter/lib/main.dart`, which is then passed through `TextToSpeech.call` → `_infer` and consumed by the denoising loop in `flutter/lib/helper.dart`.

Inside `flutter/lib/helper.dart` (lines 84–95), the core inference logic iterates exactly `totalStep` times, running the vector estimation model (`vectorEstOrt.run`) at each step:

```dart
for (var step = 0; step < totalStep; step++) {
  final result = await vectorEstOrt.run({
    // ... tensor inputs ...
    'total_step': totalStepTensor,
    'current_step': await _scalarToTensor(List.filled(bsz, step.toDouble()), [bsz]),
  });
  // Process denoised audio output
}

```

Each iteration refines the latent audio representation, meaning the 5–12 range represents a spectrum from coarse to fine-grained synthesis.

## Quality vs. Latency Tradeoffs: 5 Steps vs. 12 Steps

The choice between 5 and 12 steps creates distinct performance profiles:

- **5 Steps (Minimum Recommended):** Completes inference rapidly with minimal compute overhead. Audio output may exhibit slight hiss or reduced naturalness in complex phonemes. Suitable for real-time previews or low-power devices.
- **8 Steps (Default):** Balances quality and speed according to the Supertonic source code. Provides natural speech with acceptable latency for interactive applications.
- **12 Steps (Maximum Practical):** Delivers the highest fidelity with minimal artifacts and crisp articulation. Generation latency increases roughly 2.4× compared to 5 steps, making this suitable for offline batch processing.

## Computational Resource Impact

Raising `total_steps` from 5 to 12 affects hardware utilization across four key dimensions:

**GPU and CPU Utilization**
Each denoising step executes the full vector estimation model. Higher step counts increase GPU memory traffic and compute time proportionally. On CPU-only devices, the cost escalates further due to less parallelized tensor operations.

**Memory Footprint**
The loop in `flutter/lib/helper.dart` allocates intermediate tensors during every iteration. At 12 steps, peak RAM usage is significantly higher than at 5 steps, potentially limiting concurrent processing or maximum input length on memory-constrained mobile devices.

**Battery and Thermal Constraints**
Additional compute steps increase power draw. Running at 12 steps versus 5 generates substantially more heat and drains battery faster, a critical consideration for the Flutter mobile implementation in `flutter/lib/main.dart`.

## Implementation Details

### Configuring the UI in main.dart

The Flutter demo exposes `total_steps` as a user-adjustable slider (lines 272–286), allowing runtime tuning between the 5–12 range:

```dart
Slider(
  min: 1,
  max: 20,
  divisions: 19,
  value: _totalSteps.toDouble(),
  label: _totalSteps.toString(),
  onChanged: (value) => setState(() => _totalSteps = value.toInt()),
),

```

### Passing Parameters to the Inference Engine

When invoking the TTS engine at line 125 of `flutter/lib/main.dart`, the `_totalSteps` value flows directly into the synthesis pipeline:

```dart
final result = await _textToSpeech!.call(
  _textController.text,
  _selectedLang,
  _style!,
  _totalSteps,          // Controls quality vs. speed tradeoff
  speed: _speed,
);

```

## Selecting the Right Value for Production

Choose your `total_steps` value based on deployment constraints:

- **Real-time mobile apps:** Use **5–6 steps** to maintain responsive UI and conserve battery.
- **Standard production:** Use the default **8 steps** for balanced quality across device tiers.
- **Offline high-fidelity:** Use **10–12 steps** when latency is irrelevant and audio quality is paramount.

## Summary

- **`total_steps`** (5–12) controls denoising iterations in Supertonic's S-I2 diffusion model, directly impacting output quality and computational cost.
- **Higher values** (10–12) produce superior audio fidelity but increase latency linearly and raise GPU memory usage, limiting mobile performance.
- **Lower values** (5–6) enable real-time synthesis on constrained devices but may introduce minor artifacts.
- The default **8 steps** in `flutter/lib/main.dart` represents an optimal balance for general use cases.
- The parameter is implemented in `flutter/lib/helper.dart` via a deterministic loop that executes the vector estimation model at each step.

## Frequently Asked Questions

### What is the optimal total_steps value for mobile devices?

For battery-powered mobile deployment using the Flutter SDK, **5–6 steps** provides the best balance of intelligible speech and responsive performance. This configuration minimizes GPU thermal throttling and extends battery life while maintaining acceptable quality for most applications.

### How does total_steps affect GPU memory usage?

Each incremental step in the 5–12 range increases peak memory allocation because `flutter/lib/helper.dart` retains intermediate tensors for the duration of the denoising loop. At 12 steps, memory pressure can become significant when processing long utterances, potentially triggering garbage collection or out-of-memory errors on devices with limited RAM.

### Why does total_steps impact battery consumption?

The diffusion model executes the computationally expensive `vectorEstOrt.run` operation at every step. Increasing from 5 to 12 steps requires the CPU or GPU to perform 2.4× more work, directly translating to higher power draw and increased heat generation, which impacts battery longevity during extended TTS sessions.

### Can I use total_steps values outside the 5-12 range?

While the UI implementation in `flutter/lib/main.dart` technically supports values from 1 to 20, the **5–12 range** represents the practical operational window. Values below 5 produce noticeably degraded audio with artifacts, while values above 12 yield diminishing quality returns relative to the substantial linear increase in latency and compute cost.