What Does the Denoising Steps (`total-step`) Parameter Control in Supertonic?

The total-step parameter controls the number of diffusion denoising iterations the Supertonic TTS model executes per synthesis request, directly trading inference speed for audio fidelity with a default of 8 steps.

The total-step parameter (exposed as --total-step in CLI tools) is a critical configuration option in the supertone-inc/supertonic repository that governs synthesis quality in this diffusion-based text-to-speech engine. As implemented across the ONNX inference examples, this integer value determines how many times the model refines its latent audio representation during the reverse diffusion process.

How total-step Works in Supertonic

The Diffusion Denoising Process

Supertonic utilizes a diffusion model architecture that generates speech by iteratively denoising a latent representation. The total-step value sets the number of these iterative refinement cycles:

  • Higher values (e.g., 10–12): More denoising iterations produce clearer speech, better prosody, and fewer artifacts.
  • Lower values (e.g., 6–8): Fewer iterations reduce computational overhead and latency but may introduce subtle quality degradation.

According to the source documentation across language-specific READMEs, the parameter is described as controlling the "Number of denoising steps (higher = better quality, slower)".

Usage Examples Across Language Bindings

The total-step argument is implemented consistently across all Supertonic language bindings. Below are practical implementations from the official examples.

Python Implementation

In py/example_onnx.py (line 27), the argument is parsed and passed to the inference session:

python example_onnx.py --total-step 10 \
                       --voice-style ../assets/voice_styles/M1.json \
                       --text "Increasing the number of denoising steps improves the output's fidelity."

Swift Implementation

The Swift CLI example supports the same interface:

.build/release/example_onnx \
    --total-step 10 \
    --voice-style ../assets/voice_styles/M1.json \
    --text "Higher denoising steps give higher quality audio."

Node.js Implementation

In nodejs/example_onnx.js (line 35), the parameter is handled identically:

node example_onnx.js --total-step 12 \
                     --voice-style ../assets/voice_styles/F1.json \
                     --text "More steps = better quality."

Java Implementation

The Java example follows the same convention:

java -jar target/tts-example.jar \
    --total-step 10 \
    --text "Higher denoising steps improve fidelity."

Quality vs. Speed Trade-off

The total-step parameter represents a direct quality-latency trade-off in the Supertonic inference pipeline:

  • Audio Quality: Increasing steps from the default 8 to 10 or 12 yields measurable improvements in voice clarity and naturalness, as the diffusion model has more opportunities to refine the audio latent.
  • Inference Latency: Each additional step requires another full pass through the denoising network, linearly increasing GPU/CPU compute time and memory bandwidth usage.

The documentation in swift/README.md (line 116) and py/README.md explicitly notes this relationship under Quality vs Speed sections, recommending users tune this value based on real-time requirements versus fidelity needs.

Source File References

The --total-step argument is documented and implemented across the following repository locations:

Each implementation maintains the default value of 8 steps and accepts integer values to increase or decrease the denoising intensity.

Summary

  • total-step controls the number of diffusion denoising iterations in Supertonic's TTS synthesis pipeline.
  • Default value is 8 steps, balancing quality and speed for most use cases.
  • Higher values (10–12) improve audio fidelity and reduce artifacts at the cost of increased inference time and computational resource usage.
  • Implementation is consistent across all language bindings (Python, Swift, Node.js, Java, etc.) via the --total-step CLI argument.
  • Configuration is handled in the respective example_onnx source files and documented in each language's README.

Frequently Asked Questions

What is the default value for total-step in Supertonic?

The default value is 8 denoising steps. This default is hardcoded across all language-specific CLI examples in the repository, including py/example_onnx.py, nodejs/example_onnx.js, and the corresponding README documentation in swift/README.md and py/README.md.

How does increasing denoising steps affect GPU usage?

Each additional denoising step requires another forward pass through the diffusion model network, resulting in linearly increased GPU computation and memory bandwidth consumption. Raising the value from 8 to 12 steps increases inference time by approximately 50%, making this parameter the primary lever for trading off real-time performance against synthesis quality.

Can I set total-step to values lower than 8?

Yes, values below 8 are technically valid and will reduce latency further, though the official documentation and examples default to 8 as the minimum recommended threshold for acceptable audio quality. Setting values too low (e.g., 4 or fewer) may result in audible artifacts and degraded prosody as the diffusion process terminates before full convergence.

Which Supertonic source files handle the total-step argument parsing?

The CLI argument parsing is implemented in the example inference scripts: py/example_onnx.py (line 27), nodejs/example_onnx.js (line 35), and their Swift, Java, Rust, Go, C#, C++, and Flutter counterparts. Documentation for the parameter appears in each language-specific README (e.g., swift/README.md line 99, java/README.md) under the options table and Quality vs Speed sections.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →