What Is the No-Speech Threshold in OpenSuperWhisper and How Is It Used?

The no-speech threshold is a floating-point value between 0 and 1 that tells the underlying Whisper engine what probability of silence is acceptable before discarding an audio segment, effectively preventing hallucinated transcriptions on background noise.

OpenSuperWhisper is a Swift-based macOS application that wraps the whisper.cpp inference engine to provide local speech-to-text transcription. The no-speech threshold serves as a critical quality-control gate that determines when the model should treat audio as pure silence rather than attempt to decode it, ensuring cleaner output especially in noisy environments.

Understanding the No-Speech Threshold

Definition and Range

The threshold is defined in WhisperFullParams.swift as a Float property named noSpeechThold that accepts values from 0.0 to 1.0. It represents the probability cutoff at which the decoder considers a segment to contain no speech whatsoever. The default value shipped with OpenSuperWhisper is 0.6, meaning any segment with a greater than 60% probability of being silence is discarded.

Mapping to the Core Engine

This Swift property maps one-to-one to the C struct member no_speech_thold in the underlying whisper_full_params structure used by the whisper.cpp library. When transcription begins, the WhisperFullParams struct converts this value using its toC() method before passing it to the native whisper_full_with_state API.

How the No-Speech Threshold Flows Through the Application

User Configuration in Settings.swift

The UI exposes the threshold through a slider control defined in Settings.swift (lines 103-106). When users adjust this slider, the value binds to viewModel.noSpeechThreshold, which persists the setting via AppPreferences.shared.noSpeechThreshold using UserDefaults (as implemented in AppPreferences.swift, lines 78-79).

Runtime Application in WhisperEngine.swift

When a transcription job starts, WhisperEngine.swift (lines 38-41) copies the stored value from the Settings object into the WhisperFullParams instance before invoking the Whisper decoder. This ensures the user-configured cutoff is applied to the current decoding session.

Decoding Behavior

During inference, the Whisper engine calculates a no-speech probability for each potential segment. If this probability exceeds the configured noSpeechThold, the segment is omitted from the final transcript. This mechanism prevents "hallucinated" text—random words generated from pure background audio or silence—and reduces unnecessary timestamps in the output.

Configuring the No-Speech Threshold

You can adjust the threshold through the Settings UI or programmatically depending on your integration needs.

Adjusting via the Settings UI

The SwiftUI slider in Settings.swift provides a draggable control for the value:

// Settings.swift – binding the slider to the view model
Slider(value: $viewModel.noSpeechThreshold, in: 0.0...1.0, step: 0.1) {
    Text("No‑Speech Threshold")
}

When the slider moves, viewModel.noSpeechThreshold updates AppPreferences.shared.noSpeechThreshold, persisting the value across app launches (lines 112-119).

Setting the Threshold Programmatically

To override the default at runtime or configure it before the UI loads:

// Somewhere in your code (e.g., on launch)
import OpenSuperWhisper

AppPreferences.shared.noSpeechThreshold = 0.75   // Stricter silence detection

Passing the Threshold to Whisper

During transcription setup, the engine initializes the parameter struct:

// WhisperEngine.swift – building the WhisperFullParams
var params = WhisperFullParams()
params.noSpeechThold = Float(settings.noSpeechThreshold)   // Uses value from Settings

The params struct is later converted to the C representation (toC()) and handed to whisper_full_with_state.

Inspecting the Default Value

The model wrapper defines the fallback value directly:

// WhisperFullParams.swift – default initialization
public var noSpeechThold: Float = 0.6   // Default when user has not changed it

Summary

  • The no-speech threshold is a Float (0.0–1.0) that acts as a probability cutoff for silence detection, with a default value of 0.6 in OpenSuperWhisper.
  • It is defined in WhisperFullParams.swift and maps directly to the C struct member no_speech_thold in whisper.cpp.
  • Users configure it via a slider in Settings.swift, which persists through AppPreferences.swift using UserDefaults.
  • At runtime, WhisperEngine.swift copies the value into the parameter struct before calling whisper_full_with_state.
  • Segments with a no-speech probability exceeding the threshold are discarded, preventing hallucinations on background noise.

Frequently Asked Questions

What happens if I set the no-speech threshold to 0.0?

Setting the threshold to 0.0 effectively disables the filter, forcing Whisper to attempt transcription of every audio segment regardless of silence probability. This often results in hallucinated words during quiet portions or background noise.

Why is the default no-speech threshold set to 0.6?

The default value of 0.6 provides a balanced trade-off that filters most background noise while preserving quiet or whispered speech, based on the probability distribution of the Whisper model's noise detection capabilities.

Where is the no-speech threshold stored between app launches?

The value is persisted in UserDefaults through the AppPreferences class, specifically via the noSpeechThreshold property defined in AppPreferences.swift (lines 78-79).

Does the no-speech threshold affect word-level timestamps?

Yes, by filtering out silent segments before they reach the final transcript, the threshold indirectly improves timestamp accuracy. It removes false-positive text alignments that would otherwise occur when the model attempts to synchronize hallucinated words with non-speech audio segments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →