Can OpenSuperWhisper Handle Noisy Audio? Technical Architecture Explained
Yes, OpenSuperWhisper handles noisy audio through a multi-layered pipeline that combines Silero VAD filtering, blank audio suppression, dynamic decoding parameters, and signal normalization to extract speech from cluttered recordings.
OpenSuperWhisper is a Swift-based transcription engine built on top of OpenAI's Whisper models. According to the Starmel/OpenSuperWhisper source code, the app implements several noise-mitigation strategies directly in the core transcription pipeline, making it capable of processing audio with background chatter, music, or low-quality input without requiring manual preprocessing.
Voice Activity Detection with Silero VAD
The first line of defense against noise is a dedicated Voice Activity Detection (VAD) step that runs before the Whisper encoder processes any audio.
In OpenSuperWhisper/Engines/WhisperEngine.swift, the detectSpeech(in:) method loads a bundled Silero VAD model (ggml-silero-v5.1.2.bin) and extracts only valid speech segments from the PCM buffer. If no speech is detected, the function returns an empty string immediately, preventing the decoder from processing pure noise.
// WhisperEngine.swift – VAD gate prevents noise from reaching the encoder
private func detectSpeech(in samples: [Float]) throws -> [WhisperVadSegment] {
// Loads vadModelPath and filters samples
}
This VAD gate removes long pauses and silence sections that often cause hallucinated text in standard Whisper implementations.
Blank Audio Suppression
OpenSuperWhisper exposes a user-facing "Suppress blank audio" toggle that maps directly to the Whisper C-library's suppressBlank parameter. When enabled via Settings, the engine discards frames containing only background noise.
In WhisperEngine.swift, the configuration is applied as follows:
params.suppressBlank = settings.suppressBlankAudio // Default: true
The WhisperFullParams struct in OpenSuperWhisper/Whis/WhisperFullParams.swift defines this mapping, ensuring that spurious tokens generated from non-speech audio are filtered at the decoding stage.
Dynamic Thresholds and Temperature Control
For fine-tuning noise tolerance, the engine exposes two critical parameters through the Settings class:
- No-Speech Threshold (
noSpeechThreshold): Controls how aggressively the VAD discards low-energy frames - Temperature (
temperature): Adjusts decoder randomness to handle uncertain, noisy patterns
These are applied in WhisperEngine.swift at lines 138-141:
params.noSpeechThold = Float(settings.noSpeechThreshold) // e.g., 0.6 default
params.temperature = Float(settings.temperature) // e.g., 0.0 default
Lowering the no-speech threshold retains more frames (useful for quiet speech in noisy environments), while increasing temperature allows the decoder to explore alternative transcriptions when the audio signal is ambiguous.
Beam Search for Uncertain Audio
When dealing with particularly noisy inputs, enabling Beam Search forces the decoder to evaluate multiple transcription paths simultaneously and select the most probable sequence.
The configuration in WhisperEngine.swift (lines 174-176) implements this:
if settings.useBeamSearch {
params.beamSearchBeamSize = Int32(settings.beamSize)
}
Setting beamSize to values between 3 and 5 significantly improves accuracy on cluttered recordings by avoiding greedy decoding errors caused by isolated noise bursts.
Audio Preprocessing and Normalization
Before any neural processing occurs, the convertAudioToPCM routine normalizes the input signal to improve the signal-to-noise ratio. The pipeline performs three critical operations:
- Converts all input to 16 kHz mono float buffers
- Detects and removes silent channels from multi-channel recordings
- Mixes stereo sources and normalizes amplitude via
appendMixedSamples
This preprocessing happens in WhisperEngine.swift (lines 63-104), ensuring that the VAD and Whisper encoder receive consistent, clean waveforms regardless of the source file's original format or quality.
Practical Configuration Examples
Default Noise-Robust Transcription
The default settings optimize for general noisy environments without requiring manual tuning:
import OpenSuperWhisper
let url = URL(fileURLWithPath: "/path/to/noisy-recording.wav")
let settings = Settings() // suppressBlank = true, temperature = 0.0
Task {
do {
let text = try await TranscriptionService.shared
.transcribeAudio(url: url, settings: settings)
print("Transcription:", text)
} catch {
print("Failed:", error)
}
}
Aggressive Noise Handling
For recordings with heavy background interference, increase the temperature and enable beam search:
var noisySettings = Settings()
noisySettings.suppressBlankAudio = true
noisySettings.noSpeechThreshold = 0.4 // Lower to catch quiet speech
noisySettings.temperature = 0.2 // Increase for pattern flexibility
noisySettings.useBeamSearch = true
noisySettings.beamSize = 5
Task {
let result = try await TranscriptionService.shared
.transcribeAudio(url: noisyFileURL, settings: noisySettings)
print(result)
}
Process Cancellation for Excessive Noise
The engine supports thread-safe cancellation via an AbortFlag if transcription hangs on extremely degraded audio:
// Set flag to cancel ongoing transcription
abortFlag.isSet = true // Defined in WhisperEngine.swift
Summary
- Silero VAD filtering in
WhisperEngine.swiftstrips non-speech segments before they reach the Whisper encoder - Blank suppression via
suppressBlankremoves frames containing only background noise - Configurable thresholds (
noSpeechThreshold,temperature) allow runtime tuning for specific noise profiles - Beam search decoding explores multiple paths to avoid errors from isolated noise bursts
- 16 kHz normalization and channel mixing in
convertAudioToPCMstandardize input quality
Frequently Asked Questions
Does OpenSuperWhisper work with background music?
Yes. The Silero VAD model specifically targets human speech patterns, allowing it to distinguish vocals from instrumental background music. Combined with the suppressBlank parameter, the engine can transcribe speech even when music plays simultaneously, though accuracy improves if you lower the noSpeechThreshold to retain quieter vocal segments.
What VAD model does OpenSuperWhisper use?
OpenSuperWhisper bundles the Silero VAD model (ggml-silero-v5.1.2.bin) and loads it statically in WhisperEngine.swift. This model runs locally on device and processes the audio buffer before any Whisper inference occurs, ensuring no network latency for noise filtering.
How do I tune OpenSuperWhisper for very noisy recordings?
For challenging audio, enable useBeamSearch with a beamSize of 5, increase temperature to 0.2, and lower noSpeechThreshold to 0.4. These settings, configured through the Settings class, instruct the VAD to keep marginal speech frames while allowing the decoder to explore alternative transcriptions for ambiguous sounds.
Can I cancel transcription mid-process if the audio is too noisy?
Yes. The WhisperEngine implements an AbortFlag that checks for cancellation requests during decoding. Setting abortFlag.isSet = true immediately terminates the transcription thread without blocking the UI, preventing the app from hanging on excessively long or noisy files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →