Can OpenSuperWhisper Handle Noisy Audio? Technical Architecture Explained

Yes, OpenSuperWhisper handles noisy audio through a multi-layered pipeline that combines Silero VAD filtering, blank audio suppression, dynamic decoding parameters, and signal normalization to extract speech from cluttered recordings.

OpenSuperWhisper is a Swift-based transcription engine built on top of OpenAI's Whisper models. According to the Starmel/OpenSuperWhisper source code, the app implements several noise-mitigation strategies directly in the core transcription pipeline, making it capable of processing audio with background chatter, music, or low-quality input without requiring manual preprocessing.

Voice Activity Detection with Silero VAD

The first line of defense against noise is a dedicated Voice Activity Detection (VAD) step that runs before the Whisper encoder processes any audio.

In OpenSuperWhisper/Engines/WhisperEngine.swift, the detectSpeech(in:) method loads a bundled Silero VAD model (ggml-silero-v5.1.2.bin) and extracts only valid speech segments from the PCM buffer. If no speech is detected, the function returns an empty string immediately, preventing the decoder from processing pure noise.

// WhisperEngine.swift – VAD gate prevents noise from reaching the encoder
private func detectSpeech(in samples: [Float]) throws -> [WhisperVadSegment] { 
    // Loads vadModelPath and filters samples
}

This VAD gate removes long pauses and silence sections that often cause hallucinated text in standard Whisper implementations.

Blank Audio Suppression

OpenSuperWhisper exposes a user-facing "Suppress blank audio" toggle that maps directly to the Whisper C-library's suppressBlank parameter. When enabled via Settings, the engine discards frames containing only background noise.

In WhisperEngine.swift, the configuration is applied as follows:

params.suppressBlank = settings.suppressBlankAudio   // Default: true

The WhisperFullParams struct in OpenSuperWhisper/Whis/WhisperFullParams.swift defines this mapping, ensuring that spurious tokens generated from non-speech audio are filtered at the decoding stage.

Dynamic Thresholds and Temperature Control

For fine-tuning noise tolerance, the engine exposes two critical parameters through the Settings class:

  • No-Speech Threshold (noSpeechThreshold): Controls how aggressively the VAD discards low-energy frames
  • Temperature (temperature): Adjusts decoder randomness to handle uncertain, noisy patterns

These are applied in WhisperEngine.swift at lines 138-141:

params.noSpeechThold = Float(settings.noSpeechThreshold)   // e.g., 0.6 default
params.temperature   = Float(settings.temperature)        // e.g., 0.0 default

Lowering the no-speech threshold retains more frames (useful for quiet speech in noisy environments), while increasing temperature allows the decoder to explore alternative transcriptions when the audio signal is ambiguous.

Beam Search for Uncertain Audio

When dealing with particularly noisy inputs, enabling Beam Search forces the decoder to evaluate multiple transcription paths simultaneously and select the most probable sequence.

The configuration in WhisperEngine.swift (lines 174-176) implements this:

if settings.useBeamSearch {
    params.beamSearchBeamSize = Int32(settings.beamSize)
}

Setting beamSize to values between 3 and 5 significantly improves accuracy on cluttered recordings by avoiding greedy decoding errors caused by isolated noise bursts.

Audio Preprocessing and Normalization

Before any neural processing occurs, the convertAudioToPCM routine normalizes the input signal to improve the signal-to-noise ratio. The pipeline performs three critical operations:

  1. Converts all input to 16 kHz mono float buffers
  2. Detects and removes silent channels from multi-channel recordings
  3. Mixes stereo sources and normalizes amplitude via appendMixedSamples

This preprocessing happens in WhisperEngine.swift (lines 63-104), ensuring that the VAD and Whisper encoder receive consistent, clean waveforms regardless of the source file's original format or quality.

Practical Configuration Examples

Default Noise-Robust Transcription

The default settings optimize for general noisy environments without requiring manual tuning:

import OpenSuperWhisper

let url = URL(fileURLWithPath: "/path/to/noisy-recording.wav")
let settings = Settings()  // suppressBlank = true, temperature = 0.0

Task {
    do {
        let text = try await TranscriptionService.shared
                         .transcribeAudio(url: url, settings: settings)
        print("Transcription:", text)
    } catch {
        print("Failed:", error)
    }
}

Aggressive Noise Handling

For recordings with heavy background interference, increase the temperature and enable beam search:

var noisySettings = Settings()
noisySettings.suppressBlankAudio = true
noisySettings.noSpeechThreshold = 0.4    // Lower to catch quiet speech
noisySettings.temperature = 0.2          // Increase for pattern flexibility
noisySettings.useBeamSearch = true
noisySettings.beamSize = 5

Task {
    let result = try await TranscriptionService.shared
                     .transcribeAudio(url: noisyFileURL, settings: noisySettings)
    print(result)
}

Process Cancellation for Excessive Noise

The engine supports thread-safe cancellation via an AbortFlag if transcription hangs on extremely degraded audio:

// Set flag to cancel ongoing transcription
abortFlag.isSet = true   // Defined in WhisperEngine.swift

Summary

  • Silero VAD filtering in WhisperEngine.swift strips non-speech segments before they reach the Whisper encoder
  • Blank suppression via suppressBlank removes frames containing only background noise
  • Configurable thresholds (noSpeechThreshold, temperature) allow runtime tuning for specific noise profiles
  • Beam search decoding explores multiple paths to avoid errors from isolated noise bursts
  • 16 kHz normalization and channel mixing in convertAudioToPCM standardize input quality

Frequently Asked Questions

Does OpenSuperWhisper work with background music?

Yes. The Silero VAD model specifically targets human speech patterns, allowing it to distinguish vocals from instrumental background music. Combined with the suppressBlank parameter, the engine can transcribe speech even when music plays simultaneously, though accuracy improves if you lower the noSpeechThreshold to retain quieter vocal segments.

What VAD model does OpenSuperWhisper use?

OpenSuperWhisper bundles the Silero VAD model (ggml-silero-v5.1.2.bin) and loads it statically in WhisperEngine.swift. This model runs locally on device and processes the audio buffer before any Whisper inference occurs, ensuring no network latency for noise filtering.

How do I tune OpenSuperWhisper for very noisy recordings?

For challenging audio, enable useBeamSearch with a beamSize of 5, increase temperature to 0.2, and lower noSpeechThreshold to 0.4. These settings, configured through the Settings class, instruct the VAD to keep marginal speech frames while allowing the decoder to explore alternative transcriptions for ambiguous sounds.

Can I cancel transcription mid-process if the audio is too noisy?

Yes. The WhisperEngine implements an AbortFlag that checks for cancellation requests during decoding. Setting abortFlag.isSet = true immediately terminates the transcription thread without blocking the UI, preventing the app from hanging on excessively long or noisy files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →