# Can OpenSuperWhisper Handle Noisy Audio? Technical Architecture Explained

> Discover how OpenSuperWhisper tackles noisy audio using advanced filtering, dynamic decoding, and signal normalization. Learn its technical architecture for clear speech extraction.

- Repository: [Starmel/OpenSuperWhisper](https://github.com/Starmel/OpenSuperWhisper)
- Tags: technical-architecture
- Published: 2026-07-07

---

**Yes, OpenSuperWhisper handles noisy audio through a multi-layered pipeline that combines Silero VAD filtering, blank audio suppression, dynamic decoding parameters, and signal normalization to extract speech from cluttered recordings.**

OpenSuperWhisper is a Swift-based transcription engine built on top of OpenAI's Whisper models. According to the Starmel/OpenSuperWhisper source code, the app implements several noise-mitigation strategies directly in the core transcription pipeline, making it capable of processing audio with background chatter, music, or low-quality input without requiring manual preprocessing.

## Voice Activity Detection with Silero VAD

The first line of defense against noise is a dedicated **Voice Activity Detection (VAD)** step that runs before the Whisper encoder processes any audio.

In [`OpenSuperWhisper/Engines/WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Engines/WhisperEngine.swift), the `detectSpeech(in:)` method loads a bundled **Silero VAD model** (`ggml-silero-v5.1.2.bin`) and extracts only valid speech segments from the PCM buffer. If no speech is detected, the function returns an empty string immediately, preventing the decoder from processing pure noise.

```swift
// WhisperEngine.swift – VAD gate prevents noise from reaching the encoder
private func detectSpeech(in samples: [Float]) throws -> [WhisperVadSegment] { 
    // Loads vadModelPath and filters samples
}

```

This VAD gate removes long pauses and silence sections that often cause hallucinated text in standard Whisper implementations.

## Blank Audio Suppression

OpenSuperWhisper exposes a user-facing **"Suppress blank audio"** toggle that maps directly to the Whisper C-library's `suppressBlank` parameter. When enabled via `Settings`, the engine discards frames containing only background noise.

In [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift), the configuration is applied as follows:

```swift
params.suppressBlank = settings.suppressBlankAudio   // Default: true

```

The `WhisperFullParams` struct in [`OpenSuperWhisper/Whis/WhisperFullParams.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Whis/WhisperFullParams.swift) defines this mapping, ensuring that spurious tokens generated from non-speech audio are filtered at the decoding stage.

## Dynamic Thresholds and Temperature Control

For fine-tuning noise tolerance, the engine exposes two critical parameters through the `Settings` class:

- **No-Speech Threshold** (`noSpeechThreshold`): Controls how aggressively the VAD discards low-energy frames
- **Temperature** (`temperature`): Adjusts decoder randomness to handle uncertain, noisy patterns

These are applied in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) at lines 138-141:

```swift
params.noSpeechThold = Float(settings.noSpeechThreshold)   // e.g., 0.6 default
params.temperature   = Float(settings.temperature)        // e.g., 0.0 default

```

Lowering the no-speech threshold retains more frames (useful for quiet speech in noisy environments), while increasing temperature allows the decoder to explore alternative transcriptions when the audio signal is ambiguous.

## Beam Search for Uncertain Audio

When dealing with particularly noisy inputs, enabling **Beam Search** forces the decoder to evaluate multiple transcription paths simultaneously and select the most probable sequence.

The configuration in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) (lines 174-176) implements this:

```swift
if settings.useBeamSearch {
    params.beamSearchBeamSize = Int32(settings.beamSize)
}

```

Setting `beamSize` to values between 3 and 5 significantly improves accuracy on cluttered recordings by avoiding greedy decoding errors caused by isolated noise bursts.

## Audio Preprocessing and Normalization

Before any neural processing occurs, the `convertAudioToPCM` routine normalizes the input signal to improve the signal-to-noise ratio. The pipeline performs three critical operations:

1. Converts all input to **16 kHz mono float buffers**
2. Detects and removes silent channels from multi-channel recordings
3. Mixes stereo sources and normalizes amplitude via `appendMixedSamples`

This preprocessing happens in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) (lines 63-104), ensuring that the VAD and Whisper encoder receive consistent, clean waveforms regardless of the source file's original format or quality.

## Practical Configuration Examples

### Default Noise-Robust Transcription

The default settings optimize for general noisy environments without requiring manual tuning:

```swift
import OpenSuperWhisper

let url = URL(fileURLWithPath: "/path/to/noisy-recording.wav")
let settings = Settings()  // suppressBlank = true, temperature = 0.0

Task {
    do {
        let text = try await TranscriptionService.shared
                         .transcribeAudio(url: url, settings: settings)
        print("Transcription:", text)
    } catch {
        print("Failed:", error)
    }
}

```

### Aggressive Noise Handling

For recordings with heavy background interference, increase the temperature and enable beam search:

```swift
var noisySettings = Settings()
noisySettings.suppressBlankAudio = true
noisySettings.noSpeechThreshold = 0.4    // Lower to catch quiet speech
noisySettings.temperature = 0.2          // Increase for pattern flexibility
noisySettings.useBeamSearch = true
noisySettings.beamSize = 5

Task {
    let result = try await TranscriptionService.shared
                     .transcribeAudio(url: noisyFileURL, settings: noisySettings)
    print(result)
}

```

### Process Cancellation for Excessive Noise

The engine supports thread-safe cancellation via an `AbortFlag` if transcription hangs on extremely degraded audio:

```swift
// Set flag to cancel ongoing transcription
abortFlag.isSet = true   // Defined in WhisperEngine.swift

```

## Summary

- **Silero VAD filtering** in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) strips non-speech segments before they reach the Whisper encoder
- **Blank suppression** via `suppressBlank` removes frames containing only background noise
- **Configurable thresholds** (`noSpeechThreshold`, `temperature`) allow runtime tuning for specific noise profiles
- **Beam search decoding** explores multiple paths to avoid errors from isolated noise bursts
- **16 kHz normalization** and channel mixing in `convertAudioToPCM` standardize input quality

## Frequently Asked Questions

### Does OpenSuperWhisper work with background music?

Yes. The Silero VAD model specifically targets human speech patterns, allowing it to distinguish vocals from instrumental background music. Combined with the `suppressBlank` parameter, the engine can transcribe speech even when music plays simultaneously, though accuracy improves if you lower the `noSpeechThreshold` to retain quieter vocal segments.

### What VAD model does OpenSuperWhisper use?

OpenSuperWhisper bundles the **Silero VAD** model (`ggml-silero-v5.1.2.bin`) and loads it statically in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift). This model runs locally on device and processes the audio buffer before any Whisper inference occurs, ensuring no network latency for noise filtering.

### How do I tune OpenSuperWhisper for very noisy recordings?

For challenging audio, enable `useBeamSearch` with a `beamSize` of 5, increase `temperature` to 0.2, and lower `noSpeechThreshold` to 0.4. These settings, configured through the `Settings` class, instruct the VAD to keep marginal speech frames while allowing the decoder to explore alternative transcriptions for ambiguous sounds.

### Can I cancel transcription mid-process if the audio is too noisy?

Yes. The `WhisperEngine` implements an `AbortFlag` that checks for cancellation requests during decoding. Setting `abortFlag.isSet = true` immediately terminates the transcription thread without blocking the UI, preventing the app from hanging on excessively long or noisy files.