How Beam Search Decoding Improves Transcription Quality in OpenSuperWhisper
Beam search decoding improves transcription quality by exploring multiple token sequences in parallel and selecting the highest-scoring complete hypothesis, reducing errors from premature commits to low-probability tokens.
OpenSuperWhisper leverages the Whisper.cpp inference engine to transcribe audio, offering both greedy and beam search decoding strategies. While greedy decoding selects the single most likely token at each step, beam search decoding maintains multiple candidate sequences simultaneously, resulting in more accurate transcripts especially for noisy audio or complex sentences.
What Is Beam Search Decoding?
Beam search is a heuristic search algorithm that explores a graph by expanding the most promising nodes in a limited set (the "beam"). In the context of speech recognition, it works by keeping track of N parallel hypotheses—called beams—rather than committing to one token at a time. Each beam accumulates a log-probability score over the entire sequence, and the decoder ultimately outputs the beam with the highest overall score.
OpenSuperWhisper exposes this functionality through the WhisperSamplingStrategy enum, which defines two options:
.greedy– Selects the top token at each step (fast but prone to local maxima).beamSearch– Maintains multiple beams and selects the globally best path
How OpenSuperWhisper Implements Beam Search
The implementation spans several Swift files that bridge the UI preferences to the underlying C++ Whisper engine.
Configuration in Settings.swift
User preferences are stored in AppPreferences and exposed through the settings UI. The @Published properties useBeamSearch and beamSize control the decoding behavior:
@Published var useBeamSearch: Bool
@Published var beamSize: Int
These values persist across sessions and feed directly into the transcription engine. The relevant UI controls appear in Settings.swift, where users toggle beam search and adjust the beam width from 1 to 10.
Engine Configuration in WhisperEngine.swift
When transcription begins, WhisperEngine.transcribeAudio constructs a WhisperFullParams struct and sets the strategy based on user settings:
var params = WhisperFullParams()
params.strategy = settings.useBeamSearch ? .beamSearch : .greedy
if settings.useBeamSearch {
params.beamSearchBeamSize = Int32(settings.beamSize)
}
This logic in WhisperEngine.swift (lines 124-176) determines whether the decoder explores multiple hypotheses or follows a single greedy path.
Parameter Conversion in WhisperFullParams.swift
The Swift wrapper converts these parameters to the C-style whisper_full_params structure required by Whisper.cpp. The toC() method maps the Swift properties to the underlying C struct:
cParams.beam_search.beam_size = beamSearchBeamSize
This conversion in WhisperFullParams.swift (lines 122-124) ensures that the beam width configured in the UI reaches the core Whisper engine. The beamSearchBeamSize parameter defaults to 1 but can be increased to improve accuracy at the cost of computation time.
Why Beam Search Improves Transcription Quality
According to the OpenSuperWhisper source code, beam search enhances quality through four key mechanisms:
-
Explores multiple hypotheses – Instead of committing to the highest-probability token immediately, the decoder maintains N parallel beams. This prevents early rejection of correct words that might appear less likely in isolation but fit better in the complete sequence.
-
Global sequence scoring – Each beam accumulates log-probability across the entire token sequence. The final output selects the hypothesis with the highest cumulative score, better reflecting the true likelihood of whole sentences rather than individual tokens.
-
Robustness to noise and ambiguity – For audio with background noise or homophones (words that sound alike), greedy decoding may select incorrect tokens. Beam search can recover by maintaining alternative paths where later tokens boost the overall likelihood of the correct interpretation.
-
Configurable accuracy trade-off – The
beamSearchBeamSizeparameter allows users to balance quality against performance. Beam sizes of 3–5 typically provide significant accuracy improvements over greedy decoding without excessive latency.
Configuring Beam Search in Your Code
You can enable beam search programmatically or through the SwiftUI interface.
Enable via SwiftUI Settings
Toggle("Use Beam Search", isOn: $viewModel.useBeamSearch)
.help("Beam search can provide better results but is slower")
if viewModel.useBeamSearch {
Stepper("Beam Size: \(viewModel.beamSize)",
value: $viewModel.beamSize,
in: 1...10)
.help("Number of beams to use in beam search")
}
Programmatic Transcription
let settings = Settings()
settings.useBeamSearch = true
settings.beamSize = 4
let engine = WhisperEngine()
await engine.initialize()
let text = try await engine.transcribeAudio(url: audioURL, settings: settings)
Direct Parameter Configuration
var params = WhisperFullParams()
params.strategy = .beamSearch
params.beamSearchBeamSize = 5
let cParams = params.toC() // Passed to whisper.cpp
Summary
- Beam search decoding in OpenSuperWhisper explores multiple token sequences simultaneously through the
WhisperSamplingStrategy.beamSearchoption. - The implementation passes the
beamSearchBeamSizeparameter fromWhisperFullParams.swiftto the underlying Whisper.cpp engine via thetoC()method. - Quality improvements include fewer dropped words, better punctuation, and robust handling of long sentences compared to greedy decoding.
- Users control the trade-off between accuracy and speed through the
beamSizesetting inSettings.swift, with typical values of 3–5 providing optimal results. - The transcription engine applies these settings in
WhisperEngine.swiftwhen constructing theWhisperFullParamsstruct.
Frequently Asked Questions
What is the default beam size in OpenSuperWhisper?
The default beamSearchBeamSize is set to 1, which effectively behaves like greedy decoding. To see quality improvements, users should increase this value to between 3 and 5 through the settings UI or programmatically via the beamSize parameter.
How does beam search affect transcription speed?
Beam search increases CPU usage proportionally to the beam size because the decoder must evaluate and score multiple hypotheses simultaneously. In ProgressContext.swift, this manifests as a slower progress curve compared to greedy decoding, though the trade-off is typically worthwhile for accuracy-critical applications.
When should I use greedy decoding instead of beam search?
Use greedy decoding (.greedy) when speed is prioritized over absolute accuracy, such as for real-time transcription on low-power devices. Use beam search decoding when processing recorded audio where accuracy matters most, particularly for content with background noise, multiple speakers, or complex vocabulary.
Can I change the beam size dynamically for different transcription tasks?
Yes. Since useBeamSearch and beamSize are @Published properties in Settings.swift, you can modify these values programmatically before calling transcribeAudio. Create different Settings instances or modify the shared preferences object to apply different beam widths for different audio files or user scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →