How to Customize Whisper Engine Settings in OpenSuperWhisper: A Complete Configuration Guide
OpenSuperWhisper exposes every aspect of the Whisper transcription engine through the Settings model defined in OpenSuperWhisper/Settings.swift, allowing you to configure language detection, decoding strategies, model parameters, and output formatting either through the SwiftUI preferences panel or programmatically via the transcribeAudio(url:settings:) method.
OpenSuperWhisper provides granular control over Whisper engine settings through a comprehensive configuration system built around the Settings struct. This macOS application maps Swift properties directly to the underlying C library parameters defined in WhisperFullParams, enabling precise tuning of transcription behavior without modifying core engine code. Whether you need to adjust beam search width for higher accuracy or suppress blank audio segments for cleaner output, the repository supports both UI-driven and code-first configuration approaches.
Understanding the Configuration Architecture
The configuration system centers on two primary components that work together to translate user preferences into low-level Whisper parameters. The Settings struct in OpenSuperWhisper/Settings.swift defines all user-configurable parameters with sensible defaults and persistence logic, while WhisperEngine in OpenSuperWhisper/Engines/WhisperEngine.swift consumes these settings during the transcription pipeline initialization. When transcribeAudio(url:settings:) is invoked, the engine maps Swift properties to the WhisperFullParams C struct (defined in OpenSuperWhisper/Whis/WhisperFullParams.swift) before executing the native transcription routines.
Available Whisper Engine Customizations
Language Detection and Translation Control
You can specify the source language explicitly or enable automatic detection using the selectedLanguage and translateToEnglish properties. In WhisperEngine.swift (lines 22-24), these values configure the underlying model to either transcribe in the original language or translate non-English speech to English automatically.
Output Formatting Options
Control transcript presentation through showTimestamps and suppressBlankAudio. These boolean flags map to params.noTimestamps and params.suppressBlank within WhisperEngine.swift (lines 19-21), allowing you to include precise timecodes in the output and optionally drop silent segments from the results.
Decoding Strategy and Beam Search
Optimize for speed or accuracy using the useBeamSearch and beamSize parameters. When enabled, the engine sets params.strategy to beam search and configures params.beamSearchBeamSize (lines 17-18 and 58-59 of WhisperEngine.swift), which significantly improves transcription quality at the cost of processing speed.
Model Inference Parameters
Fine-tune the neural network behavior through temperature, noSpeechThreshold, and initialPrompt. These map directly to params.temperature, params.noSpeechThold, and params.initialPrompt (lines 25-27 of WhisperEngine.swift), controlling randomness in decoding, voice activity detection sensitivity, and providing contextual hints to the model.
Asian Language Autocorrection
For Chinese, Japanese, and Korean transcription workflows, enable useAsianAutocorrect to apply built-in post-processing. This setting is checked via Settings.shouldApplyAsianAutocorrect (line 19 of Settings.swift) and automatically corrects common recognition errors in CJK languages after the initial transcription completes.
Programmatic Configuration Without UI
To customize Whisper engine settings directly in Swift code without using the graphical interface, instantiate a Settings object and pass it to the transcription method. This approach creates the configuration (equivalent to lines 5-18 in the snippet below) and passes it directly to transcribeAudio(url:settings:) at line 79 of WhisperEngine.swift:
import OpenSuperWhisper
var custom = Settings()
custom.selectedLanguage = "en"
custom.translateToEnglish = false
custom.showTimestamps = true
custom.suppressBlankAudio = true
custom.temperature = 0.3
custom.noSpeechThreshold = 0.2
custom.initialPrompt = "Please transcribe clearly."
custom.useBeamSearch = true
custom.beamSize = 5
custom.useAsianAutocorrect = false
Task {
let engine = WhisperEngine()
try await engine.initialize()
let text = try await engine.transcribeAudio(
url: URL(fileURLWithPath: "/path/to/audio.wav"),
settings: custom
)
print("Result:", text)
}
Customizing Settings Through the User Interface
The default workflow uses SettingsView in OpenSuperWhisper/Settings.swift to present a tabbed preferences panel. SwiftUI controls bind to AppPreferences.shared, which persists user choices across application launches. The Settings initializer (lines 28-40 of Settings.swift) automatically reads these persisted preferences when constructing new instances:
Picker("Language", selection: $viewModel.selectedLanguage) {
ForEach(LanguageUtil.availableLanguages, id: \.self) { code in
Text(LanguageUtil.languageNames[code] ?? code).tag(code)
}
}
When the user clicks Done, the view invokes TranscriptionService.shared.reloadModel(with:) (lines 85-91 of Settings.swift), ensuring the engine picks up new parameters immediately without requiring an application restart.
Building Reusable Configuration Helpers
For applications requiring consistent settings across multiple transcription calls, extend WhisperEngine with a convenience wrapper that reads from AppPreferences.shared. This pattern eliminates manual Settings construction while honoring the current UI preferences:
extension WhisperEngine {
func transcribeCurrentAudio(_ file: URL) async throws -> String {
let prefs = AppPreferences.shared
var settings = Settings()
settings.selectedLanguage = prefs.whisperLanguage
settings.translateToEnglish = prefs.translateToEnglish
settings.suppressBlankAudio = prefs.suppressBlankAudio
settings.showTimestamps = prefs.showTimestamps
settings.temperature = prefs.temperature
settings.noSpeechThreshold = prefs.noSpeechThreshold
settings.initialPrompt = prefs.initialPrompt
settings.useBeamSearch = prefs.useBeamSearch
settings.beamSize = prefs.beamSize
settings.useAsianAutocorrect = prefs.useAsianAutocorrect
return try await transcribeAudio(url: file, settings: settings)
}
}
Summary
- OpenSuperWhisper centralizes Whisper configuration in the
Settingsstruct located inOpenSuperWhisper/Settings.swift. - The
WhisperEngine.transcribeAudio(url:settings:)method applies these settings by mapping them toWhisperFullParamsbefore invoking the C library. - You can customize language detection, beam search decoding, temperature controls, timestamp output, and Asian language autocorrection.
- Configuration works both programmatically via Swift code and through the
SettingsViewSwiftUI interface. - Changes made in the UI persist through
AppPreferencesand take effect when the model is reloaded viaTranscriptionService.
Frequently Asked Questions
Can I change Whisper settings without restarting the application?
Yes. When you modify settings through the SettingsView interface and click Done, the application calls TranscriptionService.shared.reloadModel(with:) to apply changes immediately. For programmatic changes, simply pass a new Settings instance to the next transcribeAudio call; the engine uses the provided configuration without requiring reinitialization.
Which file contains the actual mapping between Swift settings and the Whisper C library parameters?
The mapping occurs in OpenSuperWhisper/Engines/WhisperEngine.swift, specifically within the transcribeAudio method. Lines 17-27 construct the WhisperFullParams struct by reading properties from the Settings object passed to the function, then pass these parameters to the native whisper_full function.
Is it possible to adjust the number of CPU threads used for transcription?
Thread count is calculated automatically based on ProcessInfo.activeProcessorCount at line 14 of WhisperEngine.swift. While this isn't exposed as a user-configurable setting in the current UI, you could modify the nThreads calculation in the source code if you need to limit CPU usage for background transcription tasks.
How do I enable beam search for better transcription accuracy?
Set useBeamSearch to true and specify a beamSize value (typically 5) in your Settings instance. The engine configures params.strategy and params.beamSearchBeamSize accordingly in WhisperEngine.swift (lines 17-18 and 58-59), though this will increase processing time compared to the default greedy decoding strategy.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →