How to Customize Whisper Engine Settings in OpenSuperWhisper: A Complete Configuration Guide

OpenSuperWhisper exposes every aspect of the Whisper transcription engine through the Settings model defined in OpenSuperWhisper/Settings.swift, allowing you to configure language detection, decoding strategies, model parameters, and output formatting either through the SwiftUI preferences panel or programmatically via the transcribeAudio(url:settings:) method.

OpenSuperWhisper provides granular control over Whisper engine settings through a comprehensive configuration system built around the Settings struct. This macOS application maps Swift properties directly to the underlying C library parameters defined in WhisperFullParams, enabling precise tuning of transcription behavior without modifying core engine code. Whether you need to adjust beam search width for higher accuracy or suppress blank audio segments for cleaner output, the repository supports both UI-driven and code-first configuration approaches.

Understanding the Configuration Architecture

The configuration system centers on two primary components that work together to translate user preferences into low-level Whisper parameters. The Settings struct in OpenSuperWhisper/Settings.swift defines all user-configurable parameters with sensible defaults and persistence logic, while WhisperEngine in OpenSuperWhisper/Engines/WhisperEngine.swift consumes these settings during the transcription pipeline initialization. When transcribeAudio(url:settings:) is invoked, the engine maps Swift properties to the WhisperFullParams C struct (defined in OpenSuperWhisper/Whis/WhisperFullParams.swift) before executing the native transcription routines.

Available Whisper Engine Customizations

Language Detection and Translation Control

You can specify the source language explicitly or enable automatic detection using the selectedLanguage and translateToEnglish properties. In WhisperEngine.swift (lines 22-24), these values configure the underlying model to either transcribe in the original language or translate non-English speech to English automatically.

Output Formatting Options

Control transcript presentation through showTimestamps and suppressBlankAudio. These boolean flags map to params.noTimestamps and params.suppressBlank within WhisperEngine.swift (lines 19-21), allowing you to include precise timecodes in the output and optionally drop silent segments from the results.

Optimize for speed or accuracy using the useBeamSearch and beamSize parameters. When enabled, the engine sets params.strategy to beam search and configures params.beamSearchBeamSize (lines 17-18 and 58-59 of WhisperEngine.swift), which significantly improves transcription quality at the cost of processing speed.

Model Inference Parameters

Fine-tune the neural network behavior through temperature, noSpeechThreshold, and initialPrompt. These map directly to params.temperature, params.noSpeechThold, and params.initialPrompt (lines 25-27 of WhisperEngine.swift), controlling randomness in decoding, voice activity detection sensitivity, and providing contextual hints to the model.

Asian Language Autocorrection

For Chinese, Japanese, and Korean transcription workflows, enable useAsianAutocorrect to apply built-in post-processing. This setting is checked via Settings.shouldApplyAsianAutocorrect (line 19 of Settings.swift) and automatically corrects common recognition errors in CJK languages after the initial transcription completes.

Programmatic Configuration Without UI

To customize Whisper engine settings directly in Swift code without using the graphical interface, instantiate a Settings object and pass it to the transcription method. This approach creates the configuration (equivalent to lines 5-18 in the snippet below) and passes it directly to transcribeAudio(url:settings:) at line 79 of WhisperEngine.swift:

import OpenSuperWhisper

var custom = Settings()
custom.selectedLanguage = "en"
custom.translateToEnglish = false
custom.showTimestamps = true
custom.suppressBlankAudio = true
custom.temperature = 0.3
custom.noSpeechThreshold = 0.2
custom.initialPrompt = "Please transcribe clearly."
custom.useBeamSearch = true
custom.beamSize = 5
custom.useAsianAutocorrect = false

Task {
    let engine = WhisperEngine()
    try await engine.initialize()
    let text = try await engine.transcribeAudio(
        url: URL(fileURLWithPath: "/path/to/audio.wav"),
        settings: custom
    )
    print("Result:", text)
}

Customizing Settings Through the User Interface

The default workflow uses SettingsView in OpenSuperWhisper/Settings.swift to present a tabbed preferences panel. SwiftUI controls bind to AppPreferences.shared, which persists user choices across application launches. The Settings initializer (lines 28-40 of Settings.swift) automatically reads these persisted preferences when constructing new instances:

Picker("Language", selection: $viewModel.selectedLanguage) {
    ForEach(LanguageUtil.availableLanguages, id: \.self) { code in
        Text(LanguageUtil.languageNames[code] ?? code).tag(code)
    }
}

When the user clicks Done, the view invokes TranscriptionService.shared.reloadModel(with:) (lines 85-91 of Settings.swift), ensuring the engine picks up new parameters immediately without requiring an application restart.

Building Reusable Configuration Helpers

For applications requiring consistent settings across multiple transcription calls, extend WhisperEngine with a convenience wrapper that reads from AppPreferences.shared. This pattern eliminates manual Settings construction while honoring the current UI preferences:

extension WhisperEngine {
    func transcribeCurrentAudio(_ file: URL) async throws -> String {
        let prefs = AppPreferences.shared
        var settings = Settings()
        settings.selectedLanguage = prefs.whisperLanguage
        settings.translateToEnglish = prefs.translateToEnglish
        settings.suppressBlankAudio = prefs.suppressBlankAudio
        settings.showTimestamps = prefs.showTimestamps
        settings.temperature = prefs.temperature
        settings.noSpeechThreshold = prefs.noSpeechThreshold
        settings.initialPrompt = prefs.initialPrompt
        settings.useBeamSearch = prefs.useBeamSearch
        settings.beamSize = prefs.beamSize
        settings.useAsianAutocorrect = prefs.useAsianAutocorrect
        return try await transcribeAudio(url: file, settings: settings)
    }
}

Summary

  • OpenSuperWhisper centralizes Whisper configuration in the Settings struct located in OpenSuperWhisper/Settings.swift.
  • The WhisperEngine.transcribeAudio(url:settings:) method applies these settings by mapping them to WhisperFullParams before invoking the C library.
  • You can customize language detection, beam search decoding, temperature controls, timestamp output, and Asian language autocorrection.
  • Configuration works both programmatically via Swift code and through the SettingsView SwiftUI interface.
  • Changes made in the UI persist through AppPreferences and take effect when the model is reloaded via TranscriptionService.

Frequently Asked Questions

Can I change Whisper settings without restarting the application?

Yes. When you modify settings through the SettingsView interface and click Done, the application calls TranscriptionService.shared.reloadModel(with:) to apply changes immediately. For programmatic changes, simply pass a new Settings instance to the next transcribeAudio call; the engine uses the provided configuration without requiring reinitialization.

Which file contains the actual mapping between Swift settings and the Whisper C library parameters?

The mapping occurs in OpenSuperWhisper/Engines/WhisperEngine.swift, specifically within the transcribeAudio method. Lines 17-27 construct the WhisperFullParams struct by reading properties from the Settings object passed to the function, then pass these parameters to the native whisper_full function.

Is it possible to adjust the number of CPU threads used for transcription?

Thread count is calculated automatically based on ProcessInfo.activeProcessorCount at line 14 of WhisperEngine.swift. While this isn't exposed as a user-configurable setting in the current UI, you could modify the nThreads calculation in the source code if you need to limit CPU usage for background transcription tasks.

How do I enable beam search for better transcription accuracy?

Set useBeamSearch to true and specify a beamSize value (typically 5) in your Settings instance. The engine configures params.strategy and params.beamSearchBeamSize accordingly in WhisperEngine.swift (lines 17-18 and 58-59), though this will increase processing time compared to the default greedy decoding strategy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →