How to Extend or Modify the Functionality of the Whisper Engine in OpenSuperWhisper
Extend the Whisper engine by modifying the Swift wrapper in Whis.swift, adjusting parameters in Settings.swift, or implementing the TranscriptionEngine protocol to create custom processing pipelines.
OpenSuperWhisper provides a modular macOS transcription framework built atop the whisper.cpp C library. To extend or modify the functionality of the Whisper engine, developers work through three distinct architectural layers: the low-level C wrapper, the high-level engine implementation, and the configuration interface.
Understanding the Three-Tier Architecture
The codebase separates concerns into a C-wrapper, an engine implementation, and a configuration layer.
The C Wrapper Layer (MyWhisperContext)
At the foundation, [Whis.swift](https://github.com/Starmel/OpenSuperWhisper/blob/master/OpenSuperWhisper/Whis/Whis.swift) contains MyWhisperContext, which provides a thin Swift façade over the native whisper.cpp API. This class manages the whisper_context pointer, handles mel-spectrogram generation through pcmToMel, and performs token encoding and decoding. It maintains both a primary ctx and an optional state for stateful inference scenarios.
The Engine Implementation (WhisperEngine)
The [WhisperEngine.swift](https://github.com/Starmel/OpenSuperWhisper/blob/master/OpenSuperWhisper/Engines/WhisperEngine.swift) file implements the TranscriptionEngine protocol, orchestrating the complete transcription workflow. The engine initializes the model via MyWhisperContext.initFromFile, converts audio to PCM format, maps Settings values to WhisperFullParams, and bridges C-side progress callbacks through ProgressContext.
The Configuration Layer (Settings)
User preferences are centralized in [Settings.swift](https://github.com/Starmel/OpenSuperWhisper/blob/master/OpenSuperWhisper/Settings.swift), which defines all configurable parameters—from language selection to beam search strategies. These values flow directly into the WhisperFullParams structure during the transcription initialization phase.
Extension Strategies
You can modify the engine through several pathways depending on your use case.
Adding Custom Whisper Parameters
To expose additional whisper.cpp parameters (such as max_len), extend the configuration and mapping pipeline:
- Extend
Settingsby adding a new property with a default value:
// Settings.swift
var maxTokenLength: Int = 225
- Map the property in
WhisperEngine.transcribeAudiowhereparamsare constructed:
// WhisperEngine.swift
params.max_len = Int32(settings.maxTokenLength)
- Update the UI in
SettingsViewto expose a control bound to this property.
Modifying Audio Preprocessing
The convertAudioToPCM method in [WhisperEngine.swift](https://github.com/Starmel/OpenSuperWhisper/blob/master/OpenSuperWhisper/Engines/WhisperEngine.swift) handles audio format conversion. Insert custom DSP logic before the PCM buffer reaches the Whisper context:
private func applyNoiseReduction(to samples: [Float]) -> [Float] {
let alpha: Float = 0.9
var previous: Float = 0
var out = [Float]()
out.reserveCapacity(samples.count)
for s in samples {
let filtered = s - alpha * previous
out.append(filtered)
previous = s
}
return out
}
// Usage inside convertAudioToPCM
let cleanSamples = applyNoiseReduction(to: samples)
Creating Custom Transcription Engines
For radically different behavior (such as streaming or alternative backends), implement the TranscriptionEngine protocol:
// StreamingEngine.swift
class StreamingEngine: TranscriptionEngine {
var engineName: String { "Streaming Whisper" }
private var context: MyWhisperContext?
var onProgressUpdate: ((Float) -> Void)?
func initialize() async throws {
let path = AppPreferences.shared.selectedWhisperModelPath!
let params = WhisperContextParams()
guard let ctx = MyWhisperContext.initFromFile(path: path, params: params) else {
throw TranscriptionError.contextInitializationFailed
}
context = ctx
}
func processChunk(_ pcm: [Float]) throws -> String {
guard let ctx = context else { throw TranscriptionError.contextInitializationFailed }
// Implement encode/decode logic for streaming buffers
return ""
}
func cancelTranscription() { /* Set abort flag */ }
func getSupportedLanguages() -> [String] { LanguageUtil.availableLanguages }
}
Register the new engine in the UI layer where WhisperEngine is currently instantiated.
Implementing Post-Processing Hooks
To modify transcription output (e.g., adding punctuation restoration or text normalization), hook into the result assembly phase in WhisperEngine.transcribeAudio. After the text segments are collected but before returning, apply your transformations:
let rawText = // ... assembled from segments
let processedText = applyCustomPostProcessing(rawText)
return processedText
Summary
- Architecture: The engine consists of
MyWhisperContext(C wrapper),WhisperEngine(orchestration), andSettings(configuration). - Parameter Extension: Add fields to
Settingsand map them toWhisperFullParamsin the engine initialization. - Audio Pipeline: Override
convertAudioToPCMto insert custom DSP filters before inference. - Custom Engines: Conform to
TranscriptionEngineto implement alternative transcription strategies or streaming support. - Model Management: Extend [
WhisperModelManager.swift](https://github.com/Starmel/OpenSuperWhisper/blob/master/OpenSuperWhisper/WhisperModelManager.swift) to support new model formats or download sources.
Frequently Asked Questions
How do I add a new parameter to the Whisper engine?
Add the property to the Settings struct in [Settings.swift](https://github.com/Starmel/OpenSuperWhisper/blob/master/OpenSuperWhisper/Settings.swift), then map it to the corresponding field in WhisperFullParams within [WhisperEngine.swift](https://github.com/Starmel/OpenSuperWhisper/blob/master/OpenSuperWhisper/Engines/WhisperEngine.swift) during the transcription setup. Expose the control in the settings UI to make it user-configurable.
Can I replace the audio preprocessing pipeline?
Yes. Modify the convertAudioToPCM method in [WhisperEngine.swift](https://github.com/Starmel/OpenSuperWhisper/blob/master/OpenSuperWhisper/Engines/WhisperEngine.swift) to insert custom DSP steps—such as noise reduction, normalization, or format conversion—before the PCM data is passed to MyWhisperContext.full.
How do I create a completely new transcription engine?
Create a new class that conforms to the TranscriptionEngine protocol, implementing initialize(), transcribeAudio(), cancelTranscription(), and getSupportedLanguages(). You can reuse MyWhisperContext for low-level operations or implement a custom backend entirely.
Where are language models managed?
Model discovery, downloading, and validation are handled in [WhisperModelManager.swift](https://github.com/Starmel/OpenSuperWhisper/blob/master/OpenSuperWhisper/WhisperModelManager.swift). To support new model formats, update the download URL validation logic and ensure the model file compatibility checks align with your new format requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →