How to Integrate External Whisper Models with OpenSuperWhisper
You can integrate external Whisper models with OpenSuperWhisper by downloading a .bin file into the app's whisper-models directory, initializing a MyWhisperContext via Whis.initWith(loader:params:), and passing that context to WhisperEngine to run transcription.
OpenSuperWhisper ships with a tiny default model (ggml-tiny.en.bin) bundled inside the app and copied to the user-specific models directory on first launch. The architecture is deliberately extensible, so any community or custom-trained Whisper binary can be adopted without recompiling the app. To integrate external Whisper models with OpenSuperWhisper, you store the binary, create a WhisperModelLoader, initialize a context, and hand it to the transcription engine.
Model Storage and Naming Conventions
OpenSuperWhisper expects model files to follow Whisper's ggml-* naming convention, such as ggml-large-v2.bin. The manager class WhisperModelManager creates and manages the folder at ~/Library/Application Support/<bundle-id>/whisper-models. In OpenSuperWhisper/WhisperModelManager.swift, this class also copies the bundled default model on first launch and provides APIs for downloading, canceling, and verifying external binaries.
The Role of WhisperModelManager
WhisperModelManager handles model directory creation, default model copying, downloading, cancellation, and existence checks. You can fetch an external model using the downloadModel(url:name:progressCallback:) method, which stores the file under the models directory and tracks it as an active download so it can be cancelled via cancelDownload(name:). This design lets you integrate external Whisper models with OpenSuperWhisper on demand without manual file manipulation.
Step-by-Step Integration
Step 1: Download or Copy the External Model
Use the shared WhisperModelManager to fetch a remote model directly:
let url = URL(string: "https://huggingface.co/yourUser/ggml-large-v2/resolve/main/ggml-large-v2.bin")!
WhisperModelManager.shared.downloadModel(url: url,
name: "ggml-large-v2.bin") { progress in
print("Download progress: \(progress * 100)%")
}
The downloadModel method automatically stores the file under the correct models directory and records it as an active download.
Step 2: Verify the Model Is Present
Before initializing the context, confirm the file is present using isModelDownloaded(name:):
let isReady = WhisperModelManager.shared.isModelDownloaded(name: "ggml-large-v2.bin")
guard isReady else { fatalError("Model not downloaded") }
Step 3: Create a Loader for the File
In OpenSuperWhisper/Whis/WhisperModelLoader.swift, the loader bridges Swift to the C whisper_model_loader struct. For a standard file-based model, no custom read, eof, or close callbacks are required:
import OpenSuperWhisper.Whis
let modelPath = WhisperModelManager.shared.modelsDirectory
.appendingPathComponent("ggml-large-v2.bin")
.path
let loader = WhisperModelLoader()
Step 4: Initialize the Whisper Context
The static method Whis.initWith(loader:params:) in OpenSuperWhisper/Whis/Whis.swift wraps the underlying C call whisper_init_with_params. Supply a WhisperContextParams instance to configure thread count, sampling strategy, and other runtime behavior:
var params = WhisperContextParams()
params.setNumThreads(4)
params.setSamplingStrategy(.greedy)
guard let context = Whis.initWith(loader: loader,
params: params) else {
fatalError("Failed to create Whisper context")
}
This produces a MyWhisperContext that is ready for transcription.
Step 5: Plug the Context into the Engine
Pass the resulting context to WhisperEngine, which drives the full audio-to-text pipeline. The engine in OpenSuperWhisper/Engines/WhisperEngine.swift performs mel-generation, encoding, and decoding without requiring any model-specific changes:
let engine = WhisperEngine(context: context)
engine.transcribe(audioFileURL: myAudioURL) { result in
switch result {
case .success(let transcription):
print("🗣️ Transcription: \(transcription)")
case .failure(let err):
print("❗️ Error: \(err)")
}
}
Because WhisperEngine depends only on the context object, all downstream processing works unchanged when you swap model binaries.
Step 6: Switch Models at Runtime
OpenSuperWhisper does not require an app restart to change models. The manager can hold several binaries simultaneously. To switch, create a new WhisperModelLoader pointing at the new file, rebuild the context with Whis.initWith(loader:params:), and instantiate a fresh WhisperEngine. This makes it straightforward to benchmark different model sizes or let users select accuracy versus speed on the fly.
Summary
- Model storage: External models must be
.binfiles following theggml-*convention and stored in~/Library/Application Support/<bundle-id>/whisper-models, managed byWhisperModelManager. - Downloading: Use
WhisperModelManager.shared.downloadModel(url:name:progressCallback:)to fetch remote binaries andisModelDownloaded(name:)to confirm availability. - Loading: Create a
WhisperModelLoaderfromWhisperModelLoader.swiftto bridge Swift and the C library; file-based models require no custom callbacks. - Context creation: Call
Whis.initWith(loader:params:)inWhis.swiftto generate aMyWhisperContext, configuringWhisperContextParamsfor threads and sampling strategy. - Transcription: Pass the context to
WhisperEngineinWhisperEngine.swiftto execute the unchanged PCM-to-text pipeline. - Runtime swaps: You can hot-swap models by rebuilding the loader, context, and engine without restarting the application.
Frequently Asked Questions
What file format do external Whisper models need for OpenSuperWhisper?
OpenSuperWhisper expects standard Whisper .bin binaries that follow the ggml-* naming convention, such as ggml-large-v2.bin. The WhisperModelManager scans the whisper-models directory for these files and validates their presence with isModelDownloaded(name:).
Do I need to recompile OpenSuperWhisper to use a custom model?
No. The app is architected so that any external Whisper model binary can be loaded at runtime. You simply place the file in the models directory, create a new WhisperModelLoader, and reinitialize the context and engine as shown in the integration steps above.
How do I configure thread count and sampling strategy for a new model?
Use the WhisperContextParams struct before calling Whis.initWith(loader:params:). You can set threads with setNumThreads(_:) and choose a strategy such as .greedy or .beam_search via setSamplingStrategy(_:). These parameters are defined in WhisperContextParams.swift and passed directly to the native whisper_init_with_params C API.
Can I switch between multiple Whisper models while the app is running?
Yes. OpenSuperWhisper supports runtime model switching without requiring an app restart. Maintain multiple binaries in the models directory, then instantiate a new WhisperModelLoader pointing to the desired file, call Whis.initWith(loader:params:) to create a fresh context, and supply it to a new WhisperEngine instance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →