How to Integrate External Whisper Models with OpenSuperWhisper

You can integrate external Whisper models with OpenSuperWhisper by downloading a .bin file into the app's whisper-models directory, initializing a MyWhisperContext via Whis.initWith(loader:params:), and passing that context to WhisperEngine to run transcription.

OpenSuperWhisper ships with a tiny default model (ggml-tiny.en.bin) bundled inside the app and copied to the user-specific models directory on first launch. The architecture is deliberately extensible, so any community or custom-trained Whisper binary can be adopted without recompiling the app. To integrate external Whisper models with OpenSuperWhisper, you store the binary, create a WhisperModelLoader, initialize a context, and hand it to the transcription engine.

Model Storage and Naming Conventions

OpenSuperWhisper expects model files to follow Whisper's ggml-* naming convention, such as ggml-large-v2.bin. The manager class WhisperModelManager creates and manages the folder at ~/Library/Application Support/<bundle-id>/whisper-models. In OpenSuperWhisper/WhisperModelManager.swift, this class also copies the bundled default model on first launch and provides APIs for downloading, canceling, and verifying external binaries.

The Role of WhisperModelManager

WhisperModelManager handles model directory creation, default model copying, downloading, cancellation, and existence checks. You can fetch an external model using the downloadModel(url:name:progressCallback:) method, which stores the file under the models directory and tracks it as an active download so it can be cancelled via cancelDownload(name:). This design lets you integrate external Whisper models with OpenSuperWhisper on demand without manual file manipulation.

Step-by-Step Integration

Step 1: Download or Copy the External Model

Use the shared WhisperModelManager to fetch a remote model directly:

let url = URL(string: "https://huggingface.co/yourUser/ggml-large-v2/resolve/main/ggml-large-v2.bin")!
WhisperModelManager.shared.downloadModel(url: url,
                                         name: "ggml-large-v2.bin") { progress in
    print("Download progress: \(progress * 100)%")
}

The downloadModel method automatically stores the file under the correct models directory and records it as an active download.

Step 2: Verify the Model Is Present

Before initializing the context, confirm the file is present using isModelDownloaded(name:):

let isReady = WhisperModelManager.shared.isModelDownloaded(name: "ggml-large-v2.bin")
guard isReady else { fatalError("Model not downloaded") }

Step 3: Create a Loader for the File

In OpenSuperWhisper/Whis/WhisperModelLoader.swift, the loader bridges Swift to the C whisper_model_loader struct. For a standard file-based model, no custom read, eof, or close callbacks are required:

import OpenSuperWhisper.Whis

let modelPath = WhisperModelManager.shared.modelsDirectory
                    .appendingPathComponent("ggml-large-v2.bin")
                    .path

let loader = WhisperModelLoader()

Step 4: Initialize the Whisper Context

The static method Whis.initWith(loader:params:) in OpenSuperWhisper/Whis/Whis.swift wraps the underlying C call whisper_init_with_params. Supply a WhisperContextParams instance to configure thread count, sampling strategy, and other runtime behavior:

var params = WhisperContextParams()
params.setNumThreads(4)
params.setSamplingStrategy(.greedy)

guard let context = Whis.initWith(loader: loader,
                                 params: params) else {
    fatalError("Failed to create Whisper context")
}

This produces a MyWhisperContext that is ready for transcription.

Step 5: Plug the Context into the Engine

Pass the resulting context to WhisperEngine, which drives the full audio-to-text pipeline. The engine in OpenSuperWhisper/Engines/WhisperEngine.swift performs mel-generation, encoding, and decoding without requiring any model-specific changes:

let engine = WhisperEngine(context: context)

engine.transcribe(audioFileURL: myAudioURL) { result in
    switch result {
    case .success(let transcription):
        print("🗣️  Transcription: \(transcription)")
    case .failure(let err):
        print("❗️ Error: \(err)")
    }
}

Because WhisperEngine depends only on the context object, all downstream processing works unchanged when you swap model binaries.

Step 6: Switch Models at Runtime

OpenSuperWhisper does not require an app restart to change models. The manager can hold several binaries simultaneously. To switch, create a new WhisperModelLoader pointing at the new file, rebuild the context with Whis.initWith(loader:params:), and instantiate a fresh WhisperEngine. This makes it straightforward to benchmark different model sizes or let users select accuracy versus speed on the fly.

Summary

  • Model storage: External models must be .bin files following the ggml-* convention and stored in ~/Library/Application Support/<bundle-id>/whisper-models, managed by WhisperModelManager.
  • Downloading: Use WhisperModelManager.shared.downloadModel(url:name:progressCallback:) to fetch remote binaries and isModelDownloaded(name:) to confirm availability.
  • Loading: Create a WhisperModelLoader from WhisperModelLoader.swift to bridge Swift and the C library; file-based models require no custom callbacks.
  • Context creation: Call Whis.initWith(loader:params:) in Whis.swift to generate a MyWhisperContext, configuring WhisperContextParams for threads and sampling strategy.
  • Transcription: Pass the context to WhisperEngine in WhisperEngine.swift to execute the unchanged PCM-to-text pipeline.
  • Runtime swaps: You can hot-swap models by rebuilding the loader, context, and engine without restarting the application.

Frequently Asked Questions

What file format do external Whisper models need for OpenSuperWhisper?

OpenSuperWhisper expects standard Whisper .bin binaries that follow the ggml-* naming convention, such as ggml-large-v2.bin. The WhisperModelManager scans the whisper-models directory for these files and validates their presence with isModelDownloaded(name:).

Do I need to recompile OpenSuperWhisper to use a custom model?

No. The app is architected so that any external Whisper model binary can be loaded at runtime. You simply place the file in the models directory, create a new WhisperModelLoader, and reinitialize the context and engine as shown in the integration steps above.

How do I configure thread count and sampling strategy for a new model?

Use the WhisperContextParams struct before calling Whis.initWith(loader:params:). You can set threads with setNumThreads(_:) and choose a strategy such as .greedy or .beam_search via setSamplingStrategy(_:). These parameters are defined in WhisperContextParams.swift and passed directly to the native whisper_init_with_params C API.

Can I switch between multiple Whisper models while the app is running?

Yes. OpenSuperWhisper supports runtime model switching without requiring an app restart. Maintain multiple binaries in the models directory, then instantiate a new WhisperModelLoader pointing to the desired file, call Whis.initWith(loader:params:) to create a fresh context, and supply it to a new WhisperEngine instance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →