OpenSuperWhisper Example Usage: Complete Guide to Real-Time Transcription
TLDR: OpenSuperWhisper provides three primary usage patterns: global keyboard shortcuts for instant recording, drag-and-drop batch file processing, and direct Swift API integration via the WhisperEngine class.
OpenSuperWhisper is a macOS GUI application that runs the Whisper transcription engine locally for real-time speech-to-text. Available as a Homebrew package and a full Xcode project on GitHub, the repository contains Swift classes implementing the complete transcription pipeline. This guide demonstrates concrete OpenSuperWhisper example usage extracted directly from the source code, covering installation, interactive workflows, and programmatic embedding.
Installation and Setup
You can install the application via Homebrew or build from source using the provided scripts.
Homebrew installation (recommended):
brew update
brew install opensuperwhisper
Building from source requires the helper script run.sh at the repository root, which configures libwhisper, builds the autocorrect-swift library, and compiles the Xcode project:
# Clone the repository
git clone https://github.com/Starmel/OpenSuperWhisper.git
cd OpenSuperWhisper
# Execute the build script
./run.sh
The build script references libwhisper/CMakeLists.txt for model compilation and generates the OpenSuperWhisper.app bundle with its entry point in OpenSuperWhisper/ContentView.swift.
Interactive Usage Examples
These patterns require no code and utilize the global shortcut system defined in OpenSuperWhisper/Settings.swift.
Real-Time Recording with Global Shortcuts
When the app launches, it registers system-wide hotkeys. By default, pressing left ⌘, right ⌥, or Fn starts recording on key-down and stops on key-up. Alternatively, configure a mouse button (middle-click or thumb button) in the Settings pane.
When a shortcut triggers, the app instantiates a Recording object from OpenSuperWhisper/Models/Recording.swift and begins capturing microphone input. The AudioRecorder.swift class streams PCM frames to the transcription engine until the key is released.
Batch Transcription via Drag-and-Drop
Process existing audio files without manual recording:
- Drag any compatible audio file (WAV, MP3, M4A) onto the OpenSuperWhisper window
- The app wraps the file path in a
Recordinginstance - The file enters
TranscriptionQueue.swift, which processes items sequentially using the sharedWhisperEngineinstance
// Simplified internal logic from TranscriptionQueue.swift
let fileURL = URL(fileURLWithPath: "/Users/me/meeting.m4a")
let recorder = Recording(fileURL: fileURL)
TranscriptionQueue.shared.enqueue(recorder)
The queue automatically handles conversion to PCM via AudioUtil.swift and emits results through the same publisher pipeline used for live recording.
Programmatic Usage Example
For embedding transcription in your own Swift projects, import the Engines module and instantiate WhisperEngine directly. This example demonstrates transcribing a local audio file programmatically:
import OpenSuperWhisper
// 1️⃣ Load a Whisper model (e.g., ggml-tiny.en.bin)
let modelURL = Bundle.main.url(forResource: "ggml-tiny.en", withExtension: "bin")!
let engine = try WhisperEngine(
modelURL: modelURL,
language: .english,
temperature: 0.0
)
// 2️⃣ Load audio into PCM buffer
let audioURL = URL(fileURLWithPath: "/path/to/audio.wav")
let pcmData = try AudioUtil.loadPCM(from: audioURL)
// 3️⃣ Feed audio to the engine
try engine.appendAudio(pcmData)
// 4️⃣ Retrieve final transcription
let result = try engine.finalResult()
print("🗣️ Transcription:", result.text)
Key source files referenced:
OpenSuperWhisper/Engines/WhisperEngine.swift– Concrete implementation of theTranscriptionEngineprotocolOpenSuperWhisper/Utils/AudioUtil.swift– Helper for loading PCM data from common audio containersOpenSuperWhisper/WhisperModelManager.swift– Logic for locating and validating model binaries
Core Architecture Components
Understanding these classes helps when extending the example usage patterns:
TranscriptionEngine.swift – Defines the protocol that all transcription backends implement, specifying methods for audio appending and result retrieval.
WhisperEngine.swift – The primary implementation that wraps the C++ Whisper bindings from libwhisper. It manages model state, language detection, and emits partial results through a Combine publisher.
TranscriptionService.swift – Acts as the coordinator between the UI layer (ContentView.swift), the audio recorder, and the engine. It handles instantiation of WhisperEngine via WhisperModelManager.swift.
AudioRecorder.swift – Captures microphone samples in real-time and forwards them to the active engine instance.
Summary
- Install via
brew install opensuperwhisperor build locally usingrun.shwhich configureslibwhisperand the Xcode project - Record interactively using global shortcuts defined in
OpenSuperWhisper/Settings.swift, which triggerRecordingobjects and stream audio throughAudioRecorder.swift - Process files by dragging them onto the app window;
TranscriptionQueue.swiftmanages sequential processing through the shared engine - Integrate programmatically by importing the
Enginespackage and instantiatingWhisperEnginewith a model fromWhisperModelManager.swift - Customize language and model settings through the UI or by modifying
WhisperModelManager.swiftto load custom.binmodels
Frequently Asked Questions
How do I install OpenSuperWhisper without building from source?
Use Homebrew. Run brew install opensuperwhisper to download the precompiled binary and all dependencies. This avoids compiling the libwhisper C++ libraries and the autocorrect-swift helper manually.
Can I use OpenSuperWhisper as a library in my own Swift project?
Yes. The WhisperEngine class in OpenSuperWhisper/Engines/WhisperEngine.swift provides a public API for embedding transcription. Import the module, initialize the engine with a model URL from WhisperModelManager.swift, and feed PCM data using appendAudio(_:) to receive transcription results.
Where are the global keyboard shortcuts configured?
Shortcut definitions and modifier key mappings reside in OpenSuperWhisper/Settings.swift. You can customize single-modifier keys (like left ⌘ or right ⌥) and mouse buttons through the Settings UI, or modify the Settings struct directly in the source code to change default bindings.
What audio formats are supported for batch transcription?
The TranscriptionQueue.swift handler accepts standard formats including WAV, MP3, and M4A. The Recording model passes these files to AudioUtil.swift, which converts them to PCM buffers before feeding them into the WhisperEngine for processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →