How to Use OpenSuperWhisper for Audio Transcription: A Complete Guide
OpenSuperWhisper is a macOS-native application that captures real-time microphone input, converts audio to Whisper-compatible PCM format, and performs on-device transcription using Whisper or FluidAudio engines.
OpenSuperWhisper, hosted in the Starmel/OpenSuperWhisper repository, provides a complete Swift implementation for local audio transcription using OpenAI's Whisper models. The application processes audio entirely on-device through a three-stage pipeline involving capture, conversion, and transcription, exposing a clean API for integration into macOS workflows.
Three-Stage Transcription Pipeline
The transcription workflow follows a structured pipeline implemented across AudioRecorder.swift and WhisperEngine.swift.
Audio Capture
The AudioRecorder class manages microphone access using a private serial queue (workQueue) to prevent UI blocking. It writes 16-bit PCM audio at 16 kHz using AVFormatIDKey, AVSampleRateKey, and AVLinearPCMBitDepthKey configurations.
Key features include:
- Temporary
.wavfile creation for each recording session - Automatic cleanup of old temporary files via
cleanupOldTemporaryFiles() - Optional start-notification sounds for audio feedback
- Support for global hotkeys, modifier keys, or mouse button triggers
Audio Conversion
Before transcription, WhisperEngine.convertAudioToPCM processes recorded or user-provided audio into the format required by Whisper models. The conversion implementation:
- Resamples audio to 16 kHz mono Float32 samples using
AVAudioConverter - Builds target formats via
makeTargetFormat - Parallelizes conversion across multiple cores for large files
This ensures compatibility regardless of source audio format or sample rate.
Transcription and Post-Processing
The WhisperEngine.transcribeAudio method loads the model from AppPreferences.shared.selectedWhisperModelPath and initializes a fresh decoding state for each session. The transcription pipeline includes:
- Voice Activity Detection: Uses the built-in Silero VAD model (
ggml-silero-v5.1.2.bin) to remove silence and prevent hallucinations - Progress Reporting: Maps Whisper's 0-100% callbacks to the app's 10-95% range via
ProgressContext - Token Filtering: Removes placeholder tokens including
[MUSIC]and[BLANK_AUDIO] - Language-Specific Formatting: Applies
AutocorrectWrapper.formatfor Asian language autocorrection when enabled
The UI layer in ContentView.swift binds to AudioRecorder.isRecording, AudioRecorder.isPlaying, and TranscriptionEngine progress callbacks to display live indicators and final transcripts.
Implementation Architecture
Understanding the internal structure enables proper integration and extension of OpenSuperWhisper functionality.
Recording System
AudioRecorder operates on a dedicated serial dispatch queue to maintain responsive UI during capture. The class handles microphone permissions, temporary file management, and playback functionality. Recordings persist as temporary .wav files until transcription completes or cleanup routines execute.
Model Management
WhisperModelManager.swift handles downloading and selecting Whisper model files. The engine initializes via WhisperEngine.initialize(), which loads the model specified in AppPreferences.shared.selectedWhisperModelPath. The TranscriptionEngine protocol abstraction allows seamless switching between WhisperEngine and FluidAudioEngine implementations.
Configuration Interface
Settings.swift stores user-configurable transcription parameters including language selection, beam search width, and temperature controls. These settings pass directly to the transcription engine via the Settings.shared singleton.
Code Examples
Basic Recording and Transcription Workflow
The following Swift code demonstrates the complete workflow from hotkey-triggered recording to final transcript:
// Start recording (triggered by global hotkey, modifier key, or mouse button)
AudioRecorder.shared.startRecording()
// Stop recording and transcribe
if let wavURL = await AudioRecorder.shared.stopRecording() {
let engine: TranscriptionEngine = WhisperEngine() // or FluidAudioEngine()
try await engine.initialize()
let transcript = try await engine.transcribeAudio(
url: wavURL,
settings: Settings.shared // Contains language, beam search, etc.
)
print("Transcript: \(transcript)")
}
Manual Audio Conversion
For testing or batch processing existing audio files:
let engine = WhisperEngine()
try await engine.initialize()
if let pcmSamples = try await engine.convertAudioToPCM(fileURL: someAudioURL) {
print("Converted to \(pcmSamples.count) Float32 samples")
}
Querying Supported Languages
Populate UI language pickers using the engine's capabilities:
let supported = WhisperEngine().getSupportedLanguages()
print("Supported languages: \(supported.joined(separator: ", "))")
Integrating Progress Callbacks
Monitor transcription progress for UI updates:
// Progress reported via ProgressContext
// Maps Whisper's internal 0-100% to application's 10-95% range
// Bind to UI progress indicators in ContentView.swift
Configuration Options
OpenSuperWhisper exposes several customization points through Settings.swift:
- Language Selection: Specify source language or enable auto-detection
- Beam Search Parameters: Configure beam width for decoding accuracy
- Temperature: Control randomness in token generation
- VAD Sensitivity: Adjust Silero VAD thresholds to filter silence
The ContentView.swift interface supports drag-and-drop functionality for queuing additional audio files after initial transcription.
Summary
- OpenSuperWhisper implements a complete on-device transcription pipeline using
AudioRecorder.swiftfor capture andWhisperEngine.swiftfor processing - Audio converts to 16 kHz mono Float32 PCM format before transcription, with optional multicore parallelization for large files
- Voice Activity Detection using the Silero model (
ggml-silero-v5.1.2.bin) filters silence to reduce hallucinations and improve accuracy - The
TranscriptionEngineprotocol supports both Whisper and FluidAudio backends through a unified interface - Progress reporting maps internal Whisper callbacks to the 10-95% UI range via
ProgressContextfor accurate progress indicators
Frequently Asked Questions
How does OpenSuperWhisper handle audio format conversion?
According to the Starmel/OpenSuperWhisper source code in WhisperEngine.swift, the convertAudioToPCM method uses AVAudioConverter with a target format built by makeTargetFormat to resample any input audio to 16 kHz mono Float32 samples. For large files, the conversion can run in parallel across multiple CPU cores to improve performance.
What is the purpose of the Silero VAD model in OpenSuperWhisper?
The ggml-silero-v5.1.2.bin model performs Voice Activity Detection before transcription begins. As implemented in WhisperEngine.swift, this preprocessing step identifies and removes silent segments from the audio stream using the context.full(samples:params:) method, significantly reducing transcription hallucinations and improving processing speed by skipping non-speech content.
Can I use OpenSuperWhisper with pre-recorded audio files?
Yes. While the default workflow captures microphone input via AudioRecorder.startRecording(), you can pass any audio file URL directly to WhisperEngine.transcribeAudio or manually convert files using convertAudioToPCM. The UI in ContentView.swift supports drag-and-drop functionality for queuing existing audio files, and the system processes them through the same conversion and transcription pipeline.
How does the application prevent UI freezing during recording?
AudioRecorder.swift runs all recording operations on a private serial queue (workQueue) rather than the main thread. This architecture ensures that real-time audio capture, file I/O operations, and cleanupOldTemporaryFiles execution do not block the SwiftUI interface, allowing smooth updates to recording indicators via AudioRecorder.isRecording and AudioRecorder.isPlaying bindings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →