OpenSuperWhisper Architecture: A Deep Dive into the macOS Real-Time Transcription Engine
OpenSuperWhisper uses a modular Swift architecture that separates UI, audio capture, transcription engines, and background processing through protocol-oriented design and observable state management.
OpenSuperWhisper is a macOS-native application built in Swift that delivers real-time audio transcription using OpenAI's Whisper model. The codebase, hosted at Starmel/OpenSuperWhisper, implements a clean layered architecture that cleanly separates concerns between user interface, device management, and transcription processing. This design enables seamless swapping between transcription backends while maintaining a responsive, native macOS experience.
Layered Architecture Overview
The application organizes functionality into distinct layers, each with specific responsibilities and public interfaces:
| Layer | Responsibility | Core Types | Key Files |
|---|---|---|---|
| UI & App Lifecycle | SwiftUI windows, status-bar menu, onboarding flow | OpenSuperWhisperApp, AppDelegate, ContentView |
OpenSuperWhisper/OpenSuperWhisperApp.swift |
| Global State | Observable objects holding preferences and runtime state | AppState, AppPreferences |
OpenSuperWhisper/OpenSuperWhisperApp.swift |
| Audio Capture | Microphone selection, device monitoring, raw audio utilities | MicrophoneService, AudioUtil |
OpenSuperWhisper/MicrophoneService.swift |
| Transcription Service | Public façade loading engines and exposing progress | TranscriptionService |
OpenSuperWhisper/TranscriptionService.swift |
| Background Queue | Serial queue processing files and recordings | TranscriptionQueue, RecordingStore |
OpenSuperWhisper/TranscriptionQueue.swift |
| Engines | Interchangeable transcription backends | WhisperEngine, FluidAudioEngine |
OpenSuperWhisper/Engines/WhisperEngine.swift |
UI and App Lifecycle
The entry point resides in OpenSuperWhisperApp.swift, where the @main struct OpenSuperWhisperApp initializes the global state and delegates system events. The application uses a WindowGroup to host either the onboarding flow or the main ContentView based on user completion status.
AppDelegate handles macOS-specific integrations that SwiftUI cannot manage directly. It constructs the status-bar menu, monitors window lifecycle events, and implements file-open callbacks. When the app launches, OpenSuperWhisperApp.startTranscriptionQueue() initializes the background processing pipeline. Keyboard shortcuts and mouse-button triggers route through ShortcutManager and ModifierKeyMonitor before reaching the transcription service.
Global State Management
AppState acts as the central observable object, exposing @Published properties that the SwiftUI layer binds to for reactive updates. It tracks runtime flags like hasCompletedOnboarding and persists user preferences through AppPreferences. This singleton pattern ensures consistent state across windows while maintaining separation between volatile UI state and durable settings storage.
The preferences system stores configuration for selected models, language settings, shortcut mappings, and microphone selections, enabling the app to restore user context across launches.
Audio Capture System
MicrophoneService manages hardware discovery and selection through AVCaptureDevice.DiscoverySession. It maintains a @Published array of available microphones and broadcasts changes via Notification.Name.microphoneDidChange. The service provides methods to set the system default input device and query device capabilities, distinguishing between Bluetooth, USB, and continuity microphones.
Raw audio utilities in AudioUtil handle PCM conversion and duration calculation, supporting the transcription pipeline with standardized audio formats.
Transcription Service Layer
TranscriptionService is a singleton (shared) that serves as the public façade for all transcription operations. It owns the current TranscriptionEngine instance and manages its lifecycle based on the user's selectedEngine preference.
The primary method transcribeAudio(url:settings:) serializes access to the engine, wires progress callbacks, and returns the final transcription string. Progress updates flow back to the UI through @Published var progress: Float, enabling real-time progress bars during transcription.
import OpenSuperWhisper
func transcribeFile(at url: URL) async {
let settings = Settings() // uses app‑wide defaults
do {
let transcript = try await TranscriptionService.shared.transcribeAudio(url: url,
settings: settings)
print("Result:\n\(transcript)")
} catch {
print("Transcription failed: \(error)")
}
}
Background Processing Queue
TranscriptionQueue is a Main-actor singleton managing a FIFO queue of recordings awaiting transcription. When users drop files onto the app or temporary recordings finish, addFileToQueue(url:) creates a Recording entry (calculating duration via AudioUtil.audioDuration) and triggers startProcessingQueue().
The queue processes each recording sequentially, driving TranscriptionService.transcribeAudio and persisting results via RecordingStore. It handles cancellation, cleanup of missing source files, and removal of empty dictations without blocking the main UI thread.
func application(_ sender: NSApplication,
openFiles filenames: [String]) {
let audioURLs = filenames
.map { URL(fileURLWithPath: $0) }
.filter { isAudioFile($0) }
for url in audioURLs {
Task { @MainActor in
await TranscriptionQueue.shared.addFileToQueue(url: url)
}
}
}
Pluggable Transcription Engines
OpenSuperWhisper supports multiple transcription backends through the TranscriptionEngine protocol:
protocol TranscriptionEngine {
var isModelLoaded: Bool { get }
var engineName: String { get }
func initialize() async throws
func transcribeAudio(url: URL, settings: Settings) async throws -> String
func cancelTranscription()
func getSupportedLanguages() -> [String]
}
This protocol-oriented design allows the app to switch between WhisperEngine and FluidAudioEngine without changing consumer code.
WhisperEngine Implementation Details
WhisperEngine wraps whisper.cpp for local inference. It loads Whisper .bin models via MyWhisperContext.initFromFileNoState and initializes a Silero VAD (Voice Activity Detection) model for silence removal.
The engine performs PCM conversion using multithreaded AVAudio conversion, then applies VAD through detectSpeech to strip non-speech segments. Transcription occurs via context.full(samples:params:), with progress reported through a C-callback bridged to Swift via ProgressContext. The implementation supports beam-search decoding, timestamp generation, temperature sampling, and post-processing for Asian languages through AutocorrectWrapper.
FluidAudioEngine
The FluidAudioEngine provides an alternative backend following the same protocol contract, exposing progress updates through an onProgressUpdate closure. This demonstrates how the architecture accommodates third-party transcription services without modifying the UI or queue management layers.
Data Flow Walkthrough
Understanding the OpenSuperWhisper architecture requires tracing a complete transcription cycle:
- User Interaction – A keyboard shortcut, menu selection, or file drop triggers
AppDelegateorShortcutManager - Queue Insertion –
AppDelegateenqueues the audio URL viaTranscriptionQueue.addFileToQueue - Recording Creation –
TranscriptionQueueinstantiates aRecordingobject with metadata - Service Invocation – If idle, the queue calls
TranscriptionService.transcribeAudio, which ensures the selected engine is loaded - Engine Processing – The active engine converts audio, runs the model, and returns a transcription string
- Persistence –
TranscriptionQueuestores the result viaRecordingStore, updates UI state, and manages file system operations
All components communicate through loosely coupled protocols and @Published properties, enabling testable units and straightforward backend swaps.
Summary
- OpenSuperWhisper implements a layered architecture separating UI, state, audio capture, and transcription into distinct modules
TranscriptionServiceacts as a singleton façade, abstracting engine-specific implementations behind theTranscriptionEngineprotocolTranscriptionQueueprovides Main-actor-guarded serial processing for background transcription tasksMicrophoneServicehandles hardware abstraction throughAVCaptureDevicewith reactive@Publishedupdates- The protocol-based engine design supports both whisper.cpp (via
WhisperEngine) and external services (viaFluidAudioEngine) without UI changes - File paths like
OpenSuperWhisper/TranscriptionService.swiftandOpenSuperWhisper/Engines/WhisperEngine.swiftdemonstrate clear organizational boundaries
Frequently Asked Questions
What programming language and framework does OpenSuperWhisper use?
OpenSuperWhisper is built entirely in Swift using SwiftUI for the user interface and AppKit integrations via AppDelegate. The architecture leverages Combine and @Published properties for reactive state management, with concurrency handled through Swift's structured concurrency (async/await) and @MainActor annotations for UI safety.
How does OpenSuperWhisper handle microphone selection and audio input?
The MicrophoneService class in OpenSuperWhisper/MicrophoneService.swift discovers available input devices using AVCaptureDevice.DiscoverySession. It maintains a published list of microphones and allows programmatic selection of the system default input. The service also categorizes devices by type (Bluetooth, built-in, USB) and broadcasts changes through NotificationCenter, ensuring the UI remains synchronized with hardware state.
Can OpenSuperWhisper use different transcription models besides Whisper?
Yes. The architecture supports pluggable transcription engines through the TranscriptionEngine protocol. While WhisperEngine provides local inference via whisper.cpp, the codebase includes FluidAudioEngine as an alternative backend. Both engines implement the same interface methods (initialize(), transcribeAudio(), cancelTranscription()), allowing the TranscriptionService to swap implementations at runtime based on user preferences without modifying the UI or queue management code.
How does the background transcription queue work?
TranscriptionQueue is a Main-actor singleton that maintains a FIFO queue of Recording objects. When audio files are dropped on the app or recordings complete, addFileToQueue(url:) adds entries and triggers startProcessingQueue(). The queue processes items sequentially, invoking TranscriptionService.transcribeAudio for each file and persisting results through RecordingStore. This serial processing prevents memory pressure from concurrent model inference while providing progress updates to the UI through the service layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →