# OpenSuperWhisper Architecture: A Deep Dive into the macOS Real-Time Transcription Engine

> Explore the OpenSuperWhisper architecture, a modular Swift engine for macOS real-time transcription. Understand its UI, audio capture, and transcription engine separation via protocol-oriented design.

- Repository: [Starmel/OpenSuperWhisper](https://github.com/Starmel/OpenSuperWhisper)
- Tags: architecture
- Published: 2026-07-07

---

**OpenSuperWhisper uses a modular Swift architecture that separates UI, audio capture, transcription engines, and background processing through protocol-oriented design and observable state management.**

OpenSuperWhisper is a macOS-native application built in Swift that delivers real-time audio transcription using OpenAI's Whisper model. The codebase, hosted at `Starmel/OpenSuperWhisper`, implements a clean **layered architecture** that cleanly separates concerns between user interface, device management, and transcription processing. This design enables seamless swapping between transcription backends while maintaining a responsive, native macOS experience.

## Layered Architecture Overview

The application organizes functionality into distinct layers, each with specific responsibilities and public interfaces:

| Layer | Responsibility | Core Types | Key Files |
|-------|----------------|------------|-----------|
| **UI & App Lifecycle** | SwiftUI windows, status-bar menu, onboarding flow | `OpenSuperWhisperApp`, `AppDelegate`, `ContentView` | [`OpenSuperWhisper/OpenSuperWhisperApp.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/OpenSuperWhisperApp.swift) |
| **Global State** | Observable objects holding preferences and runtime state | `AppState`, `AppPreferences` | [`OpenSuperWhisper/OpenSuperWhisperApp.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/OpenSuperWhisperApp.swift) |
| **Audio Capture** | Microphone selection, device monitoring, raw audio utilities | `MicrophoneService`, `AudioUtil` | [`OpenSuperWhisper/MicrophoneService.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/MicrophoneService.swift) |
| **Transcription Service** | Public façade loading engines and exposing progress | `TranscriptionService` | [`OpenSuperWhisper/TranscriptionService.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/TranscriptionService.swift) |
| **Background Queue** | Serial queue processing files and recordings | `TranscriptionQueue`, `RecordingStore` | [`OpenSuperWhisper/TranscriptionQueue.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/TranscriptionQueue.swift) |
| **Engines** | Interchangeable transcription backends | `WhisperEngine`, `FluidAudioEngine` | [`OpenSuperWhisper/Engines/WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Engines/WhisperEngine.swift) |

## UI and App Lifecycle

The entry point resides in **[`OpenSuperWhisperApp.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisperApp.swift)**, where the `@main` struct `OpenSuperWhisperApp` initializes the global state and delegates system events. The application uses a `WindowGroup` to host either the onboarding flow or the main `ContentView` based on user completion status.

**`AppDelegate`** handles macOS-specific integrations that SwiftUI cannot manage directly. It constructs the **status-bar menu**, monitors window lifecycle events, and implements file-open callbacks. When the app launches, `OpenSuperWhisperApp.startTranscriptionQueue()` initializes the background processing pipeline. Keyboard shortcuts and mouse-button triggers route through `ShortcutManager` and `ModifierKeyMonitor` before reaching the transcription service.

## Global State Management

**`AppState`** acts as the central observable object, exposing `@Published` properties that the SwiftUI layer binds to for reactive updates. It tracks runtime flags like `hasCompletedOnboarding` and persists user preferences through **`AppPreferences`**. This singleton pattern ensures consistent state across windows while maintaining separation between volatile UI state and durable settings storage.

The preferences system stores configuration for selected models, language settings, shortcut mappings, and microphone selections, enabling the app to restore user context across launches.

## Audio Capture System

**`MicrophoneService`** manages hardware discovery and selection through `AVCaptureDevice.DiscoverySession`. It maintains a `@Published` array of available microphones and broadcasts changes via `Notification.Name.microphoneDidChange`. The service provides methods to set the system default input device and query device capabilities, distinguishing between Bluetooth, USB, and continuity microphones.

Raw audio utilities in `AudioUtil` handle PCM conversion and duration calculation, supporting the transcription pipeline with standardized audio formats.

## Transcription Service Layer

**`TranscriptionService`** is a **singleton** (`shared`) that serves as the public façade for all transcription operations. It owns the current **`TranscriptionEngine`** instance and manages its lifecycle based on the user's `selectedEngine` preference.

The primary method `transcribeAudio(url:settings:)` serializes access to the engine, wires progress callbacks, and returns the final transcription string. Progress updates flow back to the UI through `@Published var progress: Float`, enabling real-time progress bars during transcription.

```swift
import OpenSuperWhisper

func transcribeFile(at url: URL) async {
    let settings = Settings()               // uses app‑wide defaults
    do {
        let transcript = try await TranscriptionService.shared.transcribeAudio(url: url,
                                                                              settings: settings)
        print("Result:\n\(transcript)")
    } catch {
        print("Transcription failed: \(error)")
    }
}

```

## Background Processing Queue

**`TranscriptionQueue`** is a **Main-actor** singleton managing a FIFO queue of recordings awaiting transcription. When users drop files onto the app or temporary recordings finish, `addFileToQueue(url:)` creates a `Recording` entry (calculating duration via `AudioUtil.audioDuration`) and triggers `startProcessingQueue()`.

The queue processes each recording sequentially, driving `TranscriptionService.transcribeAudio` and persisting results via `RecordingStore`. It handles cancellation, cleanup of missing source files, and removal of empty dictations without blocking the main UI thread.

```swift
func application(_ sender: NSApplication,
                 openFiles filenames: [String]) {
    let audioURLs = filenames
        .map { URL(fileURLWithPath: $0) }
        .filter { isAudioFile($0) }

    for url in audioURLs {
        Task { @MainActor in
            await TranscriptionQueue.shared.addFileToQueue(url: url)
        }
    }
}

```

## Pluggable Transcription Engines

OpenSuperWhisper supports multiple transcription backends through the **`TranscriptionEngine`** protocol:

```swift
protocol TranscriptionEngine {
    var isModelLoaded: Bool { get }
    var engineName: String { get }
    func initialize() async throws
    func transcribeAudio(url: URL, settings: Settings) async throws -> String
    func cancelTranscription()
    func getSupportedLanguages() -> [String]
}

```

This protocol-oriented design allows the app to switch between **WhisperEngine** and **FluidAudioEngine** without changing consumer code.

### WhisperEngine Implementation Details

**`WhisperEngine`** wraps **whisper.cpp** for local inference. It loads Whisper `.bin` models via `MyWhisperContext.initFromFileNoState` and initializes a Silero VAD (Voice Activity Detection) model for silence removal.

The engine performs **PCM conversion** using multithreaded `AVAudio` conversion, then applies VAD through `detectSpeech` to strip non-speech segments. Transcription occurs via `context.full(samples:params:)`, with progress reported through a C-callback bridged to Swift via `ProgressContext`. The implementation supports beam-search decoding, timestamp generation, temperature sampling, and post-processing for Asian languages through `AutocorrectWrapper`.

### FluidAudioEngine

The **FluidAudioEngine** provides an alternative backend following the same protocol contract, exposing progress updates through an `onProgressUpdate` closure. This demonstrates how the architecture accommodates third-party transcription services without modifying the UI or queue management layers.

## Data Flow Walkthrough

Understanding the **OpenSuperWhisper architecture** requires tracing a complete transcription cycle:

1. **User Interaction** – A keyboard shortcut, menu selection, or file drop triggers `AppDelegate` or `ShortcutManager`
2. **Queue Insertion** – `AppDelegate` enqueues the audio URL via `TranscriptionQueue.addFileToQueue`
3. **Recording Creation** – `TranscriptionQueue` instantiates a `Recording` object with metadata
4. **Service Invocation** – If idle, the queue calls `TranscriptionService.transcribeAudio`, which ensures the selected engine is loaded
5. **Engine Processing** – The active engine converts audio, runs the model, and returns a transcription string
6. **Persistence** – `TranscriptionQueue` stores the result via `RecordingStore`, updates UI state, and manages file system operations

All components communicate through **loosely coupled** protocols and `@Published` properties, enabling testable units and straightforward backend swaps.

## Summary

- **OpenSuperWhisper** implements a **layered architecture** separating UI, state, audio capture, and transcription into distinct modules
- **`TranscriptionService`** acts as a singleton façade, abstracting engine-specific implementations behind the `TranscriptionEngine` protocol
- **`TranscriptionQueue`** provides Main-actor-guarded serial processing for background transcription tasks
- **`MicrophoneService`** handles hardware abstraction through `AVCaptureDevice` with reactive `@Published` updates
- The protocol-based engine design supports both **whisper.cpp** (via `WhisperEngine`) and external services (via `FluidAudioEngine`) without UI changes
- File paths like [`OpenSuperWhisper/TranscriptionService.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/TranscriptionService.swift) and [`OpenSuperWhisper/Engines/WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Engines/WhisperEngine.swift) demonstrate clear organizational boundaries

## Frequently Asked Questions

### What programming language and framework does OpenSuperWhisper use?

OpenSuperWhisper is built entirely in **Swift** using **SwiftUI** for the user interface and **AppKit** integrations via `AppDelegate`. The architecture leverages `Combine` and `@Published` properties for reactive state management, with concurrency handled through Swift's structured concurrency (`async/await`) and `@MainActor` annotations for UI safety.

### How does OpenSuperWhisper handle microphone selection and audio input?

The **`MicrophoneService`** class in [`OpenSuperWhisper/MicrophoneService.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/MicrophoneService.swift) discovers available input devices using `AVCaptureDevice.DiscoverySession`. It maintains a published list of microphones and allows programmatic selection of the system default input. The service also categorizes devices by type (Bluetooth, built-in, USB) and broadcasts changes through `NotificationCenter`, ensuring the UI remains synchronized with hardware state.

### Can OpenSuperWhisper use different transcription models besides Whisper?

Yes. The architecture supports **pluggable transcription engines** through the `TranscriptionEngine` protocol. While `WhisperEngine` provides local inference via whisper.cpp, the codebase includes `FluidAudioEngine` as an alternative backend. Both engines implement the same interface methods (`initialize()`, `transcribeAudio()`, `cancelTranscription()`), allowing the `TranscriptionService` to swap implementations at runtime based on user preferences without modifying the UI or queue management code.

### How does the background transcription queue work?

**`TranscriptionQueue`** is a Main-actor singleton that maintains a FIFO queue of `Recording` objects. When audio files are dropped on the app or recordings complete, `addFileToQueue(url:)` adds entries and triggers `startProcessingQueue()`. The queue processes items sequentially, invoking `TranscriptionService.transcribeAudio` for each file and persisting results through `RecordingStore`. This serial processing prevents memory pressure from concurrent model inference while providing progress updates to the UI through the service layer.