How to Contribute Code to OpenSuperWhisper: A Complete Developer's Guide

Fork the repository, create a feature branch, implement changes in the relevant Swift subsystem, run ./run.sh build to verify, and submit a pull request targeting the master branch.

OpenSuperWhisper is a macOS transcription application written in Swift that combines a SwiftUI frontend with a pluggable transcription engine architecture. Whether you want to contribute code to OpenSuperWhisper by adding new keyboard shortcuts, supporting additional transcription engines, or improving the model download system, understanding the repository structure is essential for submitting effective patches.

Understanding the Architecture

OpenSuperWhisper organizes functionality into three distinct subsystems that communicate through well-defined interfaces.

Core UI Layer

The Core UI layer handles all user-facing interactions using SwiftUI. This includes the transcription view, preferences panel, and global keyboard shortcuts. Key files in this layer include OpenSuperWhisper/ContentView.swift for the main interface, OpenSuperWhisper/Settings.swift for user preferences, and OpenSuperWhisper/ShortcutManager.swift for global hotkey registration.

Transcription Engine Layer

The Transcription Engine Layer abstracts transcription providers behind a unified interface. The TranscriptionService class orchestrates engines without exposing implementation details to the UI. Current implementations include OpenSuperWhisper/Whis/WhisperEngine.swift (using whisper.cpp) and OpenSuperWhisper/Whis/FluidAudioEngine.swift. The service selects engines based on the loadEngine() method logic found in OpenSuperWhisper/TranscriptionService.swift (lines 40-55).

Model and Download Manager

The Model Manager handles model file lifecycle operations. OpenSuperWhisper/WhisperModelManager.swift creates the app-support directory structure, copies bundled default models, and manages downloads with progress callbacks through the downloadModel(_:name:progressCallback:) method (lines 30-120).

Setting Up Your Development Environment

Before writing code, configure your local environment to build the project including its native dependencies.

First, fork the repository on GitHub, then clone your fork with submodules:

git clone --recursive https://github.com/<your-user>/OpenSuperWhisper.git
cd OpenSuperWhisper

Create a descriptive feature branch:

git checkout -b feature/add-custom-shortcut

The repository includes a run.sh build script that handles both the CMake compilation of the native whisper.cpp library and the Xcode project build.

The Contribution Workflow

Follow this standardized process to ensure your changes integrate smoothly with the existing codebase.

  1. Implement your changes in the appropriate subsystem files, maintaining existing protocols and interfaces.

  2. Run the local build using the provided helper script to verify compilation:

    ./run.sh build
  3. Add or update tests in the OpenSuperWhisperTests/ directory to cover new functionality.

  4. Commit and push your feature branch to your fork.

  5. Open a Pull Request targeting the upstream master branch.

The CI workflow defined in .github/workflows/build.yml automatically compiles the project on macOS, runs tests, and reports status on your pull request.

Where to Add New Functionality

Different feature types require modifications in specific files:

Feature Files to Edit Implementation Notes
New global shortcut OpenSuperWhisper/ShortcutManager.swift and OpenSuperWhisper/ContentView.swift Register the key combo in ShortcutManager, expose UI controls in ContentView
Additional transcription engine New engine class + OpenSuperWhisper/TranscriptionService.swift Conform to TranscriptionEngine protocol, update loadEngine() (lines 40-55) to recognize the new engine identifier
New model download source OpenSuperWhisper/WhisperModelManager.swift and OpenSuperWhisper/Settings.swift Extend downloadModel(_:name:progressCallback:) (lines 30-120) for new URL schemes, update UI to list the source
UI layout changes OpenSuperWhisper/ContentView.swift and related SwiftUI views All visible UI uses SwiftUI components
Audio conversion fixes OpenSuperWhisper/TranscriptionService.swift or OpenSuperWhisper/Whis/WhisperEngine.swift The service tracks conversion state via isConverting and conversionProgress properties

Code Examples

Adding a New Keyboard Shortcut

Extend OpenSuperWhisper/ShortcutManager.swift to register additional global hotkeys:

enum ShortcutAction: String {
    case toggleRecording = "Toggle Recording"
    case startListening = "Start Listening"  // New action
}

func registerShortcuts() {
    let startListening = KeyboardShortcut(key: .l, modifiers: [.command])
    ShortcutManager.shared.register(action: .startListening, shortcut: startListening) {
        TranscriptionService.shared.startRecordingIfNeeded()
    }
}

Implementing a New Transcription Engine

Create a new engine conforming to the TranscriptionEngine protocol:

// OpenSuperWhisper/Whis/MyNewEngine.swift
final class MyNewEngine: TranscriptionEngine {
    var onProgressUpdate: ((Float) -> Void)?

    func initialize() async throws { 
        // Load model and allocate resources
    }

    func transcribeAudio(url: URL, settings: Settings) async throws -> String {
        onProgressUpdate?(0.5)  // Report progress
        return "Transcribed text from MyNewEngine"
    }

    func cancelTranscription() {
        // Abort running inference
    }
}

Then update OpenSuperWhisper/TranscriptionService.swift to instantiate your engine:

// In loadEngine() method (lines 40-55)
if selectedEngine == "mynewengine" {
    engine = await MyNewEngine()
}

Testing Your Implementation

Add tests in OpenSuperWhisperTests/OpenSuperWhisperTests.swift:

func testMyNewEngineTranscription() async throws {
    let engine = MyNewEngine()
    try await engine.initialize()
    let dummyURL = Bundle.main.url(forResource: "jfk", withExtension: "wav")!
    let result = try await engine.transcribeAudio(url: dummyURL, settings: Settings.default)
    XCTAssertFalse(result.isEmpty, "Engine should return non-empty transcription")
}

Summary

Frequently Asked Questions

Do I need to manually compile the whisper.cpp library?

No. The ./run.sh build script automatically invokes CMake to compile the native whisper.cpp library before building the Xcode project. However, ensure you cloned the repository with --recursive to include the whisper.cpp submodule.

What protocol must new transcription engines implement?

New engines must conform to the TranscriptionEngine protocol defined in the codebase, implementing initialize(), transcribeAudio(url:settings:), and cancelTranscription() methods. The onProgressUpdate callback property enables progress reporting to the UI.

How do I add a new model download source?

Extend the downloadModel(_:name:progressCallback:) method in OpenSuperWhisper/WhisperModelManager.swift (lines 30-120) to handle your new URL scheme, then update OpenSuperWhisper/Settings.swift to expose the new source in the preferences UI.

Where does the CI validate my contributions?

The repository uses GitHub Actions defined in .github/workflows/build.yml to automatically compile the project on macOS and run the test suite whenever you open or update a pull request against the master branch.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →