Does OpenSuperWhisper Support Real-Time Audio Transcription? A Technical Deep Dive

Yes, OpenSuperWhisper captures live microphone input continuously and processes it through asynchronous transcription engines, delivering text updates to the UI within seconds of speech completion.

OpenSuperWhisper enables real-time audio transcription by combining immediate microphone capture with non-blocking background processing. As implemented in the Starmel/OpenSuperWhisper repository, the app records temporary audio files while you speak, then pipelines that data to either Whisper CPP or FluidAudio engines without freezing the interface.

How Real-Time Transcription Works in OpenSuperWhisper

The architecture follows a five-stage pipeline that keeps the UI responsive while minimizing latency between speech and text.

1. Audio Capture with AudioRecorder

The AudioRecorder class handles live microphone input in AudioRecorder.swift. When you trigger a hotkey or mouse button, it begins writing raw audio to a temporary .wav file on a background queue. Critically, the recorder adds a short audio "tail" after you release the stop command to prevent clipping of the final word.

2. Queue Management via TranscriptionQueue

Once recording stops, AudioRecorder passes the file URL to TranscriptionQueue (TranscriptionQueue.swift). This queue stores a Recording object containing the file path, duration, and status, then schedules it for immediate processing. The queue manages file lifecycle operations, moving temporary recordings to permanent storage after successful transcription.

3. Engine Selection and Configuration

In TranscriptionService.swift (lines 40-55), the TranscriptionService loads your preferred engine based on AppPreferences settings. You can choose between:

  • WhisperEngine: Wraps Whisper CPP for high-accuracy offline transcription
  • FluidAudioEngine: Interfaces with the Parakeet model for low-latency processing

4. Asynchronous Transcription Processing

The transcribeAudio(_:settings:) method in TranscriptionService.swift (lines 83-112) serializes engine access to prevent conflicts while running the actual transcription on a detached task. This method:

  • Applies voice activity detection (VAD) to filter silence
  • Streams progress updates via @Published properties (progress, transcribedText)
  • Returns results without blocking the main thread

5. Result Handling and File Management

After transcription completes, TranscriptionQueue handles cleanup (lines 50-78). The temporary file is copied or moved to your permanent recordings folder, and the UI updates immediately with the finalized text.

Implementation Details: Key Components

Understanding the source code reveals why the real-time pipeline feels instantaneous:

  • AudioRecorder.swift: Manages AVAudioEngine instance and connection status monitoring, ensuring continuous capture even during system interruptions
  • TranscriptionService.swift: Acts as a singleton coordinator that prevents multiple simultaneous transcription jobs while exposing live progress through Combine publishers
  • Engines/WhisperEngine.swift: Implements the transcribeAudio(url:settings:) protocol method with VAD preprocessing
  • Engines/FluidAudioEngine.swift: Provides alternative low-latency transcription for the Parakeet model
  • Utils/AudioUtil.swift: Measures audio duration and validates file integrity before processing

Code Example: Implementing Live Transcription

To integrate OpenSuperWhisper's real-time capabilities into your own Swift code:

import OpenSuperWhisper

// Start capture when user presses hotkey
AudioRecorder.shared.startRecording()

// ... user speaks ...

// Stop and transcribe immediately
Task {
    if let url = await AudioRecorder.shared.stopRecording() {
        let settings = Settings()  // Reads current UI preferences
        do {
            let text = try await TranscriptionService.shared.transcribeAudio(
                url: url, 
                settings: settings
            )
            print("🗣️ Transcription:", text)
        } catch {
            print("❌ Transcription failed:", error)
        }
    }
}

For automatic queue handling (the default UI behavior), manually add files to the pipeline:

TranscriptionQueue.shared.addFileToQueue(url: temporaryRecordingURL)

Engine Comparison: Whisper CPP vs FluidAudio

WhisperEngine provides robust accuracy through the Whisper CPP implementation, making it ideal for longer dictation sessions where precision matters more than speed.

FluidAudioEngine leverages the Parakeet model specifically optimized for latency-sensitive applications, processing shorter audio segments faster while maintaining acceptable accuracy for command-and-control scenarios.

Summary

  • OpenSuperWhisper implements real-time audio transcription through a background queue architecture that records temporary files while you speak
  • The AudioRecorder adds protective audio "tails" to prevent word clipping, while TranscriptionQueue manages file lifecycle and processing order
  • TranscriptionService coordinates between Whisper CPP and FluidAudio engines, streaming progress via Combine publishers at lines 83-112
  • All processing occurs on detached tasks, keeping the UI responsive during transcription
  • The feature is explicitly documented in the README (lines 11-12) as "Real-time audio recording and transcription"

Frequently Asked Questions

How does OpenSuperWhisper handle audio buffering during real-time transcription?

The app writes directly to a temporary .wav file rather than keeping audio in memory buffers. This approach, implemented in AudioRecorder.swift, prevents memory pressure during long recording sessions while ensuring the TranscriptionQueue receives a complete, seekable file immediately upon stop.

What audio formats does OpenSuperWhisper support for live recording?

The AudioRecorder generates standard WAV files with PCM encoding. The AudioUtil.swift helper validates file headers and calculates duration before passing data to the transcription engines, ensuring compatibility with both Whisper CPP and FluidAudio requirements.

Can I switch between transcription engines without restarting the app?

Yes. TranscriptionService reads the engine preference from AppPreferences at the start of each transcription job (lines 40-55). Changing the engine in Settings affects the next queued recording immediately, allowing real-time A/B testing between Whisper CPP and FluidAudio without app restarts.

How does the app prevent the last word from being clipped?

AudioRecorder implements a "tail" mechanism that continues recording for a brief moment (approximately 100-300ms) after receiving the stop command. This buffer captures the decay of final consonants and trailing silence that human reaction time would otherwise miss, ensuring complete transcription of the final spoken word.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →