Performance Characteristics of the Whisper Engine in OpenSuperWhisper: 7 Real-Time Optimizations

The Whisper engine in OpenSuperWhisper delivers sub-second transcription latency on macOS by combining a dynamic thread pool, parallel audio conversion, RMS-based channel mixing, and C-level abort callbacks.

OpenSuperWhisper is an open-source macOS transcription app maintained by Starmel that delivers real-time speech-to-text through a custom Whisper engine. Understanding the performance characteristics of the Whisper engine is essential for developers building low-latency audio pipelines. The repository's source code—particularly OpenSuperWhisper/Engines/WhisperEngine.swift—exposes seven concrete optimizations that bound CPU usage, minimize memory churn, and keep cancellation responsive on modern Apple Silicon Macs.

CPU and Threading Optimizations

Dynamic Thread Pool for Inference

The engine computes the thread count for Whisper inference in WhisperEngine.swift (lines 14–15) by taking the lesser of the active processor count and eight, with a floor of two:

let nThreads = max(2, min(ProcessInfo.processInfo.activeProcessorCount, 8))

This guarantees consistent multi-core speed-up without overwhelming the system scheduler.

Parallel Audio Conversion for Long Files

For inputs exceeding ten seconds, audio conversion to 16 kHz PCM is parallelized across a worker per CPU core. Comments in WhisperEngine.swift (lines 57–59) note measured speed-ups of +339% on four cores and +609% on eight cores:

// Use parallel processing for large files (> 10 seconds of audio)
// Benchmarked: 4 cores = +339%, 8 cores = +609% improvement

This optimization improves end-to-end latency by keeping pre-processing off the main thread.

Audio Pre-Processing Optimizations

Efficient Mixed-Channel Handling via RMS Detection

Instead of mixing every channel blindly, the engine detects activity via RMS and only mixes active channels. In WhisperEngine.swift (lines 59–80), silent channels are filtered out using a threshold of 0.0001 and active channels are normalized:

let activityThreshold: Float = 0.0001
let normalization = 1.0 / Float(activeChannels.count)

Skipping inactive data reduces the payload fed into the model, cutting inference time for multi-channel sources.

Optimized 16 kHz Mono Float32 Target Format

The engine converts all input to the exact format Whisper expects—16 kHz mono float32—using AVAudioConverter at medium quality. The return statement in WhisperEngine.swift (lines 94–101) illustrates this:

return AVAudioFormat(commonFormat: .pcmFormatFloat32, sampleRate: 16000, ...)

By front-loading resampling, the model avoids expensive internal format conversion and runs at peak throughput.

Cache-Friendly Buffer Allocation

To minimize memory churn, the engine pre-allocates the output float array based on the expected frame count and fills it in-place. Lines 77–79 of WhisperEngine.swift show the allocation:

var result = [Float](repeating: 0, count: outputFrameCount)

Pre-allocation lowers garbage collection pressure, which is critical for maintaining smooth real-time performance.

Runtime Responsiveness and Progress Feedback

Instant Cancellation with GGMLAbortCallback

A C-level GGMLAbortCallback is registered directly with Whisper’s inference loop. Defined in WhisperEngine.swift (lines 29–31), the callback type allows the engine to break out of native inference immediately:

typealias GGMLAbortCallback = @convention(c) (UnsafeMutableRawPointer?) -> Bool

When cancelTranscription() is triggered, the UI remains responsive because the abort does not wait for the current segment to finish.

Progress Mapping for Perceived Performance

Whisper reports progress from 0–100%, but the engine remaps it to a 10–95% range so the UI never appears stuck. Lines 42–44 of WhisperEngine.swift contain the remap formula:

let normalizedProgress = 0.10 + (Float(progressPercent) / 100.0) * 0.85

This mapping provides continuous visual feedback during both audio conversion and model inference, improving the perceived speed of transcription.

Practical Code Examples

Transcribing a File with Progress Monitoring

let engine = WhisperEngine()
try await engine.initialize()

engine.onProgressUpdate = { progress in
    print("Transcription progress: \(Int(progress * 100)) %")
}

let settings = Settings()
settings.useBeamSearch = false
settings.showTimestamps = true

let result = try await engine.transcribeAudio(
    url: URL(fileURLWithPath: "/path/to/audio.m4a"),
    settings: settings
)

print("Result:\n\(result)")

Cancelling an Ongoing Transcription

Task {
    do {
        let text = try await engine.transcribeAudio(url: audioURL, settings: settings)
        print(text)
    } catch {
        print("Transcription stopped or failed:", error)
    }
}

// Later, on user request
engine.cancelTranscription()

Key Source Files

File Role
OpenSuperWhisper/Engines/WhisperEngine.swift Core transcription engine: model loading, audio conversion, progress handling, and abort callbacks.
OpenSuperWhisper/Whis/WhisperFullParams.swift Swift wrapper for Whisper’s C-struct inference parameters, including thread count and sampling strategy.
OpenSuperWhisper/Whis/WhisperContextParams.swift Context-level configuration such as DTW heads and alignment presets.
OpenSuperWhisper/Whis/WhisperModelLoader.swift Loads the Whisper binary (*.bin) into native memory via MyWhisperContext.
OpenSuperWhisper/Settings.swift User-facing transcription settings: beam size, temperature, language, and more.

Summary

  • The Whisper engine limits inference threads to max(2, min(activeProcessorCount, 8)), preventing CPU saturation while exploiting multi-core Macs.
  • Parallel audio conversion for inputs over ten seconds achieves up to +609% speed-up on eight-core machines.
  • RMS-based channel detection strips silent channels before mixing, shrinking the data passed to the model.
  • Audio is normalized to 16 kHz mono float32 ahead of inference to eliminate costly runtime format conversion.
  • Pre-allocated float buffers reduce GC pressure and memory churn.
  • The GGMLAbortCallback integration enables instantaneous cancellation without blocking the main thread.
  • Progress values are remapped from 0–100% to 10–95% to give users smooth, uninterrupted visual feedback.

Frequently Asked Questions

How does the Whisper engine handle multi-core CPUs?

The engine queries ProcessInfo.processInfo.activeProcessorCount and caps inference threads at eight. This ensures the Whisper engine scales across multi-core Apple Silicon Macs without monopolizing system resources.

Can the Whisper engine transcribe long audio files efficiently?

Yes. Files longer than ten seconds trigger parallel audio conversion across all CPU cores. Benchmarks embedded in WhisperEngine.swift show a +339% improvement on four cores and +609% on eight cores, dramatically lowering end-to-end latency.

What happens when a user cancels transcription mid-stream?

The engine invokes a GGMLAbortCallback that terminates the underlying C-level inference loop immediately. This avoids waiting for the current Whisper segment to complete and keeps the macOS UI fully responsive.

Why does the engine remap progress percentages to 10–95%?

Raw Whisper progress spans 0–100%, but the engine reserves headroom for audio conversion and finalization steps. The remap formula 0.10 + (progressPercent / 100.0) * 0.85 ensures the progress bar never stalls at zero or jumps abruptly to completion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →