# Performance Characteristics of the Whisper Engine in OpenSuperWhisper: 7 Real-Time Optimizations

> Discover the Whisper engine's impressive performance in OpenSuperWhisper. Learn how dynamic thread pools and parallel processing achieve sub-second transcription latency on macOS.

- Repository: [Starmel/OpenSuperWhisper](https://github.com/Starmel/OpenSuperWhisper)
- Tags: performance
- Published: 2026-07-05

---

**The Whisper engine in OpenSuperWhisper delivers sub-second transcription latency on macOS by combining a dynamic thread pool, parallel audio conversion, RMS-based channel mixing, and C-level abort callbacks.**

OpenSuperWhisper is an open-source macOS transcription app maintained by Starmel that delivers real-time speech-to-text through a custom Whisper engine. Understanding the performance characteristics of the Whisper engine is essential for developers building low-latency audio pipelines. The repository's source code—particularly [`OpenSuperWhisper/Engines/WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Engines/WhisperEngine.swift)—exposes seven concrete optimizations that bound CPU usage, minimize memory churn, and keep cancellation responsive on modern Apple Silicon Macs.

## CPU and Threading Optimizations

### Dynamic Thread Pool for Inference

The engine computes the thread count for Whisper inference in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) (lines 14–15) by taking the lesser of the active processor count and eight, with a floor of two:

```swift
let nThreads = max(2, min(ProcessInfo.processInfo.activeProcessorCount, 8))

```

This guarantees consistent multi-core speed-up without overwhelming the system scheduler.

### Parallel Audio Conversion for Long Files

For inputs exceeding ten seconds, audio conversion to 16 kHz PCM is parallelized across a worker per CPU core. Comments in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) (lines 57–59) note measured speed-ups of +339% on four cores and +609% on eight cores:

```swift
// Use parallel processing for large files (> 10 seconds of audio)
// Benchmarked: 4 cores = +339%, 8 cores = +609% improvement

```

This optimization improves end-to-end latency by keeping pre-processing off the main thread.

## Audio Pre-Processing Optimizations

### Efficient Mixed-Channel Handling via RMS Detection

Instead of mixing every channel blindly, the engine detects activity via RMS and only mixes active channels. In [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) (lines 59–80), silent channels are filtered out using a threshold of `0.0001` and active channels are normalized:

```swift
let activityThreshold: Float = 0.0001
let normalization = 1.0 / Float(activeChannels.count)

```

Skipping inactive data reduces the payload fed into the model, cutting inference time for multi-channel sources.

### Optimized 16 kHz Mono Float32 Target Format

The engine converts all input to the exact format Whisper expects—16 kHz mono float32—using `AVAudioConverter` at medium quality. The return statement in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) (lines 94–101) illustrates this:

```swift
return AVAudioFormat(commonFormat: .pcmFormatFloat32, sampleRate: 16000, ...)

```

By front-loading resampling, the model avoids expensive internal format conversion and runs at peak throughput.

### Cache-Friendly Buffer Allocation

To minimize memory churn, the engine pre-allocates the output float array based on the expected frame count and fills it in-place. Lines 77–79 of [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) show the allocation:

```swift
var result = [Float](repeating: 0, count: outputFrameCount)

```

Pre-allocation lowers garbage collection pressure, which is critical for maintaining smooth real-time performance.

## Runtime Responsiveness and Progress Feedback

### Instant Cancellation with GGMLAbortCallback

A C-level `GGMLAbortCallback` is registered directly with Whisper’s inference loop. Defined in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) (lines 29–31), the callback type allows the engine to break out of native inference immediately:

```swift
typealias GGMLAbortCallback = @convention(c) (UnsafeMutableRawPointer?) -> Bool

```

When `cancelTranscription()` is triggered, the UI remains responsive because the abort does not wait for the current segment to finish.

### Progress Mapping for Perceived Performance

Whisper reports progress from 0–100%, but the engine remaps it to a 10–95% range so the UI never appears stuck. Lines 42–44 of [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) contain the remap formula:

```swift
let normalizedProgress = 0.10 + (Float(progressPercent) / 100.0) * 0.85

```

This mapping provides continuous visual feedback during both audio conversion and model inference, improving the perceived speed of transcription.

## Practical Code Examples

### Transcribing a File with Progress Monitoring

```swift
let engine = WhisperEngine()
try await engine.initialize()

engine.onProgressUpdate = { progress in
    print("Transcription progress: \(Int(progress * 100)) %")
}

let settings = Settings()
settings.useBeamSearch = false
settings.showTimestamps = true

let result = try await engine.transcribeAudio(
    url: URL(fileURLWithPath: "/path/to/audio.m4a"),
    settings: settings
)

print("Result:\n\(result)")

```

### Cancelling an Ongoing Transcription

```swift
Task {
    do {
        let text = try await engine.transcribeAudio(url: audioURL, settings: settings)
        print(text)
    } catch {
        print("Transcription stopped or failed:", error)
    }
}

// Later, on user request
engine.cancelTranscription()

```

## Key Source Files

| File | Role |
|------|------|
| [`OpenSuperWhisper/Engines/WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Engines/WhisperEngine.swift) | Core transcription engine: model loading, audio conversion, progress handling, and abort callbacks. |
| [`OpenSuperWhisper/Whis/WhisperFullParams.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Whis/WhisperFullParams.swift) | Swift wrapper for Whisper’s C-struct inference parameters, including thread count and sampling strategy. |
| [`OpenSuperWhisper/Whis/WhisperContextParams.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Whis/WhisperContextParams.swift) | Context-level configuration such as DTW heads and alignment presets. |
| [`OpenSuperWhisper/Whis/WhisperModelLoader.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Whis/WhisperModelLoader.swift) | Loads the Whisper binary (`*.bin`) into native memory via `MyWhisperContext`. |
| [`OpenSuperWhisper/Settings.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/OpenSuperWhisper/Settings.swift) | User-facing transcription settings: beam size, temperature, language, and more. |

## Summary

- The Whisper engine limits inference threads to `max(2, min(activeProcessorCount, 8))`, preventing CPU saturation while exploiting multi-core Macs.
- Parallel audio conversion for inputs over ten seconds achieves up to +609% speed-up on eight-core machines.
- RMS-based channel detection strips silent channels before mixing, shrinking the data passed to the model.
- Audio is normalized to 16 kHz mono float32 ahead of inference to eliminate costly runtime format conversion.
- Pre-allocated float buffers reduce GC pressure and memory churn.
- The `GGMLAbortCallback` integration enables instantaneous cancellation without blocking the main thread.
- Progress values are remapped from 0–100% to 10–95% to give users smooth, uninterrupted visual feedback.

## Frequently Asked Questions

### How does the Whisper engine handle multi-core CPUs?

The engine queries `ProcessInfo.processInfo.activeProcessorCount` and caps inference threads at eight. This ensures the Whisper engine scales across multi-core Apple Silicon Macs without monopolizing system resources.

### Can the Whisper engine transcribe long audio files efficiently?

Yes. Files longer than ten seconds trigger parallel audio conversion across all CPU cores. Benchmarks embedded in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) show a +339% improvement on four cores and +609% on eight cores, dramatically lowering end-to-end latency.

### What happens when a user cancels transcription mid-stream?

The engine invokes a `GGMLAbortCallback` that terminates the underlying C-level inference loop immediately. This avoids waiting for the current Whisper segment to complete and keeps the macOS UI fully responsive.

### Why does the engine remap progress percentages to 10–95%?

Raw Whisper progress spans 0–100%, but the engine reserves headroom for audio conversion and finalization steps. The remap formula `0.10 + (progressPercent / 100.0) * 0.85` ensures the progress bar never stalls at zero or jumps abruptly to completion.