What Are WhisperAheads and How They Enable Accurate Token Timestamps in OpenSuperWhisper
WhisperAheads are Swift wrapper structures that map specific decoder attention heads to the underlying Whisper C library, enabling dynamic time warping (DTW) for precise token-level timestamp generation during transcription.
OpenSuperWhisper provides a Swift interface to OpenAI's Whisper speech recognition models. When transcribing audio, WhisperAheads serve as the configuration bridge that tells the model exactly which attention layers to consult for alignment data, directly impacting the accuracy of word-level timestamps.
Understanding the WhisperAheads Structure
The Individual Alignment Head (WhisperAhead)
In WhisperAhead.swift, the WhisperAhead struct represents a single alignment head corresponding to the C library's whisper_ahead struct. Each head specifies two critical indices: nTextLayer (the decoder layer index) and nHead (the attention head index within that layer). These indices determine which specific attention weights the dynamic time warping algorithm examines when correlating audio segments with text tokens.
The Collection Wrapper (WhisperAheads)
Defined in WhisperAheads.swift, the WhisperAheads struct bundles multiple WhisperAhead instances into a collection compatible with the C whisper_aheads structure. This wrapper manages the underlying array of heads and provides toC() and fromC(_:) methods for marshalling data between Swift and the underlying C library.
Preset Configurations via WhisperAlignmentHeadsPreset
Rather than manually specifying individual heads, developers can use WhisperAlignmentHeadsPreset defined in WhisperAlignmentHeadsPreset.swift. This enum provides pre-calibrated head sets optimized for specific model variants like tinyEn, base, and largeV3Turbo. The preset automatically expands into the correct WhisperAheads configuration for the selected model architecture.
How WhisperAheads Drive the Transcription Process
Configuration in WhisperContextParams
The transcription pipeline configures alignment heads through WhisperContextParams.swift. This file defines the WhisperContextParams struct, which contains two key fields: dtwAheadsPreset for selecting a built-in preset, and dtwAheads for supplying a custom WhisperAheads instance. According to the source code in Whis.swift (lines 460-466), these parameters are initialized when creating a new Whisper context.
Marshalling to the C Library
When initializing transcription, the Swift layer converts configuration data to C structures. The toC() method in WhisperAheads transforms the Swift array into a whisper_aheads C struct containing a pointer to whisper_ahead elements. This marshalling occurs within WhisperContextParams.toC(), which passes the resulting pointer to the whisper_context_params configuration used by the underlying library.
DTW Alignment During Transcription
During the actual transcription process, the Whisper C library consults the supplied heads to compute alignment scores for each token. By examining specific decoder layers and attention heads as defined in the WhisperAheads collection, the dynamic time warping algorithm identifies precise token boundaries within the audio stream. This enables higher-accuracy timestamps, particularly crucial for large models with multiple decoder layers where not all heads contain relevant alignment information.
Implementing WhisperAheads in Code
Selecting a Preset
Use a built-in preset when you want optimized alignment heads for standard model variants:
var params = WhisperContextParams()
params.dtwAheadsPreset = .tinyEn
let ctx = try WhisperContext(modelPath: "ggml-tiny.en.bin", params: params)
Supplying Custom Heads
Define specific layers and head indices when you need fine-grained control over the alignment process:
let customHeads = WhisperAheads(heads: [
WhisperAhead(nTextLayer: 0, nHead: 4),
WhisperAhead(nTextLayer: 1, nHead: 2)
])
var params = WhisperContextParams()
params.dtwAheads = customHeads
let ctx = try WhisperContext(modelPath: "ggml-base.bin", params: params)
Accessing Effective Heads
Inspect which heads are active after context creation:
let effectiveHeads = ctx.params.dtwAheads
print("Using \(effectiveHeads.heads.count) alignment heads")
Summary
- WhisperAheads are Swift wrappers for the C library's alignment head structures, enabling DTW-based timestamping.
- Individual heads specify layer and head indices via
WhisperAhead, while collections manage multiple heads viaWhisperAheads. - Configuration occurs through
WhisperContextParamsusing either presets (dtwAheadsPreset) or custom head arrays (dtwAheads). - The
toC()method marshals Swift structures to C for consumption by the underlying transcription engine. - Key source files include
WhisperAhead.swift,WhisperAheads.swift, andWhisperContextParams.swiftin the OpenSuperWhisper repository.
Frequently Asked Questions
What is the difference between WhisperAhead and WhisperAheads?
WhisperAhead represents a single alignment head specifying one text layer and head index, while WhisperAheads is a collection wrapper that manages multiple heads and handles C interoperability through toC() and fromC(_:) methods. You typically configure the collection, not individual heads, when setting up transcription parameters.
Can I use WhisperAheads with any Whisper model variant?
Yes, but you should match the preset to your model size. The WhisperAlignmentHeadsPreset enum provides optimized head configurations for specific variants like tinyEn, base, and largeV3Turbo, ensuring the selected layers exist in your chosen model architecture. Using mismatched presets may result in invalid layer references.
How do WhisperAheads improve transcription timestamp accuracy?
They enable dynamic time warping by identifying which specific decoder attention layers contain alignment information. By consulting only relevant heads rather than all layers, the DTW algorithm correlates audio segments with text tokens more precisely, generating accurate word-level timestamps especially critical for long-form transcription.
Where are the default WhisperAheads values defined in the source code?
Default values are read from the C library via whisper_context_default_params() in Whis.swift (lines 460-466), then reconstructed into Swift WhisperAheads instances using fromC(_:) to maintain synchronization between Swift configuration and native defaults. This ensures the Swift wrapper starts with the same alignment behavior as the underlying C implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →