OpenSuperWhisper Audio Formats: Supported File Types and Conversion Guide

OpenSuperWhisper supports any audio format that macOS or iOS can decode via AVFoundation, including WAV, MP3, M4A/AAC, AIFF, and CAF, with automatic conversion to 16 kHz mono PCM for transcription.

OpenSuperWhisper is a Swift-based transcription engine that leverages Apple's AVFoundation framework to handle diverse audio inputs. Unlike rigid transcription tools that enforce specific container formats, this open-source project accepts virtually any audio file your system can play. The core architecture delegates format detection and decoding to AVAudioFile, ensuring broad compatibility without manual codec management.

Supported Audio Formats

OpenSuperWhisper does not restrict you to a single hard-coded audio type. The transcription pipeline accepts any file that Apple's AVAudioFile class can initialize, which includes the most common consumer formats:

  • WAV – The default format for recordings created by the app's built-in recorder
  • MP3 – Compressed MPEG audio supported natively by AVFoundation
  • M4A/AAC – Common Apple-specific containers and codecs
  • AIFF – Audio Interchange File Format for uncompressed audio
  • CAF – Core Audio Format supporting various encodings
  • Any AVFoundation-compatible codec – System-wide decoder support extends to additional formats

How OpenSuperWhisper Processes Audio Files

The engine's flexibility stems from its reliance on system-level frameworks rather than custom decoders. This approach minimizes code complexity while maximizing format support.

AVFoundation Integration via AVAudioFile

At the heart of the conversion pipeline is AVAudioFile, which handles reading and decoding. According to the OpenSuperWhisper source code in WhisperEngine.swift, the engine instantiates AVAudioFile with the provided URL, allowing the system to handle the underlying codec complexity. This means the same code path processes WAV, MP3, M4A, and other formats without format-specific branches.

MP4 Container Handling

For container formats that are not natively PCM—specifically MP4-based files—the engine implements minimal plumbing. In WhisperEngine.swift (lines 80-96), the code inspects the first few bytes of the file to detect an MP4 header. When detected, it copies the file to a temporary .m4a wrapper before feeding it to the converter. This temporary file creation ensures compatibility with AVAudioFile while preserving the original data integrity.

PCM Conversion Pipeline

Once loaded, all audio routes through convertAudioToPCM, which standardizes the input to a 16 kHz mono PCM stream. This normalization is crucial for the Whisper model's input requirements. The conversion happens regardless of the source format, ensuring consistent transcription quality whether the input is a high-fidelity AIFF or a compressed MP3.

Code Examples for Audio Processing

The following Swift examples demonstrate how to work with various audio formats in OpenSuperWhisper:

import OpenSuperWhisper

// 1️⃣ Record a new WAV file (the app’s built‑in recorder always writes .wav)
let recorder = AudioRecorder()
await recorder.startRecording()   // → creates a temporary file “…/timestamp.wav”
await recorder.stopRecording()
let wavURL = recorder.recordingURL   // URL to a .wav file
// 2️⃣ Transcribe an existing audio file of any supported format
let engine = WhisperEngine()
let pcmSamples = try await engine.convertAudioToPCM(fileURL: wavURL)   // works for .wav, .mp3, .m4a, etc.
let transcription = try await engine.transcribe(pcmSamples)
// 3️⃣ Quick one‑liner for ad‑hoc testing (any supported file)
let result = try await WhisperEngine().transcribeFile(at: URL(fileURLWithPath: "/path/to/audio.mp3"))
print(result.text)

The transcribeFile(at:) helper method wraps the conversion and transcription calls into a single asynchronous operation.

Key Implementation Files

Understanding the source architecture helps clarify how OpenSuperWhisper maintains format flexibility:

  • WhisperEngine.swift – Located in OpenSuperWhisper/Engines/, this file contains the core logic for MP4 header detection (lines 80-96), temporary file creation, and the convertAudioToPCM method that standardizes audio to 16 kHz mono PCM.

  • AudioUtil.swift – This utility file demonstrates format-agnostic handling by using AVURLAsset to determine audio duration for any supported URL, proving the app can query metadata without knowing the specific codec beforehand.

  • AudioRecorder.swift – Shows the default recording implementation, which always writes WAV files to ensure lossless capture during microphone input.

  • TranscriptionQueue.swift – Manages transcription job scheduling and relies on AudioUtil.audioDuration to queue work, reinforcing the format-agnostic architecture throughout the pipeline.

Summary

  • OpenSuperWhisper supports any AVFoundation-compatible format, including WAV, MP3, M4A/AAC, AIFF, and CAF.
  • MP4-based files receive special handling via temporary wrapper creation in WhisperEngine.swift before processing.
  • All audio converts to 16 kHz mono PCM through the convertAudioToPCM method, ensuring consistent Whisper model input.
  • Built-in recording defaults to WAV format for maximum quality during capture.
  • System-level decoding via AVAudioFile eliminates the need for manual codec management or third-party dependencies.

Frequently Asked Questions

Does OpenSuperWhisper require audio files to be in a specific format?

No. OpenSuperWhisper accepts any audio format that macOS or iOS can decode through AVFoundation. While the built-in recorder creates WAV files, the transcription engine handles MP3, M4A, AIFF, CAF, and other formats automatically without requiring manual conversion.

How does OpenSuperWhisper handle compressed audio like MP3 or M4A?

The engine uses AVAudioFile to decode compressed formats automatically. For MP4-based containers, it first checks the file header in WhisperEngine.swift and creates a temporary .m4a wrapper if needed. All audio then passes through convertAudioToPCM to normalize to 16 kHz mono PCM suitable for the Whisper model.

What audio format does the built-in recorder use?

The built-in recorder in AudioRecorder.swift always writes WAV files. This ensures lossless, uncompressed capture during recording sessions, though you can transcribe the resulting files alongside other supported formats.

Can I transcribe audio files created by other applications?

Yes. OpenSuperWhisper can transcribe audio files from any source, including voice memos, screen recordings, or third-party recording apps, provided the format is supported by Apple's AVFoundation framework. The transcribeFile(at:) method handles format detection and conversion automatically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →