# OpenSuperWhisper Audio Formats: Supported File Types and Conversion Guide

> Discover OpenSuperWhisper supported audio formats like WAV, MP3, and M4A. Learn how it automatically converts files for seamless transcription and get our conversion guide.

- Repository: [Starmel/OpenSuperWhisper](https://github.com/Starmel/OpenSuperWhisper)
- Tags: api-reference
- Published: 2026-07-07

---

**OpenSuperWhisper supports any audio format that macOS or iOS can decode via AVFoundation, including WAV, MP3, M4A/AAC, AIFF, and CAF, with automatic conversion to 16 kHz mono PCM for transcription.**

OpenSuperWhisper is a Swift-based transcription engine that leverages Apple's AVFoundation framework to handle diverse audio inputs. Unlike rigid transcription tools that enforce specific container formats, this open-source project accepts virtually any audio file your system can play. The core architecture delegates format detection and decoding to `AVAudioFile`, ensuring broad compatibility without manual codec management.

## Supported Audio Formats

OpenSuperWhisper does not restrict you to a single hard-coded audio type. The transcription pipeline accepts any file that Apple's `AVAudioFile` class can initialize, which includes the most common consumer formats:

- **WAV** – The default format for recordings created by the app's built-in recorder
- **MP3** – Compressed MPEG audio supported natively by AVFoundation
- **M4A/AAC** – Common Apple-specific containers and codecs
- **AIFF** – Audio Interchange File Format for uncompressed audio
- **CAF** – Core Audio Format supporting various encodings
- **Any AVFoundation-compatible codec** – System-wide decoder support extends to additional formats

## How OpenSuperWhisper Processes Audio Files

The engine's flexibility stems from its reliance on system-level frameworks rather than custom decoders. This approach minimizes code complexity while maximizing format support.

### AVFoundation Integration via AVAudioFile

At the heart of the conversion pipeline is `AVAudioFile`, which handles reading and decoding. According to the OpenSuperWhisper source code in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift), the engine instantiates `AVAudioFile` with the provided URL, allowing the system to handle the underlying codec complexity. This means the same code path processes WAV, MP3, M4A, and other formats without format-specific branches.

### MP4 Container Handling

For container formats that are not natively PCM—specifically MP4-based files—the engine implements minimal plumbing. In [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) (lines 80-96), the code inspects the first few bytes of the file to detect an MP4 header. When detected, it copies the file to a temporary `.m4a` wrapper before feeding it to the converter. This temporary file creation ensures compatibility with `AVAudioFile` while preserving the original data integrity.

### PCM Conversion Pipeline

Once loaded, all audio routes through `convertAudioToPCM`, which standardizes the input to a 16 kHz mono PCM stream. This normalization is crucial for the Whisper model's input requirements. The conversion happens regardless of the source format, ensuring consistent transcription quality whether the input is a high-fidelity AIFF or a compressed MP3.

## Code Examples for Audio Processing

The following Swift examples demonstrate how to work with various audio formats in OpenSuperWhisper:

```swift
import OpenSuperWhisper

// 1️⃣ Record a new WAV file (the app’s built‑in recorder always writes .wav)
let recorder = AudioRecorder()
await recorder.startRecording()   // → creates a temporary file “…/timestamp.wav”
await recorder.stopRecording()
let wavURL = recorder.recordingURL   // URL to a .wav file

```

```swift
// 2️⃣ Transcribe an existing audio file of any supported format
let engine = WhisperEngine()
let pcmSamples = try await engine.convertAudioToPCM(fileURL: wavURL)   // works for .wav, .mp3, .m4a, etc.
let transcription = try await engine.transcribe(pcmSamples)

```

```swift
// 3️⃣ Quick one‑liner for ad‑hoc testing (any supported file)
let result = try await WhisperEngine().transcribeFile(at: URL(fileURLWithPath: "/path/to/audio.mp3"))
print(result.text)

```

The `transcribeFile(at:)` helper method wraps the conversion and transcription calls into a single asynchronous operation.

## Key Implementation Files

Understanding the source architecture helps clarify how OpenSuperWhisper maintains format flexibility:

- **[`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift)** – Located in `OpenSuperWhisper/Engines/`, this file contains the core logic for MP4 header detection (lines 80-96), temporary file creation, and the `convertAudioToPCM` method that standardizes audio to 16 kHz mono PCM.

- **[`AudioUtil.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/AudioUtil.swift)** – This utility file demonstrates format-agnostic handling by using `AVURLAsset` to determine audio duration for any supported URL, proving the app can query metadata without knowing the specific codec beforehand.

- **[`AudioRecorder.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/AudioRecorder.swift)** – Shows the default recording implementation, which always writes WAV files to ensure lossless capture during microphone input.

- **[`TranscriptionQueue.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/TranscriptionQueue.swift)** – Manages transcription job scheduling and relies on `AudioUtil.audioDuration` to queue work, reinforcing the format-agnostic architecture throughout the pipeline.

## Summary

- **OpenSuperWhisper supports any AVFoundation-compatible format**, including WAV, MP3, M4A/AAC, AIFF, and CAF.
- **MP4-based files** receive special handling via temporary wrapper creation in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) before processing.
- **All audio converts to 16 kHz mono PCM** through the `convertAudioToPCM` method, ensuring consistent Whisper model input.
- **Built-in recording** defaults to WAV format for maximum quality during capture.
- **System-level decoding** via `AVAudioFile` eliminates the need for manual codec management or third-party dependencies.

## Frequently Asked Questions

### Does OpenSuperWhisper require audio files to be in a specific format?

No. OpenSuperWhisper accepts any audio format that macOS or iOS can decode through AVFoundation. While the built-in recorder creates WAV files, the transcription engine handles MP3, M4A, AIFF, CAF, and other formats automatically without requiring manual conversion.

### How does OpenSuperWhisper handle compressed audio like MP3 or M4A?

The engine uses `AVAudioFile` to decode compressed formats automatically. For MP4-based containers, it first checks the file header in [`WhisperEngine.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/WhisperEngine.swift) and creates a temporary `.m4a` wrapper if needed. All audio then passes through `convertAudioToPCM` to normalize to 16 kHz mono PCM suitable for the Whisper model.

### What audio format does the built-in recorder use?

The built-in recorder in [`AudioRecorder.swift`](https://github.com/Starmel/OpenSuperWhisper/blob/main/AudioRecorder.swift) always writes WAV files. This ensures lossless, uncompressed capture during recording sessions, though you can transcribe the resulting files alongside other supported formats.

### Can I transcribe audio files created by other applications?

Yes. OpenSuperWhisper can transcribe audio files from any source, including voice memos, screen recordings, or third-party recording apps, provided the format is supported by Apple's AVFoundation framework. The `transcribeFile(at:)` method handles format detection and conversion automatically.