How to Run Insanely-Fast-Whisper on Apple Silicon (MPS) for GPU-Accelerated Transcription
Pass --device-id mps to the CLI to enable the Metal Performance Shaders backend on M1, M2, or M3 Macs, or set device="mps" in the Python pipeline constructor for native GPU acceleration without manual tensor placement.
Vaibhavs10/insanely-fast-whisper is a command-line wrapper around Hugging Face Transformers that automates Whisper transcription workflows. When running on Apple Silicon, the tool leverages the Metal Performance Shaders (MPS) backend to execute inference directly on the integrated GPU, delivering significantly faster transcription speeds compared to CPU-only execution.
How MPS Device Routing Works
Device selection is handled transparently in src/insanely_fast_whisper/cli.py (lines 30‑35). The CLI parses the --device-id argument and injects the appropriate device string into the Transformers pipeline constructor.
When you specify --device-id mps, the code builds the pipeline as follows:
pipe = pipeline(
"automatic-speech-recognition",
model=args.model_name,
torch_dtype=torch.float16,
device="mps" if args.device_id == "mps" else f"cuda:{args.device_id}",
model_kwargs={"attn_implementation": "flash_attention_2"} if args.flash else {"attn_implementation": "sdpa"},
)
Setting device="mps" routes all torch tensors to Apple’s Metal backend, utilizing the GPU cores on Apple Silicon chips. The pipeline automatically handles device placement for the Whisper model, tokenizer, and feature extractor.
Speaker Diarization on Apple Silicon
If you enable speaker diarization, the PyAnnote pipeline is also moved to the MPS device. In src/insanely_fast_whisper/utils/diarization_pipeline.py (lines 15‑16), the code explicitly casts the diarization model to the same backend:
device = torch.device("mps" if device_id == "mps" else f"cuda:{device_id}")
self.pipeline = pipeline.to(device)
This ensures that both the ASR and diarization models reside on the same accelerator, eliminating costly CPU-GPU transfer overhead during the alignment phase implemented in utils/diarize.py.
CLI Usage Examples
Basic Transcription with MPS
Run transcription on a local audio file or URL using Apple Silicon GPU acceleration:
insanely-fast-whisper \
--file-name audio.wav \
--device-id mps \
--batch-size 4 \
--flash True \
--timestamp word \
--transcript-path result.json
The --flash True flag attempts to use Flash Attention 2; if the library is unavailable, it gracefully falls back to scaled dot-product attention (SDPA) according to the source logic in cli.py.
Transcription with Speaker Diarization
To identify speakers, provide a Hugging Face access token (required for the PyAnnote model) and specify the number of speakers:
insanely-fast-whisper \
--file-name audio.wav \
--device-id mps \
--hf-token $HF_TOKEN \
--diarization_model pyannote/speaker-diarization-3.1 \
--num-speakers 2 \
--transcript-path result.json
The final JSON output—formatted by utils/result.py—contains word-level timestamps and speaker labels mapped to each transcription segment.
Programmatic Python Usage
You can also invoke the MPS backend directly in Python without the CLI wrapper:
import torch
from transformers import pipeline
pipe = pipeline(
"automatic-speech-recognition",
model="openai/whisper-large-v3",
torch_dtype=torch.float16,
device="mps", # ← Enables Apple Silicon GPU
model_kwargs={"attn_implementation": "flash_attention_2"},
)
outputs = pipe(
"audio.wav",
chunk_length_s=30,
batch_size=4,
return_timestamps="word",
)
print(outputs["text"])
This pattern is functionally identical to the CLI’s internal implementation, giving you full control over batch size, chunking, and model selection while retaining MPS acceleration.
Optimizing Performance on MPS
Flash Attention 2: The --flash flag is supported on MPS devices. According to the source code, if Flash Attention 2 kernels are not compiled for your environment, the pipeline automatically falls back to sdpa (scaled dot-product attention), ensuring compatibility across PyTorch versions.
Memory Efficiency: The codebase uses torch.float16 by default when MPS is active, reducing memory bandwidth pressure on Apple Silicon unified memory architectures. For long audio files, adjust --batch-size based on your Mac’s RAM (4‑8 works well for 16 GB systems).
Diarization overhead: When using --num-speakers, note that the PyAnnote model runs sequentially after ASR. Both stages utilize MPS, but the diarization step may temporarily spike memory usage due to the segmentation algorithm in utils/diarize.py.
Summary
- Device selection: Pass
--device-id mpsto the CLI ordevice="mps"in Python to activate the Metal backend. - Automatic fallback: Flash Attention 2 requests gracefully degrade to SDPA if unavailable.
- Diarization support: Speaker segmentation runs on MPS when enabled, with device handling in
diarization_pipeline.py. - No manual casting: The pipeline automatically places all tensors on the Apple Silicon GPU without explicit
.to("mps")calls. - Output format: Results are serialized to JSON via
utils/result.py, containing text, timestamps, and optional speaker labels.
Frequently Asked Questions
Does insanely-fast-whisper support M1, M2, and M3 Macs natively?
Yes. The tool detects Apple Silicon via PyTorch’s MPS backend and routes computation to the GPU automatically when --device-id mps is provided. No Rosetta translation or Docker containers are required.
Can I use Flash Attention 2 with MPS on Apple Silicon?
Yes, you can pass --flash True or set attn_implementation="flash_attention_2" in Python. If the Flash Attention 2 library is not installed or lacks MPS kernels, the code falls back to SDPA automatically, maintaining functionality without crashes.
How do I enable speaker diarization when using MPS?
Supply a Hugging Face token via --hf-token and specify --diarization_model (e.g., pyannote/speaker-diarization-3.1). The diarization pipeline in utils/diarization_pipeline.py automatically moves the PyAnnote model to the MPS device to match the ASR backend.
Why am I seeing CPU fallback during transcription?
Ensure you are passing --device-id mps explicitly; the default device selection may fall back to CPU if the flag is omitted. Also verify that your PyTorch installation supports MPS (torch.backends.mps.is_available() should return True). If diarization is enabled, confirm that the PyAnnote model successfully loaded onto the GPU by checking activity in Activity Monitor’s GPU tab.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →