# How to Run Insanely-Fast-Whisper on Apple Silicon (MPS) for GPU-Accelerated Transcription

> Accelerate Whisper transcription on Apple Silicon Macs. Learn how to run insanely-fast-whisper using MPS for native GPU acceleration with the mps device ID or parameter.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: how-to-guide
- Published: 2026-03-27

---

**Pass `--device-id mps` to the CLI to enable the Metal Performance Shaders backend on M1, M2, or M3 Macs, or set `device="mps"` in the Python pipeline constructor for native GPU acceleration without manual tensor placement.**

Vaibhavs10/insanely-fast-whisper is a command-line wrapper around Hugging Face Transformers that automates Whisper transcription workflows. When running on Apple Silicon, the tool leverages the Metal Performance Shaders (MPS) backend to execute inference directly on the integrated GPU, delivering significantly faster transcription speeds compared to CPU-only execution.

## How MPS Device Routing Works

Device selection is handled transparently in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) (lines 30‑35). The CLI parses the `--device-id` argument and injects the appropriate device string into the Transformers pipeline constructor.

When you specify `--device-id mps`, the code builds the pipeline as follows:

```python
pipe = pipeline(
    "automatic-speech-recognition",
    model=args.model_name,
    torch_dtype=torch.float16,
    device="mps" if args.device_id == "mps" else f"cuda:{args.device_id}",
    model_kwargs={"attn_implementation": "flash_attention_2"} if args.flash else {"attn_implementation": "sdpa"},
)

```

Setting `device="mps"` routes all **torch tensors** to Apple’s Metal backend, utilizing the GPU cores on Apple Silicon chips. The pipeline automatically handles device placement for the Whisper model, tokenizer, and feature extractor.

## Speaker Diarization on Apple Silicon

If you enable speaker diarization, the PyAnnote pipeline is also moved to the MPS device. In [`src/insanely_fast_whisper/utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarization_pipeline.py) (lines 15‑16), the code explicitly casts the diarization model to the same backend:

```python
device = torch.device("mps" if device_id == "mps" else f"cuda:{device_id}")
self.pipeline = pipeline.to(device)

```

This ensures that both the ASR and diarization models reside on the same accelerator, eliminating costly CPU-GPU transfer overhead during the alignment phase implemented in [`utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/diarize.py).

## CLI Usage Examples

### Basic Transcription with MPS

Run transcription on a local audio file or URL using Apple Silicon GPU acceleration:

```bash
insanely-fast-whisper \
    --file-name audio.wav \
    --device-id mps \
    --batch-size 4 \
    --flash True \
    --timestamp word \
    --transcript-path result.json

```

The `--flash True` flag attempts to use **Flash Attention 2**; if the library is unavailable, it gracefully falls back to scaled dot-product attention (SDPA) according to the source logic in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py).

### Transcription with Speaker Diarization

To identify speakers, provide a Hugging Face access token (required for the PyAnnote model) and specify the number of speakers:

```bash
insanely-fast-whisper \
    --file-name audio.wav \
    --device-id mps \
    --hf-token $HF_TOKEN \
    --diarization_model pyannote/speaker-diarization-3.1 \
    --num-speakers 2 \
    --transcript-path result.json

```

The final JSON output—formatted by [`utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/result.py)—contains word-level timestamps and speaker labels mapped to each transcription segment.

## Programmatic Python Usage

You can also invoke the MPS backend directly in Python without the CLI wrapper:

```python
import torch
from transformers import pipeline

pipe = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large-v3",
    torch_dtype=torch.float16,
    device="mps",                          # ← Enables Apple Silicon GPU

    model_kwargs={"attn_implementation": "flash_attention_2"},
)

outputs = pipe(
    "audio.wav",
    chunk_length_s=30,
    batch_size=4,
    return_timestamps="word",
)

print(outputs["text"])

```

This pattern is functionally identical to the CLI’s internal implementation, giving you full control over batch size, chunking, and model selection while retaining MPS acceleration.

## Optimizing Performance on MPS

**Flash Attention 2:** The `--flash` flag is supported on MPS devices. According to the source code, if Flash Attention 2 kernels are not compiled for your environment, the pipeline automatically falls back to `sdpa` (scaled dot-product attention), ensuring compatibility across PyTorch versions.

**Memory Efficiency:** The codebase uses `torch.float16` by default when MPS is active, reducing memory bandwidth pressure on Apple Silicon unified memory architectures. For long audio files, adjust `--batch-size` based on your Mac’s RAM (4‑8 works well for 16 GB systems).

**Diarization overhead:** When using `--num-speakers`, note that the PyAnnote model runs sequentially after ASR. Both stages utilize MPS, but the diarization step may temporarily spike memory usage due to the segmentation algorithm in [`utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/diarize.py).

## Summary

- **Device selection:** Pass `--device-id mps` to the CLI or `device="mps"` in Python to activate the Metal backend.
- **Automatic fallback:** Flash Attention 2 requests gracefully degrade to SDPA if unavailable.
- **Diarization support:** Speaker segmentation runs on MPS when enabled, with device handling in [`diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/diarization_pipeline.py).
- **No manual casting:** The pipeline automatically places all tensors on the Apple Silicon GPU without explicit `.to("mps")` calls.
- **Output format:** Results are serialized to JSON via [`utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/result.py), containing text, timestamps, and optional speaker labels.

## Frequently Asked Questions

### Does insanely-fast-whisper support M1, M2, and M3 Macs natively?

Yes. The tool detects Apple Silicon via PyTorch’s MPS backend and routes computation to the GPU automatically when `--device-id mps` is provided. No Rosetta translation or Docker containers are required.

### Can I use Flash Attention 2 with MPS on Apple Silicon?

Yes, you can pass `--flash True` or set `attn_implementation="flash_attention_2"` in Python. If the Flash Attention 2 library is not installed or lacks MPS kernels, the code falls back to SDPA automatically, maintaining functionality without crashes.

### How do I enable speaker diarization when using MPS?

Supply a Hugging Face token via `--hf-token` and specify `--diarization_model` (e.g., `pyannote/speaker-diarization-3.1`). The diarization pipeline in [`utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/diarization_pipeline.py) automatically moves the PyAnnote model to the MPS device to match the ASR backend.

### Why am I seeing CPU fallback during transcription?

Ensure you are passing `--device-id mps` explicitly; the default device selection may fall back to CPU if the flag is omitted. Also verify that your PyTorch installation supports MPS (`torch.backends.mps.is_available()` should return `True`). If diarization is enabled, confirm that the PyAnnote model successfully loaded onto the GPU by checking activity in Activity Monitor’s GPU tab.