# Resolving CUDA Out-of-Memory Errors with Insanely-Fast-Whisper: Optimization Guide

> Fix CUDA out of memory errors in insanely-fast-whisper by lowering batch size, enabling Flash Attention 2, and optimizing your device. Transcribe audio without crashes.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: optimization-guide
- Published: 2026-03-27

---

**Lower the `--batch-size` parameter to reduce parallel processing load, enable Flash Attention 2 with `--flash True`, and ensure you're using the optimal device backend to prevent CUDA out-of-memory crashes while transcribing audio.**

The `insanely-fast-whisper` project by Vaibhavs10 accelerates OpenAI Whisper transcription by leveraging the 🤗 Transformers pipeline with batched processing on CUDA or Apple Silicon. While this delivers significant speed improvements, the default configuration can exhaust GPU memory on consumer cards with 8 GB or 12 GB of VRAM. Understanding how to tune the pipeline's memory parameters in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) is essential for stable inference on limited hardware.

## Understanding GPU Memory Consumption

The pipeline loads Whisper models in **FP16** (`torch_dtype=torch.float16`) and streams audio in 30-second chunks (`chunk_length_s=30`). It parallelizes these chunks using a configurable `batch_size`, where each batch holds simultaneous forward passes in memory. The default `--batch-size 24` creates 24 parallel audio chunks during inference, multiplying the intermediate activation memory required for each forward pass. On GPUs with limited VRAM, this parallel processing exceeds available memory, triggering the **CUDA out-of-memory error** before transcription completes.

## Three Methods to Fix CUDA OOM Errors

### Reduce the Batch Size

The most direct lever for controlling memory is the `--batch-size` argument defined in [[`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py)](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L55-L63). Lowering this value reduces the number of simultaneous forward passes and cuts peak VRAM usage roughly linearly.

For CUDA GPUs with 8–12 GB of VRAM, reduce the batch size from the default 24 to 4 or 8:

```bash
insanely-fast-whisper \
  --file-name my_audio.wav \
  --device-id 0 \
  --batch-size 4 \
  --model-name openai/whisper-large-v3 \
  --transcript-path result.json

```

### Enable Flash Attention 2

Flash Attention 2 implements a memory-efficient attention kernel that reduces the memory overhead compared to the default SDPA implementation. Enable this by passing `--flash True`, which sets `attn_implementation="flash_attention_2"` in the pipeline's `model_kwargs` as implemented in [[`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py)](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L30-L36).

This optimization often allows you to maintain higher batch sizes without encountering OOM errors:

```bash
insanely-fast-whisper \
  --file-name my_audio.wav \
  --device-id 0 \
  --flash True \
  --batch-size 24 \
  --model-name openai/whisper-large-v3 \
  --transcript-path result.json

```

### Select the Appropriate Device Backend

The device selection logic in [[`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py)](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L30-L35) supports both CUDA and Apple Silicon (MPS). However, the MPS backend is currently less optimized and consumes more VRAM than CUDA for equivalent batch sizes. If you have multiple CUDA devices, specify the less-utilized GPU with `--device-id <gpu-id>`.

For macOS users experiencing OOM errors, the documentation recommends starting with `--batch-size 4` regardless of available memory, as MPS requires approximately 12 GB of VRAM even for modest batch sizes:

```bash
insanely-fast-whisper \
  --file-name my_audio.wav \
  --device-id mps \
  --batch-size 4 \
  --model-name openai/whisper-large-v3 \
  --transcript-path result.json

```

## Python Pipeline Implementation

If you prefer scripting over the CLI, replicate these memory optimizations directly in the Transformers pipeline:

```python
import torch
from transformers import pipeline

pipe = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large-v3",
    torch_dtype=torch.float16,
    device="cuda:0",                     # or "mps" for Mac

    model_kwargs={"attn_implementation": "flash_attention_2"},
)

outputs = pipe(
    "my_audio.wav",
    chunk_length_s=30,
    batch_size=4,                        # adjust to fit VRAM

    return_timestamps=True,
)

```

## Key Source Files

Understanding the codebase structure helps when debugging memory issues:

- **[`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py)**: Parses command-line arguments (`--batch-size`, `--flash`, `--device-id`) and constructs the Whisper pipeline with the specified optimizations.
- **[`src/insanely_fast_whisper/utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/result.py)**: Formats the pipeline output into the JSON structure written to `--transcript-path`.
- **[`README.md`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/README.md)**: Documents additional OOM avoidance strategies and hardware-specific recommendations.
- **[`pyproject.toml`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/pyproject.toml)**: Declares dependencies including Transformers and Torch that control CUDA support and memory management.

## Summary

- **Batch size** is the primary control for memory usage; reduce `--batch-size` from 24 to 4–8 for GPUs with limited VRAM.
- **Flash Attention 2** significantly reduces memory footprint without sacrificing speed; enable with `--flash True` in the CLI or `attn_implementation="flash_attention_2"` in Python.
- **Device selection** matters; MPS consumes more memory than CUDA, requiring lower batch sizes on Apple Silicon.
- Memory consumption scales with parallel chunks (batch size × activation memory), not just model size.

## Frequently Asked Questions

### What is the default batch size in insanely-fast-whisper?

The CLI defaults to `--batch-size 24` as defined in [[`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py)](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L55-L63). This value is optimized for speed on high-end GPUs with 16 GB or more of VRAM, but it typically causes out-of-memory errors on consumer cards with 8 GB or 12 GB.

### Does enabling Flash Attention 2 reduce transcription quality?

No, Flash Attention 2 is a memory-efficient implementation of the same attention mechanism that produces numerically identical results to the default SDPA implementation. It changes only the computational kernel, not the model weights or inference logic, so transcription accuracy remains unchanged while VRAM usage decreases.

### Why does the MPS backend require smaller batch sizes than CUDA?

The Metal Performance Shaders (MPS) backend for Apple Silicon is currently less optimized for the Whisper architecture's memory access patterns compared to NVIDIA's CUDA implementation. As documented in the repository's README, MPS requires approximately 12 GB of VRAM even when using `--batch-size 4`, whereas CUDA can often handle larger batches with equivalent or less memory.

### Will reducing batch size make transcription slower?

Yes, lowering the batch size reduces GPU utilization by decreasing the number of parallel audio chunks processed simultaneously. This increases total transcription time, but combining a moderately reduced batch size with Flash Attention 2 typically balances speed and memory safety for consumer hardware.