Resolving CUDA Out-of-Memory Errors with Insanely-Fast-Whisper: Optimization Guide

Lower the --batch-size parameter to reduce parallel processing load, enable Flash Attention 2 with --flash True, and ensure you're using the optimal device backend to prevent CUDA out-of-memory crashes while transcribing audio.

The insanely-fast-whisper project by Vaibhavs10 accelerates OpenAI Whisper transcription by leveraging the 🤗 Transformers pipeline with batched processing on CUDA or Apple Silicon. While this delivers significant speed improvements, the default configuration can exhaust GPU memory on consumer cards with 8 GB or 12 GB of VRAM. Understanding how to tune the pipeline's memory parameters in src/insanely_fast_whisper/cli.py is essential for stable inference on limited hardware.

Understanding GPU Memory Consumption

The pipeline loads Whisper models in FP16 (torch_dtype=torch.float16) and streams audio in 30-second chunks (chunk_length_s=30). It parallelizes these chunks using a configurable batch_size, where each batch holds simultaneous forward passes in memory. The default --batch-size 24 creates 24 parallel audio chunks during inference, multiplying the intermediate activation memory required for each forward pass. On GPUs with limited VRAM, this parallel processing exceeds available memory, triggering the CUDA out-of-memory error before transcription completes.

Three Methods to Fix CUDA OOM Errors

Reduce the Batch Size

The most direct lever for controlling memory is the --batch-size argument defined in [cli.py](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L55-L63). Lowering this value reduces the number of simultaneous forward passes and cuts peak VRAM usage roughly linearly.

For CUDA GPUs with 8–12 GB of VRAM, reduce the batch size from the default 24 to 4 or 8:

insanely-fast-whisper \
  --file-name my_audio.wav \
  --device-id 0 \
  --batch-size 4 \
  --model-name openai/whisper-large-v3 \
  --transcript-path result.json

Enable Flash Attention 2

Flash Attention 2 implements a memory-efficient attention kernel that reduces the memory overhead compared to the default SDPA implementation. Enable this by passing --flash True, which sets attn_implementation="flash_attention_2" in the pipeline's model_kwargs as implemented in [cli.py](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L30-L36).

This optimization often allows you to maintain higher batch sizes without encountering OOM errors:

insanely-fast-whisper \
  --file-name my_audio.wav \
  --device-id 0 \
  --flash True \
  --batch-size 24 \
  --model-name openai/whisper-large-v3 \
  --transcript-path result.json

Select the Appropriate Device Backend

The device selection logic in [cli.py](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L30-L35) supports both CUDA and Apple Silicon (MPS). However, the MPS backend is currently less optimized and consumes more VRAM than CUDA for equivalent batch sizes. If you have multiple CUDA devices, specify the less-utilized GPU with --device-id <gpu-id>.

For macOS users experiencing OOM errors, the documentation recommends starting with --batch-size 4 regardless of available memory, as MPS requires approximately 12 GB of VRAM even for modest batch sizes:

insanely-fast-whisper \
  --file-name my_audio.wav \
  --device-id mps \
  --batch-size 4 \
  --model-name openai/whisper-large-v3 \
  --transcript-path result.json

Python Pipeline Implementation

If you prefer scripting over the CLI, replicate these memory optimizations directly in the Transformers pipeline:

import torch
from transformers import pipeline

pipe = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large-v3",
    torch_dtype=torch.float16,
    device="cuda:0",                     # or "mps" for Mac

    model_kwargs={"attn_implementation": "flash_attention_2"},
)

outputs = pipe(
    "my_audio.wav",
    chunk_length_s=30,
    batch_size=4,                        # adjust to fit VRAM

    return_timestamps=True,
)

Key Source Files

Understanding the codebase structure helps when debugging memory issues:

  • src/insanely_fast_whisper/cli.py: Parses command-line arguments (--batch-size, --flash, --device-id) and constructs the Whisper pipeline with the specified optimizations.
  • src/insanely_fast_whisper/utils/result.py: Formats the pipeline output into the JSON structure written to --transcript-path.
  • README.md: Documents additional OOM avoidance strategies and hardware-specific recommendations.
  • pyproject.toml: Declares dependencies including Transformers and Torch that control CUDA support and memory management.

Summary

  • Batch size is the primary control for memory usage; reduce --batch-size from 24 to 4–8 for GPUs with limited VRAM.
  • Flash Attention 2 significantly reduces memory footprint without sacrificing speed; enable with --flash True in the CLI or attn_implementation="flash_attention_2" in Python.
  • Device selection matters; MPS consumes more memory than CUDA, requiring lower batch sizes on Apple Silicon.
  • Memory consumption scales with parallel chunks (batch size × activation memory), not just model size.

Frequently Asked Questions

What is the default batch size in insanely-fast-whisper?

The CLI defaults to --batch-size 24 as defined in [src/insanely_fast_whisper/cli.py](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L55-L63). This value is optimized for speed on high-end GPUs with 16 GB or more of VRAM, but it typically causes out-of-memory errors on consumer cards with 8 GB or 12 GB.

Does enabling Flash Attention 2 reduce transcription quality?

No, Flash Attention 2 is a memory-efficient implementation of the same attention mechanism that produces numerically identical results to the default SDPA implementation. It changes only the computational kernel, not the model weights or inference logic, so transcription accuracy remains unchanged while VRAM usage decreases.

Why does the MPS backend require smaller batch sizes than CUDA?

The Metal Performance Shaders (MPS) backend for Apple Silicon is currently less optimized for the Whisper architecture's memory access patterns compared to NVIDIA's CUDA implementation. As documented in the repository's README, MPS requires approximately 12 GB of VRAM even when using --batch-size 4, whereas CUDA can often handle larger batches with equivalent or less memory.

Will reducing batch size make transcription slower?

Yes, lowering the batch size reduces GPU utilization by decreasing the number of parallel audio chunks processed simultaneously. This increases total transcription time, but combining a moderately reduced batch size with Flash Attention 2 typically balances speed and memory safety for consumer hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →