Resolving CUDA Out-of-Memory Errors with Insanely-Fast-Whisper: Optimization Guide
Lower the --batch-size parameter to reduce parallel processing load, enable Flash Attention 2 with --flash True, and ensure you're using the optimal device backend to prevent CUDA out-of-memory crashes while transcribing audio.
The insanely-fast-whisper project by Vaibhavs10 accelerates OpenAI Whisper transcription by leveraging the 🤗 Transformers pipeline with batched processing on CUDA or Apple Silicon. While this delivers significant speed improvements, the default configuration can exhaust GPU memory on consumer cards with 8 GB or 12 GB of VRAM. Understanding how to tune the pipeline's memory parameters in src/insanely_fast_whisper/cli.py is essential for stable inference on limited hardware.
Understanding GPU Memory Consumption
The pipeline loads Whisper models in FP16 (torch_dtype=torch.float16) and streams audio in 30-second chunks (chunk_length_s=30). It parallelizes these chunks using a configurable batch_size, where each batch holds simultaneous forward passes in memory. The default --batch-size 24 creates 24 parallel audio chunks during inference, multiplying the intermediate activation memory required for each forward pass. On GPUs with limited VRAM, this parallel processing exceeds available memory, triggering the CUDA out-of-memory error before transcription completes.
Three Methods to Fix CUDA OOM Errors
Reduce the Batch Size
The most direct lever for controlling memory is the --batch-size argument defined in [cli.py](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L55-L63). Lowering this value reduces the number of simultaneous forward passes and cuts peak VRAM usage roughly linearly.
For CUDA GPUs with 8–12 GB of VRAM, reduce the batch size from the default 24 to 4 or 8:
insanely-fast-whisper \
--file-name my_audio.wav \
--device-id 0 \
--batch-size 4 \
--model-name openai/whisper-large-v3 \
--transcript-path result.json
Enable Flash Attention 2
Flash Attention 2 implements a memory-efficient attention kernel that reduces the memory overhead compared to the default SDPA implementation. Enable this by passing --flash True, which sets attn_implementation="flash_attention_2" in the pipeline's model_kwargs as implemented in [cli.py](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L30-L36).
This optimization often allows you to maintain higher batch sizes without encountering OOM errors:
insanely-fast-whisper \
--file-name my_audio.wav \
--device-id 0 \
--flash True \
--batch-size 24 \
--model-name openai/whisper-large-v3 \
--transcript-path result.json
Select the Appropriate Device Backend
The device selection logic in [cli.py](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L30-L35) supports both CUDA and Apple Silicon (MPS). However, the MPS backend is currently less optimized and consumes more VRAM than CUDA for equivalent batch sizes. If you have multiple CUDA devices, specify the less-utilized GPU with --device-id <gpu-id>.
For macOS users experiencing OOM errors, the documentation recommends starting with --batch-size 4 regardless of available memory, as MPS requires approximately 12 GB of VRAM even for modest batch sizes:
insanely-fast-whisper \
--file-name my_audio.wav \
--device-id mps \
--batch-size 4 \
--model-name openai/whisper-large-v3 \
--transcript-path result.json
Python Pipeline Implementation
If you prefer scripting over the CLI, replicate these memory optimizations directly in the Transformers pipeline:
import torch
from transformers import pipeline
pipe = pipeline(
"automatic-speech-recognition",
model="openai/whisper-large-v3",
torch_dtype=torch.float16,
device="cuda:0", # or "mps" for Mac
model_kwargs={"attn_implementation": "flash_attention_2"},
)
outputs = pipe(
"my_audio.wav",
chunk_length_s=30,
batch_size=4, # adjust to fit VRAM
return_timestamps=True,
)
Key Source Files
Understanding the codebase structure helps when debugging memory issues:
src/insanely_fast_whisper/cli.py: Parses command-line arguments (--batch-size,--flash,--device-id) and constructs the Whisper pipeline with the specified optimizations.src/insanely_fast_whisper/utils/result.py: Formats the pipeline output into the JSON structure written to--transcript-path.README.md: Documents additional OOM avoidance strategies and hardware-specific recommendations.pyproject.toml: Declares dependencies including Transformers and Torch that control CUDA support and memory management.
Summary
- Batch size is the primary control for memory usage; reduce
--batch-sizefrom 24 to 4–8 for GPUs with limited VRAM. - Flash Attention 2 significantly reduces memory footprint without sacrificing speed; enable with
--flash Truein the CLI orattn_implementation="flash_attention_2"in Python. - Device selection matters; MPS consumes more memory than CUDA, requiring lower batch sizes on Apple Silicon.
- Memory consumption scales with parallel chunks (batch size × activation memory), not just model size.
Frequently Asked Questions
What is the default batch size in insanely-fast-whisper?
The CLI defaults to --batch-size 24 as defined in [src/insanely_fast_whisper/cli.py](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L55-L63). This value is optimized for speed on high-end GPUs with 16 GB or more of VRAM, but it typically causes out-of-memory errors on consumer cards with 8 GB or 12 GB.
Does enabling Flash Attention 2 reduce transcription quality?
No, Flash Attention 2 is a memory-efficient implementation of the same attention mechanism that produces numerically identical results to the default SDPA implementation. It changes only the computational kernel, not the model weights or inference logic, so transcription accuracy remains unchanged while VRAM usage decreases.
Why does the MPS backend require smaller batch sizes than CUDA?
The Metal Performance Shaders (MPS) backend for Apple Silicon is currently less optimized for the Whisper architecture's memory access patterns compared to NVIDIA's CUDA implementation. As documented in the repository's README, MPS requires approximately 12 GB of VRAM even when using --batch-size 4, whereas CUDA can often handle larger batches with equivalent or less memory.
Will reducing batch size make transcription slower?
Yes, lowering the batch size reduces GPU utilization by decreasing the number of parallel audio chunks processed simultaneously. This increases total transcription time, but combining a moderately reduced batch size with Flash Attention 2 typically balances speed and memory safety for consumer hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →