How to Configure Batch Size to Avoid OOM Errors on Different GPUs with Insanely Fast Whisper

Set --batch-size based on your GPU VRAM: use 8–12 for ≤6 GB cards, 12–20 for 8–12 GB cards, and 24–48 for ≥16 GB cards to prevent out-of-memory crashes while maintaining optimal transcription throughput.

Insanely Fast Whisper accelerates OpenAI’s Whisper model by processing audio in parallel batches, but this parallelism consumes GPU memory rapidly. Configuring the correct batch size for your specific GPU is essential to prevent CUDA out-of-memory (OOM) errors while maximizing transcription speed. The tool exposes a --batch-size parameter in src/insanely_fast_whisper/cli.py that directly controls how many 30-second audio chunks are processed simultaneously.

Understanding the --batch-size Parameter

In src/insanely_fast_whisper/cli.py, the CLI defines the --batch-size argument at lines 53‑58 with a default value of 24. This value is passed directly to the Hugging Face pipeline call at lines 159‑164, where the Whisper model processes audio chunks of 30 seconds each (chunk_length_s=30).

outputs = pipe(
    args.file_name,
    chunk_length_s=30,
    batch_size=args.batch_size,   # Controls GPU memory consumption

    generate_kwargs=generate_kwargs,
    return_timestamps=ts,
)

The batch_size parameter determines how many audio chunks are kept in GPU memory simultaneously during the forward pass. Larger batches improve GPU utilization and throughput but require proportionally more VRAM.

GPU memory capacity is the primary constraint when configuring batch size. The following recommendations balance throughput against memory safety for different hardware tiers.

GPUs with ≤ 6 GB VRAM (e.g., GTX 1660, RTX 3050):

  • Use --batch-size 8 to --batch-size 12
  • This keeps the per-batch tensor footprint low enough to avoid OOM while still leveraging modest parallelism.

GPUs with 8 GB–12 GB VRAM (e.g., RTX 3060, RTX 3070):

  • Use --batch-size 12 to --batch-size 20
  • You can safely test up to 24 when using smaller Whisper models or shorter audio files.

GPUs with ≥ 16 GB VRAM (e.g., RTX 3080, RTX 3090, A100):

  • Use --batch-size 24 to --batch-size 48
  • High VRAM headroom allows aggressive batching for maximum throughput.

Apple Silicon (MPS device):

  • Use --batch-size 4 to --batch-size 12
  • The MPS backend has different memory semantics than CUDA; start low and increment gradually until stable.

Interactions with Other Performance Flags

Batch size does not operate in isolation. Several other flags in cli.py affect GPU memory usage and must be considered when tuning.

Flash Attention (--flash):

Enabling Flash Attention via --flash true switches the model to the memory-efficient flash_attention_2 implementation. This can allow a larger batch size on the same GPU without triggering OOM errors.

Model Size:

Larger Whisper checkpoints (e.g., openai/whisper-large-v3) consume significantly more VRAM per batch than base or small variants. When upgrading to a larger model, reduce --batch-size proportionally.

Device Selection (--device-id):

The batch size logic applies uniformly across CUDA devices (cuda:0, cuda:1) and Apple Silicon (mps). However, multi-GPU setups require setting --device-id to specify which GPU handles the batch.

Practical Configuration Examples

Use these CLI patterns as starting points for your hardware.

8 GB GPU without Flash Attention:

python -m insanely_fast_whisper.cli \
    --file-name podcast.wav \
    --device-id 0 \
    --batch-size 16 \
    --model-name openai/whisper-large-v3 \
    --timestamp word

12 GB GPU with Flash Attention (allows larger batch):

python -m insanely_fast_whisper.cli \
    --file-name interview.mp3 \
    --device-id 0 \
    --batch-size 24 \
    --flash true \
    --timestamp chunk

Automated Batch Size Search:

If you are unsure of your GPU limit, script a descending search to find the maximum stable batch size:

for b in 24 20 16 12 8; do
  echo "Testing batch size $b"
  python -m insanely_fast_whisper.cli --file-name audio.wav --batch-size $b
  if [ $? -eq 0 ]; then break; fi
done

Monitor VRAM usage in real-time with watch -n 1 nvidia-smi (CUDA) or Activity Monitor (macOS) while testing.

Summary

  • The --batch-size flag in src/insanely_fast_whisper/cli.py defaults to 24 and controls how many 30-second audio chunks are processed in parallel on the GPU.
  • OOM errors occur when batch size exceeds available GPU memory; reduce the value incrementally until the process runs stable.
  • Memory guidelines: Use 8–12 for ≤6 GB GPUs, 12–20 for 8–12 GB GPUs, 24–48 for ≥16 GB GPUs, and 4–12 for Apple Silicon.
  • Flash Attention (--flash) reduces memory pressure per chunk, potentially allowing higher batch sizes on the same hardware.
  • Related files like src/insanely_fast_whisper/utils/result.py and diarization_pipeline.py handle post-processing and do not affect GPU memory during the batch transcription stage.

Frequently Asked Questions

What is the default batch size in insanely-fast-whisper?

The default batch size is 24, defined in the CLI argument parser at lines 53‑58 of src/insanely_fast_whisper/cli.py. This default assumes a GPU with at least 12 GB of VRAM; users with less memory should reduce this value immediately to avoid OOM errors.

How do I know if my batch size is too large?

The process will crash with a CUDA out of memory error or a segmentation fault during the forward pass. To prevent this, monitor nvidia-smi (for CUDA) or Activity Monitor (for MPS) while processing a test file; if memory utilization approaches 100%, lower --batch-size before running the full workload.

Does enabling Flash Attention allow me to use a larger batch size?

Yes. When you pass --flash true, the pipeline uses the flash_attention_2 implementation, which has a reduced memory footprint per attention operation. This efficiency often allows you to increase --batch-size by 4–8 units on the same GPU without triggering OOM errors.

Can I use the same batch size on Apple Silicon (MPS) as on NVIDIA GPUs?

No. Apple Silicon devices using the mps backend typically require smaller batch sizes (4–12) compared to CUDA GPUs with equivalent nominal VRAM due to different memory management semantics in PyTorch’s MPS backend. Start with --batch-size 4 and increase gradually while monitoring system memory pressure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →