# How to Configure Batch Size to Avoid OOM Errors on Different GPUs with Insanely Fast Whisper

> Configure batch size to avoid OOM errors on different GPUs with Insanely Fast Whisper. Learn optimal settings for VRAM to prevent crashes and boost transcription speed.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: performance
- Published: 2026-03-27

---

**Set `--batch-size` based on your GPU VRAM: use 8–12 for ≤6 GB cards, 12–20 for 8–12 GB cards, and 24–48 for ≥16 GB cards to prevent out-of-memory crashes while maintaining optimal transcription throughput.**

Insanely Fast Whisper accelerates OpenAI’s Whisper model by processing audio in parallel batches, but this parallelism consumes GPU memory rapidly. Configuring the correct batch size for your specific GPU is essential to prevent CUDA out-of-memory (OOM) errors while maximizing transcription speed. The tool exposes a `--batch-size` parameter in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) that directly controls how many 30-second audio chunks are processed simultaneously.

## Understanding the `--batch-size` Parameter

In [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py), the CLI defines the `--batch-size` argument at lines 53‑58 with a default value of **24**. This value is passed directly to the Hugging Face `pipeline` call at lines 159‑164, where the Whisper model processes audio chunks of 30 seconds each (`chunk_length_s=30`).

```python
outputs = pipe(
    args.file_name,
    chunk_length_s=30,
    batch_size=args.batch_size,   # Controls GPU memory consumption

    generate_kwargs=generate_kwargs,
    return_timestamps=ts,
)

```

The `batch_size` parameter determines how many audio chunks are kept in GPU memory simultaneously during the forward pass. Larger batches improve GPU utilization and throughput but require proportionally more VRAM.

## Recommended Batch Sizes by GPU Class

GPU memory capacity is the primary constraint when configuring batch size. The following recommendations balance throughput against memory safety for different hardware tiers.

**GPUs with ≤ 6 GB VRAM** (e.g., GTX 1660, RTX 3050):

- Use **`--batch-size 8`** to **`--batch-size 12`**
- This keeps the per-batch tensor footprint low enough to avoid OOM while still leveraging modest parallelism.

**GPUs with 8 GB–12 GB VRAM** (e.g., RTX 3060, RTX 3070):

- Use **`--batch-size 12`** to **`--batch-size 20`**
- You can safely test up to 24 when using smaller Whisper models or shorter audio files.

**GPUs with ≥ 16 GB VRAM** (e.g., RTX 3080, RTX 3090, A100):

- Use **`--batch-size 24`** to **`--batch-size 48`**
- High VRAM headroom allows aggressive batching for maximum throughput.

**Apple Silicon (MPS device):**

- Use **`--batch-size 4`** to **`--batch-size 12`**
- The MPS backend has different memory semantics than CUDA; start low and increment gradually until stable.

## Interactions with Other Performance Flags

Batch size does not operate in isolation. Several other flags in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py) affect GPU memory usage and must be considered when tuning.

**Flash Attention** (`--flash`):

Enabling Flash Attention via `--flash true` switches the model to the memory-efficient `flash_attention_2` implementation. This can allow a larger batch size on the same GPU without triggering OOM errors.

**Model Size**:

Larger Whisper checkpoints (e.g., `openai/whisper-large-v3`) consume significantly more VRAM per batch than `base` or `small` variants. When upgrading to a larger model, reduce `--batch-size` proportionally.

**Device Selection** (`--device-id`):

The batch size logic applies uniformly across CUDA devices (`cuda:0`, `cuda:1`) and Apple Silicon (`mps`). However, multi-GPU setups require setting `--device-id` to specify which GPU handles the batch.

## Practical Configuration Examples

Use these CLI patterns as starting points for your hardware.

**8 GB GPU without Flash Attention:**

```bash
python -m insanely_fast_whisper.cli \
    --file-name podcast.wav \
    --device-id 0 \
    --batch-size 16 \
    --model-name openai/whisper-large-v3 \
    --timestamp word

```

**12 GB GPU with Flash Attention** (allows larger batch):

```bash
python -m insanely_fast_whisper.cli \
    --file-name interview.mp3 \
    --device-id 0 \
    --batch-size 24 \
    --flash true \
    --timestamp chunk

```

**Automated Batch Size Search:**

If you are unsure of your GPU limit, script a descending search to find the maximum stable batch size:

```bash
for b in 24 20 16 12 8; do
  echo "Testing batch size $b"
  python -m insanely_fast_whisper.cli --file-name audio.wav --batch-size $b
  if [ $? -eq 0 ]; then break; fi
done

```

Monitor VRAM usage in real-time with `watch -n 1 nvidia-smi` (CUDA) or Activity Monitor (macOS) while testing.

## Summary

- The `--batch-size` flag in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) defaults to 24 and controls how many 30-second audio chunks are processed in parallel on the GPU.
- **OOM errors** occur when batch size exceeds available GPU memory; reduce the value incrementally until the process runs stable.
- **Memory guidelines**: Use 8–12 for ≤6 GB GPUs, 12–20 for 8–12 GB GPUs, 24–48 for ≥16 GB GPUs, and 4–12 for Apple Silicon.
- **Flash Attention** (`--flash`) reduces memory pressure per chunk, potentially allowing higher batch sizes on the same hardware.
- Related files like [`src/insanely_fast_whisper/utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/result.py) and [`diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/diarization_pipeline.py) handle post-processing and do not affect GPU memory during the batch transcription stage.

## Frequently Asked Questions

### What is the default batch size in insanely-fast-whisper?

The default batch size is **24**, defined in the CLI argument parser at lines 53‑58 of [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py). This default assumes a GPU with at least 12 GB of VRAM; users with less memory should reduce this value immediately to avoid OOM errors.

### How do I know if my batch size is too large?

The process will crash with a `CUDA out of memory` error or a segmentation fault during the forward pass. To prevent this, monitor `nvidia-smi` (for CUDA) or Activity Monitor (for MPS) while processing a test file; if memory utilization approaches 100%, lower `--batch-size` before running the full workload.

### Does enabling Flash Attention allow me to use a larger batch size?

Yes. When you pass `--flash true`, the pipeline uses the `flash_attention_2` implementation, which has a reduced memory footprint per attention operation. This efficiency often allows you to increase `--batch-size` by 4–8 units on the same GPU without triggering OOM errors.

### Can I use the same batch size on Apple Silicon (MPS) as on NVIDIA GPUs?

No. Apple Silicon devices using the `mps` backend typically require smaller batch sizes (4–12) compared to CUDA GPUs with equivalent nominal VRAM due to different memory management semantics in PyTorch’s MPS backend. Start with `--batch-size 4` and increase gradually while monitoring system memory pressure.