# How to Select Different Whisper Model Variants (large-v3, distil-large-v2) with insanely-fast-whisper

> Easily select Whisper model variants like large-v3 or distil-large-v2 with insanely-fast-whisper using the model-name flag or Python API for flexible transcription options.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: how-to-guide
- Published: 2026-03-27

---

**Use the `--model-name` CLI flag or `model_name` parameter in the Python API to load any Hugging Face Transformers-compatible Whisper checkpoint, including `openai/whisper-large-v3` (default) or `distil-whisper/large-v2`.**

insanely-fast-whisper is a high-performance wrapper around the Hugging Face Transformers automatic speech recognition (ASR) pipeline. The repository `Vaibhavs10/insanely-fast-whisper` exposes a streamlined interface for selecting different Whisper model variants without modifying underlying inference logic. Whether you require maximum accuracy with OpenAI's latest checkpoint or faster inference with distilled models, the selection happens through a single configuration parameter.

## How Model Selection Works

### CLI Argument Definition

In [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) (lines 32-36), the `--model-name` argument is defined with a default value pointing to OpenAI's large-v3 checkpoint:

```python
parser.add_argument(
    "--model-name",
    required=False,
    default="openai/whisper-large-v3",
    type=str,
    help="Name of the pretrained model/ checkpoint to perform ASR. (default: openai/whisper-large-v3)",
)

```

This argument accepts any valid Hugging Face model identifier, allowing you to swap between OpenAI's official releases and community-distilled variants.

### Pipeline Instantiation

The CLI passes `args.model_name` directly to the Transformers `pipeline` constructor. As implemented in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py):

```python
pipe = pipeline(
    "automatic-speech-recognition",
    model=args.model_name,
    torch_dtype=torch.float16,
    device="mps" if args.device_id == "mps" else f"cuda:{args.device_id}",
    model_kwargs={"attn_implementation": "flash_attention_2"} if args.flash else {"attn_implementation": "sdpa"},
)

```

## Available Whisper Model Variants

The wrapper supports any Whisper-compatible checkpoint resolvable by the Transformers library. Common configurations include:

- **openai/whisper-large-v3**: The default model offering highest accuracy. Typical FP16 batch size of 24 processes audio at approximately 2 seconds per minute on an A100 with Flash Attention 2 enabled.
- **distil-whisper/large-v2**: A distilled variant providing ~2× speedup with slightly lower word error rate (WER). Maintains the same batch size capabilities but reduces computational overhead.
- **openai/whisper-medium** and **openai/whisper-small**: Smaller checkpoints for resource-constrained environments, trading accuracy for lower GPU memory requirements.

Because the model is cached locally after the first download, switching variants only affects the initial load time and subsequent inference graph.

## Selecting Models via Command Line

### Using distil-whisper/large-v2 with Flash Attention

For maximum speed with minimal accuracy loss, specify the distilled checkpoint and enable Flash Attention 2:

```bash
insanely-fast-whisper \
  --file-name my_talk.wav \
  --model-name distil-whisper/large-v2 \
  --flash True \
  --batch-size 24 \
  --transcript-path result.json

```

### Running the Default large-v3 Model

To explicitly use the default high-accuracy model without Flash Attention:

```bash
insanely-fast-whisper \
  --file-name lecture.mp3 \
  --model-name openai/whisper-large-v3 \
  --flash False

```

## Python API Model Selection

When using the Python API directly, pass the model identifier to the `model` parameter in the pipeline constructor:

```python
from transformers import pipeline, is_flash_attn_2_available
import torch

def transcribe(audio_path: str, model: str = "openai/whisper-large-v3", use_flash: bool = True):
    pipe = pipeline(
        "automatic-speech-recognition",
        model=model,
        torch_dtype=torch.float16,
        device="cuda:0",
        model_kwargs={"attn_implementation": "flash_attention_2"} if use_flash and is_flash_attn_2_available()
        else {"attn_implementation": "sdpa"},
    )
    return pipe(
        audio_path,
        chunk_length_s=30,
        batch_size=24,
        return_timestamps=True,
    )

result = transcribe("interview.wav", model="distil-whisper/large-v2")
print(result["text"])

```

This implementation mirrors the CLI behavior exactly, exposing the same performance characteristics and model flexibility.

## Flash Attention Compatibility

When you enable `--flash True`, the tool injects `attn_implementation="flash_attention_2"` into the model kwargs. This works for any model variant that supports PyTorch 2.0 Flash Attention kernels, including both standard and distilled Whisper checkpoints. The attention implementation is independent of model selection, meaning you can combine Flash Attention 2 with `large-v3`, `distil-large-v2`, or custom fine-tuned models.

## Summary

- **Model selection** in insanely-fast-whisper is controlled entirely by the `--model-name` CLI flag or the `model` parameter in the Python API.
- The default checkpoint is `openai/whisper-large-v3`, but you can substitute any Hugging Face Transformers-compatible Whisper model such as `distil-whisper/large-v2`.
- Model variants are passed directly to the underlying Transformers pipeline in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py), ensuring full compatibility with the Hugging Face ecosystem.
- Flash Attention 2 can be enabled independently of model choice using the `--flash` flag or `attn_implementation` model kwargs.
- Models are cached locally after first download, making subsequent switches between variants efficient.

## Frequently Asked Questions

### Can I use custom fine-tuned Whisper models with insanely-fast-whisper?

Yes. Any Whisper-compatible checkpoint hosted on Hugging Face or stored locally can be specified using the `--model-name` argument. Simply pass the repository ID (e.g., `your-username/whisper-finetuned-en`) or local path to the model weights.

### Does switching between model variants require reinstalling the package?

No. The model weights are downloaded and cached by the Transformers library in `~/.cache/huggingface/hub` upon first use. Switching models only changes which checkpoint gets loaded into the pipeline; the insanely-fast-whisper code remains unchanged.

### Which Whisper model variant offers the best speed-accuracy trade-off?

According to the repository benchmarks, `distil-whisper/large-v2` provides approximately 2× faster inference than `large-v3` with only a minor degradation in word error rate. For production scenarios where latency is critical, the distilled variant is recommended, while `large-v3` remains optimal for maximum transcription accuracy.

### Is Flash Attention 2 compatible with all Whisper model variants?

Flash Attention 2 works with any Whisper model that supports the PyTorch SDPA (Scaled Dot Product Attention) interface, which includes all standard OpenAI and Distil-Whisper checkpoints. However, Flash Attention 2 requires specific hardware (NVIDIA Ampere GPUs or newer) and proper environment setup with `flash-attn` installed.