How to Select Different Whisper Model Variants (large-v3, distil-large-v2) with insanely-fast-whisper

Use the --model-name CLI flag or model_name parameter in the Python API to load any Hugging Face Transformers-compatible Whisper checkpoint, including openai/whisper-large-v3 (default) or distil-whisper/large-v2.

insanely-fast-whisper is a high-performance wrapper around the Hugging Face Transformers automatic speech recognition (ASR) pipeline. The repository Vaibhavs10/insanely-fast-whisper exposes a streamlined interface for selecting different Whisper model variants without modifying underlying inference logic. Whether you require maximum accuracy with OpenAI's latest checkpoint or faster inference with distilled models, the selection happens through a single configuration parameter.

How Model Selection Works

CLI Argument Definition

In src/insanely_fast_whisper/cli.py (lines 32-36), the --model-name argument is defined with a default value pointing to OpenAI's large-v3 checkpoint:

parser.add_argument(
    "--model-name",
    required=False,
    default="openai/whisper-large-v3",
    type=str,
    help="Name of the pretrained model/ checkpoint to perform ASR. (default: openai/whisper-large-v3)",
)

This argument accepts any valid Hugging Face model identifier, allowing you to swap between OpenAI's official releases and community-distilled variants.

Pipeline Instantiation

The CLI passes args.model_name directly to the Transformers pipeline constructor. As implemented in src/insanely_fast_whisper/cli.py:

pipe = pipeline(
    "automatic-speech-recognition",
    model=args.model_name,
    torch_dtype=torch.float16,
    device="mps" if args.device_id == "mps" else f"cuda:{args.device_id}",
    model_kwargs={"attn_implementation": "flash_attention_2"} if args.flash else {"attn_implementation": "sdpa"},
)

Available Whisper Model Variants

The wrapper supports any Whisper-compatible checkpoint resolvable by the Transformers library. Common configurations include:

  • openai/whisper-large-v3: The default model offering highest accuracy. Typical FP16 batch size of 24 processes audio at approximately 2 seconds per minute on an A100 with Flash Attention 2 enabled.
  • distil-whisper/large-v2: A distilled variant providing ~2× speedup with slightly lower word error rate (WER). Maintains the same batch size capabilities but reduces computational overhead.
  • openai/whisper-medium and openai/whisper-small: Smaller checkpoints for resource-constrained environments, trading accuracy for lower GPU memory requirements.

Because the model is cached locally after the first download, switching variants only affects the initial load time and subsequent inference graph.

Selecting Models via Command Line

Using distil-whisper/large-v2 with Flash Attention

For maximum speed with minimal accuracy loss, specify the distilled checkpoint and enable Flash Attention 2:

insanely-fast-whisper \
  --file-name my_talk.wav \
  --model-name distil-whisper/large-v2 \
  --flash True \
  --batch-size 24 \
  --transcript-path result.json

Running the Default large-v3 Model

To explicitly use the default high-accuracy model without Flash Attention:

insanely-fast-whisper \
  --file-name lecture.mp3 \
  --model-name openai/whisper-large-v3 \
  --flash False

Python API Model Selection

When using the Python API directly, pass the model identifier to the model parameter in the pipeline constructor:

from transformers import pipeline, is_flash_attn_2_available
import torch

def transcribe(audio_path: str, model: str = "openai/whisper-large-v3", use_flash: bool = True):
    pipe = pipeline(
        "automatic-speech-recognition",
        model=model,
        torch_dtype=torch.float16,
        device="cuda:0",
        model_kwargs={"attn_implementation": "flash_attention_2"} if use_flash and is_flash_attn_2_available()
        else {"attn_implementation": "sdpa"},
    )
    return pipe(
        audio_path,
        chunk_length_s=30,
        batch_size=24,
        return_timestamps=True,
    )

result = transcribe("interview.wav", model="distil-whisper/large-v2")
print(result["text"])

This implementation mirrors the CLI behavior exactly, exposing the same performance characteristics and model flexibility.

Flash Attention Compatibility

When you enable --flash True, the tool injects attn_implementation="flash_attention_2" into the model kwargs. This works for any model variant that supports PyTorch 2.0 Flash Attention kernels, including both standard and distilled Whisper checkpoints. The attention implementation is independent of model selection, meaning you can combine Flash Attention 2 with large-v3, distil-large-v2, or custom fine-tuned models.

Summary

  • Model selection in insanely-fast-whisper is controlled entirely by the --model-name CLI flag or the model parameter in the Python API.
  • The default checkpoint is openai/whisper-large-v3, but you can substitute any Hugging Face Transformers-compatible Whisper model such as distil-whisper/large-v2.
  • Model variants are passed directly to the underlying Transformers pipeline in src/insanely_fast_whisper/cli.py, ensuring full compatibility with the Hugging Face ecosystem.
  • Flash Attention 2 can be enabled independently of model choice using the --flash flag or attn_implementation model kwargs.
  • Models are cached locally after first download, making subsequent switches between variants efficient.

Frequently Asked Questions

Can I use custom fine-tuned Whisper models with insanely-fast-whisper?

Yes. Any Whisper-compatible checkpoint hosted on Hugging Face or stored locally can be specified using the --model-name argument. Simply pass the repository ID (e.g., your-username/whisper-finetuned-en) or local path to the model weights.

Does switching between model variants require reinstalling the package?

No. The model weights are downloaded and cached by the Transformers library in ~/.cache/huggingface/hub upon first use. Switching models only changes which checkpoint gets loaded into the pipeline; the insanely-fast-whisper code remains unchanged.

Which Whisper model variant offers the best speed-accuracy trade-off?

According to the repository benchmarks, distil-whisper/large-v2 provides approximately 2× faster inference than large-v3 with only a minor degradation in word error rate. For production scenarios where latency is critical, the distilled variant is recommended, while large-v3 remains optimal for maximum transcription accuracy.

Is Flash Attention 2 compatible with all Whisper model variants?

Flash Attention 2 works with any Whisper model that supports the PyTorch SDPA (Scaled Dot Product Attention) interface, which includes all standard OpenAI and Distil-Whisper checkpoints. However, Flash Attention 2 requires specific hardware (NVIDIA Ampere GPUs or newer) and proper environment setup with flash-attn installed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →