# Using Distil-Whisper Models for Faster Inference in Insanely-Fast-Whisper

> Achieve 2-3x faster transcription with Distil-Whisper models in Insanely-Fast-Whisper. Leverage optimized performance with fewer parameters. Get started now.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: performance
- Published: 2026-03-27

---

**You can achieve 2-3x faster transcription speeds in Insanely-Fast-Whisper by simply passing a Distil-Whisper checkpoint (e.g., `distil-whisper/large-v2`) to the `--model-name` argument, which leverages the same Flash Attention and batching optimizations while utilizing roughly half the parameters of standard Whisper models.**

Insanely-Fast-Whisper is a high-performance CLI and Python wrapper around the Hugging Face Transformers automatic speech recognition pipeline. Because the tool treats model identifiers as generic inputs to the `pipeline()` constructor, swapping standard Whisper weights for Distil-Whisper variants requires zero code changes while delivering significant latency reductions. This guide examines the specific source code paths that enable this drop-in acceleration according to the Vaibhavs10/insanely-fast-whisper repository.

## Architecture of the Inference Pipeline

The repository organizes execution into three logical layers contained within `src/insanely_fast_whisper/`.

### CLI Argument Parsing

The entry point in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) captures user inputs, storing the `--model-name` value in `args.model_name`. When you specify a Distil-Whisper identifier such as `distil-whisper/large-v2`, the CLI validates the string but does not differentiate between distilled and full-size architectures during parsing.

### Pipeline Construction (Lines 30-36)

The core instantiation occurs in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py) where the script builds the Hugging Face `pipeline` object. The constructor receives the model identifier directly, alongside dtype and attention implementations:

```python
pipe = pipeline(
    "automatic-speech-recognition",
    model=args.model_name,                 # Accepts any Whisper or Distil-Whisper checkpoint

    torch_dtype=torch.float16,
    device="mps" if args.device_id == "mps" else f"cuda:{args.device_id}",
    model_kwargs={"attn_implementation": "flash_attention_2"} if args.flash else {"attn_implementation": "sdpa"},
)

```

This configuration is agnostic to model size, meaning Distil-Whisper checkpoints load automatically when referenced.

### Inference and Post-Processing

Transcription runs inside the `with Progress(...)` block (lines 52-66) which calls the pipeline with `chunk_length_s=30` and `batch_size=args.batch_size` (defaulting to 24). Optional speaker diarisation triggers via [`utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/diarization_pipeline.py) when a Hugging Face token is provided, while [`utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/result.py) formats the final JSON output.

## Why Distil-Whisper Accelerates Performance

Distil-Whisper models are knowledge-distilled variants of OpenAI's original architectures, containing approximately half the parameters of `whisper-large-v3`. Within Insanely-Fast-Whisper, this reduction yields three concrete advantages:

- **Faster Initialization**: Smaller weight files reduce the time spent in the `pipeline()` constructor during model download and loading.
- **Increased Batch Capacity**: Lower VRAM consumption allows you to increase `--batch-size` beyond 24 on consumer GPUs without out-of-memory errors.
- **Optimization Compatibility**: Distil-Whisper fully supports the Flash Attention 2 and SDPA implementations injected via `model_kwargs`, ensuring you retain all hardware acceleration benefits.

## Running Distil-Whisper Models

Switching to a distilled checkpoint requires only changing the model identifier argument.

### Command-Line Usage

Execute the fastest possible configuration using the following pattern:

```bash
insanely-fast-whisper \
    --model-name distil-whisper/large-v2 \
    --file-name /path/to/audio.wav \
    --batch-size 24 \
    --flash True \
    --timestamp word \
    --transcript-path distil_output.json

```

The `--flash True` flag activates Flash Attention 2, which is particularly effective with the FP16-scaled Distil-Whisper weights.

### Python API Usage

For programmatic access, replicate the CLI's pipeline construction exactly:

```python
from transformers import pipeline
import torch
import json

# Build pipeline identical to CLI behavior

pipe = pipeline(
    "automatic-speech-recognition",
    model="distil-whisper/large-v2",
    torch_dtype=torch.float16,
    device="cuda:0",                     # Use "mps" for Apple Silicon

    model_kwargs={"attn_implementation": "flash_attention_2"},
)

# Run inference

outputs = pipe(
    "audio.wav",
    chunk_length_s=30,
    batch_size=24,
    return_timestamps="word",
)

# Export results

result = {"segments": outputs, "model": "distil-whisper/large-v2"}
with open("output.json", "w", encoding="utf8") as f:
    json.dump(result, f, ensure_ascii=False, indent=2)

```

To add speaker diarization, wrap this pipeline with the helper functions in [`utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/diarization_pipeline.py) and [`utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/result.py).

## Key Source Files

Understanding these modules helps debug or extend Distil-Whisper deployments:

- **[`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py)**: Parses arguments and orchestrates the pipeline instantiation.
- **[`src/insanely_fast_whisper/utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/result.py)**: Formats raw transcription outputs into the final JSON structure.
- **[`src/insanely_fast_whisper/utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarization_pipeline.py)**: Handles speaker segmentation using PyAnnote models when `--hf-token` is supplied.
- **[`pyproject.toml`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/pyproject.toml)**: Declares dependencies including Transformers, Optimum, and Flash Attention libraries required for acceleration.

## Summary

- Distil-Whisper models function as drop-in replacements in Insanely-Fast-Whisper by changing only the `--model-name` argument to a distilled identifier.
- The pipeline construction logic in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py) automatically applies Flash Attention and FP16 optimizations to these smaller checkpoints.
- Reduced parameter counts enable higher batch sizes and faster initialization without sacrificing word-level timestamps or speaker diarization features.

## Frequently Asked Questions

### Does Distil-Whisper support word-level timestamps in Insanely-Fast-Whisper?

Yes. Distil-Whisper models generate word-level timestamps when you pass `--timestamp word` via CLI or `return_timestamps="word"` in Python. The [`utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/result.py) module processes these granular timestamps identically to standard Whisper outputs, placing them in the exported JSON.

### Can I use Flash Attention 2 with Distil-Whisper models?

Absolutely. The CLI's `--flash` flag injects `attn_implementation="flash_attention_2"` into the `model_kwargs` dictionary before pipeline construction. Because Distil-Whisper architectures are compatible with Flash Attention 2, this significantly accelerates computation on NVIDIA Ampere GPUs and newer.

### How does batch size affect Distil-Whisper performance compared to standard Whisper?

Distil-Whisper's smaller memory footprint allows you to increase `--batch-size` well beyond the default 24 while staying within VRAM limits. On a 16GB GPU, you may run batches of 32 or 48 with Distil-Whisper where standard Whisper would exhaust memory, further reducing total transcription time.

### Is speaker diarization compatible with Distil-Whisper checkpoints?

Yes. When you provide a Hugging Face token via `--hf-token`, the [`utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/diarization_pipeline.py) module performs speaker segmentation using PyAnnote models independently of the ASR backend. Distil-Whisper transcriptions integrate seamlessly with these diarization results, producing speaker-attributed segments in the final JSON.