Using Distil-Whisper Models for Faster Inference in Insanely-Fast-Whisper
You can achieve 2-3x faster transcription speeds in Insanely-Fast-Whisper by simply passing a Distil-Whisper checkpoint (e.g., distil-whisper/large-v2) to the --model-name argument, which leverages the same Flash Attention and batching optimizations while utilizing roughly half the parameters of standard Whisper models.
Insanely-Fast-Whisper is a high-performance CLI and Python wrapper around the Hugging Face Transformers automatic speech recognition pipeline. Because the tool treats model identifiers as generic inputs to the pipeline() constructor, swapping standard Whisper weights for Distil-Whisper variants requires zero code changes while delivering significant latency reductions. This guide examines the specific source code paths that enable this drop-in acceleration according to the Vaibhavs10/insanely-fast-whisper repository.
Architecture of the Inference Pipeline
The repository organizes execution into three logical layers contained within src/insanely_fast_whisper/.
CLI Argument Parsing
The entry point in src/insanely_fast_whisper/cli.py captures user inputs, storing the --model-name value in args.model_name. When you specify a Distil-Whisper identifier such as distil-whisper/large-v2, the CLI validates the string but does not differentiate between distilled and full-size architectures during parsing.
Pipeline Construction (Lines 30-36)
The core instantiation occurs in cli.py where the script builds the Hugging Face pipeline object. The constructor receives the model identifier directly, alongside dtype and attention implementations:
pipe = pipeline(
"automatic-speech-recognition",
model=args.model_name, # Accepts any Whisper or Distil-Whisper checkpoint
torch_dtype=torch.float16,
device="mps" if args.device_id == "mps" else f"cuda:{args.device_id}",
model_kwargs={"attn_implementation": "flash_attention_2"} if args.flash else {"attn_implementation": "sdpa"},
)
This configuration is agnostic to model size, meaning Distil-Whisper checkpoints load automatically when referenced.
Inference and Post-Processing
Transcription runs inside the with Progress(...) block (lines 52-66) which calls the pipeline with chunk_length_s=30 and batch_size=args.batch_size (defaulting to 24). Optional speaker diarisation triggers via utils/diarization_pipeline.py when a Hugging Face token is provided, while utils/result.py formats the final JSON output.
Why Distil-Whisper Accelerates Performance
Distil-Whisper models are knowledge-distilled variants of OpenAI's original architectures, containing approximately half the parameters of whisper-large-v3. Within Insanely-Fast-Whisper, this reduction yields three concrete advantages:
- Faster Initialization: Smaller weight files reduce the time spent in the
pipeline()constructor during model download and loading. - Increased Batch Capacity: Lower VRAM consumption allows you to increase
--batch-sizebeyond 24 on consumer GPUs without out-of-memory errors. - Optimization Compatibility: Distil-Whisper fully supports the Flash Attention 2 and SDPA implementations injected via
model_kwargs, ensuring you retain all hardware acceleration benefits.
Running Distil-Whisper Models
Switching to a distilled checkpoint requires only changing the model identifier argument.
Command-Line Usage
Execute the fastest possible configuration using the following pattern:
insanely-fast-whisper \
--model-name distil-whisper/large-v2 \
--file-name /path/to/audio.wav \
--batch-size 24 \
--flash True \
--timestamp word \
--transcript-path distil_output.json
The --flash True flag activates Flash Attention 2, which is particularly effective with the FP16-scaled Distil-Whisper weights.
Python API Usage
For programmatic access, replicate the CLI's pipeline construction exactly:
from transformers import pipeline
import torch
import json
# Build pipeline identical to CLI behavior
pipe = pipeline(
"automatic-speech-recognition",
model="distil-whisper/large-v2",
torch_dtype=torch.float16,
device="cuda:0", # Use "mps" for Apple Silicon
model_kwargs={"attn_implementation": "flash_attention_2"},
)
# Run inference
outputs = pipe(
"audio.wav",
chunk_length_s=30,
batch_size=24,
return_timestamps="word",
)
# Export results
result = {"segments": outputs, "model": "distil-whisper/large-v2"}
with open("output.json", "w", encoding="utf8") as f:
json.dump(result, f, ensure_ascii=False, indent=2)
To add speaker diarization, wrap this pipeline with the helper functions in utils/diarization_pipeline.py and utils/result.py.
Key Source Files
Understanding these modules helps debug or extend Distil-Whisper deployments:
src/insanely_fast_whisper/cli.py: Parses arguments and orchestrates the pipeline instantiation.src/insanely_fast_whisper/utils/result.py: Formats raw transcription outputs into the final JSON structure.src/insanely_fast_whisper/utils/diarization_pipeline.py: Handles speaker segmentation using PyAnnote models when--hf-tokenis supplied.pyproject.toml: Declares dependencies including Transformers, Optimum, and Flash Attention libraries required for acceleration.
Summary
- Distil-Whisper models function as drop-in replacements in Insanely-Fast-Whisper by changing only the
--model-nameargument to a distilled identifier. - The pipeline construction logic in
cli.pyautomatically applies Flash Attention and FP16 optimizations to these smaller checkpoints. - Reduced parameter counts enable higher batch sizes and faster initialization without sacrificing word-level timestamps or speaker diarization features.
Frequently Asked Questions
Does Distil-Whisper support word-level timestamps in Insanely-Fast-Whisper?
Yes. Distil-Whisper models generate word-level timestamps when you pass --timestamp word via CLI or return_timestamps="word" in Python. The utils/result.py module processes these granular timestamps identically to standard Whisper outputs, placing them in the exported JSON.
Can I use Flash Attention 2 with Distil-Whisper models?
Absolutely. The CLI's --flash flag injects attn_implementation="flash_attention_2" into the model_kwargs dictionary before pipeline construction. Because Distil-Whisper architectures are compatible with Flash Attention 2, this significantly accelerates computation on NVIDIA Ampere GPUs and newer.
How does batch size affect Distil-Whisper performance compared to standard Whisper?
Distil-Whisper's smaller memory footprint allows you to increase --batch-size well beyond the default 24 while staying within VRAM limits. On a 16GB GPU, you may run batches of 32 or 48 with Distil-Whisper where standard Whisper would exhaust memory, further reducing total transcription time.
Is speaker diarization compatible with Distil-Whisper checkpoints?
Yes. When you provide a Hugging Face token via --hf-token, the utils/diarization_pipeline.py module performs speaker segmentation using PyAnnote models independently of the ASR backend. Distil-Whisper transcriptions integrate seamlessly with these diarization results, producing speaker-attributed segments in the final JSON.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →