insanely-fast-whisper

19 articles 11.6k View on GitHub ↗
19 articles
How to Use `--min-speakers` and `--max-speakers` for Flexible Diarization in Insanely-Fast-Whisper

Control speaker count in insanely-fast-whisper using min-speakers and max-speakers parameters for flexible diarization. Automatically discover optimal speaker numbers within your specified range.

how-to-guide
Mar 27, 2026
Internal Configuration of the Transformers Pipeline in Insanely‑Fast‑Whisper

Explore the internal configuration of the insanely-fast-whisper Transformers pipeline. Optimize ASR with float16, FlashAttention 2, and device-specific settings for faster performance.

internals
Mar 27, 2026
How to Provide a HuggingFace Token for pyannote.audio Access in insanely-fast-whisper

Learn how to provide your HuggingFace token to insanely-fast-whisper using the --hf-token argument. Securely access pyannote.audio for speaker diarization.

how-to-guide
Mar 27, 2026
How to Get Word-Level Timestamps in Insanely-Fast-Whisper for Precise Timing

Unlock precise timing with word-level timestamps in insanely-fast-whisper. Simply use the --timestamp word CLI flag for accurate audio analysis. Get started now!

how-to-guide
Mar 27, 2026
Tuning Batch Size for Specific GPU Memory Constraints in Insanely-Fast-Whisper

Optimize Whisper GPU memory with precise batch size tuning. Calculate your max parallel processing limit using torch.cuda.max_memory_allocated() and 85% of your VRAM for peak performance.

how-to-guide
Mar 27, 2026
BetterTransformer Integration in insanely-fast-whisper: A 5× Speedup Deep Dive

Discover how BetterTransformer integration in insanely-fast-whisper achieves 5x faster inference by optimizing compute graphs and fusing attention kernels on CUDA devices without Flash Attention 2.

deep-dive
Mar 27, 2026
How to Fix "Torch not compiled with CUDA" Error on Windows with insanely-fast-whisper

Fix the Torch not compiled with CUDA error for insanely-fast-whisper on Windows. Resolve PyTorch GPU memory issues and get your Whisper model running on the GPU quickly.

how-to-guide
Mar 27, 2026
How to Select Different Whisper Model Variants (large-v3, distil-large-v2) with insanely-fast-whisper

Easily select Whisper model variants like large-v3 or distil-large-v2 with insanely-fast-whisper using the model-name flag or Python API for flexible transcription options.

how-to-guide
Mar 27, 2026
How to Specify a Custom Output File Path for Transcriptions in insanely-fast-whisper

Learn to specify a custom output file path for transcriptions in insanely-fast-whisper using the --transcript-path argument. Control where your transcription results are saved with ease.

how-to-guide
Mar 27, 2026
How to Install flash-attn in a pipx Environment for insanely-fast-whisper

Install flash-attn in pipx for insanely-fast-whisper hardware acceleration. Boost inference speed with this quick setup guide.

how-to-guide
Mar 27, 2026
Using the Translate Task Instead of Transcribe in Insanely-Fast-Whisper: A Complete Guide

Learn how to use the translate task in insanely-fast-whisper with the --task translate flag. Get translated text instead of transcripts while keeping the same JSON output. Optimize your audio translation workflow.

how-to-guide
Mar 27, 2026
Resolving CUDA Out-of-Memory Errors with Insanely-Fast-Whisper: Optimization Guide

Fix CUDA out of memory errors in insanely-fast-whisper by lowering batch size, enabling Flash Attention 2, and optimizing your device. Transcribe audio without crashes.

optimization-guide
Mar 27, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →