# insanely-fast-whisper | vb | Knowledge Base | Instagit

GitHub Stars: 11.6k

Repository: https://github.com/Vaibhavs10/insanely-fast-whisper

---

## Articles

### [How to Use `--min-speakers` and `--max-speakers` for Flexible Diarization in Insanely-Fast-Whisper](/Vaibhavs10/insanely-fast-whisper/how-to-use-min-speakers-and-max-speakers-for-flexible-diarization)

Control speaker count in insanely-fast-whisper using min-speakers and max-speakers parameters for flexible diarization. Automatically discover optimal speaker numbers within your specified range.

- Tags: how-to-guide
- Published: 2026-03-27

### [Internal Configuration of the Transformers Pipeline in Insanely‑Fast‑Whisper](/Vaibhavs10/insanely-fast-whisper/how-transformers-pipeline-is-configured-internally)

Explore the internal configuration of the insanely-fast-whisper Transformers pipeline. Optimize ASR with float16, FlashAttention 2, and device-specific settings for faster performance.

- Tags: internals
- Published: 2026-03-27

### [How to Provide a HuggingFace Token for pyannote.audio Access in insanely-fast-whisper](/Vaibhavs10/insanely-fast-whisper/how-to-provide-huggingface-token-for-pyannote-audio-access)

Learn how to provide your HuggingFace token to insanely-fast-whisper using the --hf-token argument. Securely access pyannote.audio for speaker diarization.

- Tags: how-to-guide
- Published: 2026-03-27

### [How to Get Word-Level Timestamps in Insanely-Fast-Whisper for Precise Timing](/Vaibhavs10/insanely-fast-whisper/how-to-get-word-level-timestamps-for-precise-timing)

Unlock precise timing with word-level timestamps in insanely-fast-whisper. Simply use the --timestamp word CLI flag for accurate audio analysis. Get started now!

- Tags: how-to-guide
- Published: 2026-03-27

### [Tuning Batch Size for Specific GPU Memory Constraints in Insanely-Fast-Whisper](/Vaibhavs10/insanely-fast-whisper/how-to-tune-batch-size-for-specific-gpu-memory-constraints)

Optimize Whisper GPU memory with precise batch size tuning. Calculate your max parallel processing limit using torch.cuda.max_memory_allocated() and 85% of your VRAM for peak performance.

- Tags: how-to-guide
- Published: 2026-03-27

### [BetterTransformer Integration in insanely-fast-whisper: A 5× Speedup Deep Dive](/Vaibhavs10/insanely-fast-whisper/how-bettertransformer-integration-improves-performance)

Discover how BetterTransformer integration in insanely-fast-whisper achieves 5x faster inference by optimizing compute graphs and fusing attention kernels on CUDA devices without Flash Attention 2.

- Tags: deep-dive
- Published: 2026-03-27

### [How to Fix "Torch not compiled with CUDA" Error on Windows with insanely-fast-whisper](/Vaibhavs10/insanely-fast-whisper/how-to-fix-torch-not-compiled-with-cuda-error-on-windows)

Fix the Torch not compiled with CUDA error for insanely-fast-whisper on Windows. Resolve PyTorch GPU memory issues and get your Whisper model running on the GPU quickly.

- Tags: how-to-guide
- Published: 2026-03-27

### [How to Select Different Whisper Model Variants (large-v3, distil-large-v2) with insanely-fast-whisper](/Vaibhavs10/insanely-fast-whisper/how-to-select-different-whisper-model-variants)

Easily select Whisper model variants like large-v3 or distil-large-v2 with insanely-fast-whisper using the model-name flag or Python API for flexible transcription options.

- Tags: how-to-guide
- Published: 2026-03-27

### [How to Specify a Custom Output File Path for Transcriptions in insanely-fast-whisper](/Vaibhavs10/insanely-fast-whisper/how-to-specify-custom-output-file-path-for-transcription)

Learn to specify a custom output file path for transcriptions in insanely-fast-whisper using the --transcript-path argument. Control where your transcription results are saved with ease.

- Tags: how-to-guide
- Published: 2026-03-27

### [How to Install flash-attn in a pipx Environment for insanely-fast-whisper](/Vaibhavs10/insanely-fast-whisper/how-to-correctly-install-flash-attn-in-pipx-environment)

Install flash-attn in pipx for insanely-fast-whisper hardware acceleration. Boost inference speed with this quick setup guide.

- Tags: how-to-guide
- Published: 2026-03-27

### [Using the Translate Task Instead of Transcribe in Insanely-Fast-Whisper: A Complete Guide](/Vaibhavs10/insanely-fast-whisper/how-to-use-translate-task-instead-of-transcribe)

Learn how to use the translate task in insanely-fast-whisper with the --task translate flag. Get translated text instead of transcripts while keeping the same JSON output. Optimize your audio translation workflow.

- Tags: how-to-guide
- Published: 2026-03-27

### [Resolving CUDA Out-of-Memory Errors with Insanely-Fast-Whisper: Optimization Guide](/Vaibhavs10/insanely-fast-whisper/how-to-resolve-cuda-out-of-memory-errors)

Fix CUDA out of memory errors in insanely-fast-whisper by lowering batch size, enabling Flash Attention 2, and optimizing your device. Transcribe audio without crashes.

- Tags: optimization-guide
- Published: 2026-03-27

### [Using Distil-Whisper Models for Faster Inference in Insanely-Fast-Whisper](/Vaibhavs10/insanely-fast-whisper/how-to-use-distil-whisper-models-for-faster-inference)

Achieve 2-3x faster transcription with Distil-Whisper models in Insanely-Fast-Whisper. Leverage optimized performance with fewer parameters. Get started now.

- Tags: performance
- Published: 2026-03-27

### [How to Specify the Exact Number of Speakers in Diarization with Insanely-Fast-Whisper](/Vaibhavs10/insanely-fast-whisper/how-to-specify-exact-number-of-speakers-in-diarization)

Control speaker count in audio diarization with insanely-fast-whisper. Easily specify the exact number of speakers using the --num-speakers N flag for precise results.

- Tags: how-to-guide
- Published: 2026-03-27

### [Using the Python API Instead of CLI for Custom Pipelines in Insanely‑Fast‑Whisper](/Vaibhavs10/insanely-fast-whisper/how-to-use-python-api-instead-of-cli-for-custom-pipelines)

Leverage the insanely-fast-whisper Python API for custom ASR pipelines. Gain full control over model selection, batching, Flash-Attention 2, and diarization. Avoid CLI subprocesses for efficient audio processing.

- Tags: how-to-guide
- Published: 2026-03-27

### [Choosing Between Chunked and Word-Level Timestamps in insanely-fast-whisper](/Vaibhavs10/insanely-fast-whisper/how-to-choose-between-chunked-and-word-level-timestamps)

Learn to choose between chunked and word-level timestamps in insanely-fast-whisper. Get fast processing or precise timing for subtitles and editing workflows. Optimize your audio transcription.

- Tags: comparison
- Published: 2026-03-27

### [How to Run Insanely-Fast-Whisper on Apple Silicon (MPS) for GPU-Accelerated Transcription](/Vaibhavs10/insanely-fast-whisper/how-to-run-insanely-fast-whisper-on-apple-silicon-mps-devices)

Accelerate Whisper transcription on Apple Silicon Macs. Learn how to run insanely-fast-whisper using MPS for native GPU acceleration with the mps device ID or parameter.

- Tags: how-to-guide
- Published: 2026-03-27

### [How to Configure Batch Size to Avoid OOM Errors on Different GPUs with Insanely Fast Whisper](/Vaibhavs10/insanely-fast-whisper/how-to-configure-batch-size-to-avoid-oom-errors-on-different-gpus)

Configure batch size to avoid OOM errors on different GPUs with Insanely Fast Whisper. Learn optimal settings for VRAM to prevent crashes and boost transcription speed.

- Tags: performance
- Published: 2026-03-27

### [How to Enable Flash Attention 2 for Maximum Transcription Speed with insanely-fast-whisper](/Vaibhavs10/insanely-fast-whisper/how-to-enable-flash-attention-2-for-maximum-transcription-speed)

Boost transcription speed with insanely-fast-whisper by enabling Flash Attention 2. Learn how to install and configure it for maximum performance.

- Tags: how-to-guide
- Published: 2026-03-27

