Using the Python API Instead of CLI for Custom Pipelines in Insanely‑Fast‑Whisper
The insanely-fast-whisper CLI is a thin wrapper around Hugging Face Transformers ASR pipelines that you can import directly from the source modules and orchestrate in Python, giving you full control over model selection, batch processing, Flash-Attention 2, and speaker diarization without subprocess calls.
The insanely-fast-whisper repository ships with a command-line interface for rapid Whisper transcription, but the underlying implementation is fully exposed as a Python API. By importing the pipeline builders and utility functions from src/insanely_fast_whisper/, you can construct custom transcription workflows that bypass the CLI entirely while retaining access to all performance optimizations including Flash-Attention and PyAnnote speaker segmentation.
Core Architecture and Source Modules
The project separates concerns across four key modules that the CLI merely wires together. Understanding these components allows you to reconstruct the pipeline programmatically:
-
CLI Entry Point – Located in
src/insanely_fast_whisper/cli.py, this file parses arguments and constructs the basepipeline("automatic-speech-recognition", …)object with your selected model, dtype, device, and attention implementation. -
Result Builder – The
build_result(transcript, outputs)function insrc/insanely_fast_whisper/utils/result.pynormalizes the raw Whisper output into a JSON-compatible dictionary, copying thechunksandtextfields and appending optional diarization data. -
Diarization Pipeline –
src/insanely_fast_whisper/utils/diarization_pipeline.pyprovides the high-leveldiarize()wrapper that loads apyannote.audiopipeline, moves it to the correct device, and aligns speaker segments with ASR timestamps. -
Low-Level Diarization Helpers –
src/insanely_fast_whisper/utils/diarize.pycontains the actual preprocessing and alignment logic, includingpreprocess_inputs,diarize_audio, andpost_process_segments_and_transcripts.
Because these functions are public, you can swap models, adjust batch sizes on the fly, inject custom preprocessing, or replace the diarization backend entirely.
Basic ASR Pipeline Implementation
To transcribe audio without invoking the CLI, instantiate the Transformers pipeline directly using the same defaults found in src/insanely_fast_whisper/cli.py. The following snippet demonstrates word-level timestamps with Flash-Attention 2 enabled:
import torch
from transformers import pipeline
from insanely_fast_whisper.utils.result import build_result
# Build the pipeline – mirrors the CLI defaults
pipe = pipeline(
"automatic-speech-recognition",
model="openai/whisper-large-v3",
torch_dtype=torch.float16,
device="cuda:0", # or "mps" on macOS
model_kwargs={"attn_implementation": "flash_attention_2"},
)
# Run inference – chunk size = 30s, batch size = 24, word-level timestamps
outputs = pipe(
"my_audio_file.wav",
chunk_length_s=30,
batch_size=24,
return_timestamps="word",
)
# Build the JSON result (no diarisation)
result = build_result([], outputs)
# Write to disk (optional)
import json
with open("output.json", "w", encoding="utf-8") as fp:
json.dump(result, fp, ensure_ascii=False, indent=2)
The model_kwargs dictionary enables Flash-Attention 2 via "flash_attention_2", while return_timestamps="word" generates the same granular timestamps produced by the CLI's --timestamp word flag. The build_result utility ensures the output schema matches the CLI's JSON format exactly.
Adding Speaker Diarization Programmatically
To include speaker segmentation without the CLI, import the diarization utilities and provide a configuration object matching the CLI's argument structure. The diarize() function expects an object with attributes like file_name, device_id, and hf_token:
import torch
from transformers import pipeline
from pyannote.audio import Pipeline
from insanely_fast_whisper.utils.diarization_pipeline import diarize
from insanely_fast_whisper.utils.result import build_result
# 1. ASR pipeline (same as before)
pipe = pipeline(
"automatic-speech-recognition",
model="openai/whisper-large-v3",
torch_dtype=torch.float16,
device="cuda:0",
model_kwargs={"attn_implementation": "flash_attention_2"},
)
# 2. Run ASR
outputs = pipe(
"my_audio_file.wav",
chunk_length_s=30,
batch_size=24,
return_timestamps="word",
)
# 3. Prepare configuration object for diarization
class Args:
file_name = "my_audio_file.wav"
device_id = "0"
hf_token = "hf_XXXXXXXXXXXXXXXX" # Your HF token with PyAnnote access
diarization_model = "pyannote/speaker-diarization-3.1"
num_speakers = None
min_speakers = None
max_speakers = None
args = Args()
# 4. Run diarisation – returns speaker-annotated chunks
speakers_transcript = diarize(args, outputs)
# 5. Combine ASR+diarisation into final JSON
result = build_result(speakers_transcript, outputs)
with open("output_with_speakers.json", "w", encoding="utf-8") as fp:
import json
json.dump(result, fp, ensure_ascii=False, indent=2)
The diarize function handles loading the PyAnnote pipeline via Pipeline.from_pretrained, moving it to the specified CUDA or MPS device, and aligning timestamps using the helpers in src/insanely_fast_whisper/utils/diarize.py.
Custom Models and Batch Configuration
You can swap the underlying Whisper model or adjust memory constraints by modifying the pipeline arguments. This example uses distil-whisper/large-v2 with a reduced batch size:
import torch
from transformers import pipeline
from insanely_fast_whisper.utils.result import build_result
pipe = pipeline(
"automatic-speech-recognition",
model="distil-whisper/large-v2",
torch_dtype=torch.float16,
device="cuda:0",
model_kwargs={"attn_implementation": "flash_attention_2"},
)
outputs = pipe(
"meeting.wav",
chunk_length_s=30,
batch_size=8, # Smaller batch to fit GPU memory
return_timestamps=True,
)
result = build_result([], outputs)
with open("distil_output.json", "w", encoding="utf-8") as fp:
import json
json.dump(result, fp, ensure_ascii=False, indent=2)
Changing the model parameter automatically swaps the checkpoint without requiring code changes to the pipeline logic. You can further customize the workflow by importing convert_output.py utilities to generate SRT or VTT files from the resulting JSON.
Summary
- The insanely-fast-whisper CLI is optional – all functionality resides in importable Python modules under
src/insanely_fast_whisper/. - Use
pipeline()from Transformers withmodel_kwargs={"attn_implementation": "flash_attention_2"}to enable Flash-Attention 2. - Normalize outputs with
build_result()fromutils/result.pyto maintain CLI-compatible JSON schemas. - Add diarization by calling
diarize()fromutils/diarization_pipeline.pywith a configuration object containing yourhf_tokenand device settings. - Customize freely – swap models (e.g.,
distil-whisper), adjustbatch_size, or loop over file batches without subprocess overhead.
Frequently Asked Questions
How do I enable Flash-Attention 2 when using the Python API?
Pass model_kwargs={"attn_implementation": "flash_attention_2"} when constructing the Transformers pipeline in src/insanely_fast_whisper/cli.py. This mirrors the CLI's --flash flag and requires the flash-attn package installed in your environment.
Can I run speaker diarization without using the CLI token arguments?
Yes. Import diarize from src/insanely_fast_whisper/utils/diarization_pipeline.py and pass a custom class or argparse.Namespace object containing file_name, device_id, hf_token, and speaker count constraints. The function loads the PyAnnote pipeline internally using your token.
What file contains the pipeline logic that the CLI uses?
The core pipeline instantiation logic resides in src/insanely_fast_whisper/cli.py, which creates the pipeline("automatic-speech-recognition", …) object. You can import and reuse this logic directly, or copy the pattern into your own scripts for custom batch processing workflows.
How do I process multiple audio files in a loop using the Python API?
Instantiate the pipeline once outside your loop, then iterate over file paths calling pipe() on each audio file. Collect the outputs and pass them to build_result() individually or aggregate them into a list. Since the pipeline persists in memory, this avoids the initialization overhead of repeated CLI calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →