Using the Python API Instead of CLI for Custom Pipelines in Insanely‑Fast‑Whisper

The insanely-fast-whisper CLI is a thin wrapper around Hugging Face Transformers ASR pipelines that you can import directly from the source modules and orchestrate in Python, giving you full control over model selection, batch processing, Flash-Attention 2, and speaker diarization without subprocess calls.

The insanely-fast-whisper repository ships with a command-line interface for rapid Whisper transcription, but the underlying implementation is fully exposed as a Python API. By importing the pipeline builders and utility functions from src/insanely_fast_whisper/, you can construct custom transcription workflows that bypass the CLI entirely while retaining access to all performance optimizations including Flash-Attention and PyAnnote speaker segmentation.

Core Architecture and Source Modules

The project separates concerns across four key modules that the CLI merely wires together. Understanding these components allows you to reconstruct the pipeline programmatically:

  • CLI Entry Point – Located in src/insanely_fast_whisper/cli.py, this file parses arguments and constructs the base pipeline("automatic-speech-recognition", …) object with your selected model, dtype, device, and attention implementation.

  • Result Builder – The build_result(transcript, outputs) function in src/insanely_fast_whisper/utils/result.py normalizes the raw Whisper output into a JSON-compatible dictionary, copying the chunks and text fields and appending optional diarization data.

  • Diarization Pipeline – src/insanely_fast_whisper/utils/diarization_pipeline.py provides the high-level diarize() wrapper that loads a pyannote.audio pipeline, moves it to the correct device, and aligns speaker segments with ASR timestamps.

  • Low-Level Diarization Helpers – src/insanely_fast_whisper/utils/diarize.py contains the actual preprocessing and alignment logic, including preprocess_inputs, diarize_audio, and post_process_segments_and_transcripts.

Because these functions are public, you can swap models, adjust batch sizes on the fly, inject custom preprocessing, or replace the diarization backend entirely.

Basic ASR Pipeline Implementation

To transcribe audio without invoking the CLI, instantiate the Transformers pipeline directly using the same defaults found in src/insanely_fast_whisper/cli.py. The following snippet demonstrates word-level timestamps with Flash-Attention 2 enabled:

import torch
from transformers import pipeline
from insanely_fast_whisper.utils.result import build_result

# Build the pipeline – mirrors the CLI defaults

pipe = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large-v3",
    torch_dtype=torch.float16,
    device="cuda:0",                     # or "mps" on macOS

    model_kwargs={"attn_implementation": "flash_attention_2"},
)

# Run inference – chunk size = 30s, batch size = 24, word-level timestamps

outputs = pipe(
    "my_audio_file.wav",
    chunk_length_s=30,
    batch_size=24,
    return_timestamps="word",
)

# Build the JSON result (no diarisation)

result = build_result([], outputs)

# Write to disk (optional)

import json
with open("output.json", "w", encoding="utf-8") as fp:
    json.dump(result, fp, ensure_ascii=False, indent=2)

The model_kwargs dictionary enables Flash-Attention 2 via "flash_attention_2", while return_timestamps="word" generates the same granular timestamps produced by the CLI's --timestamp word flag. The build_result utility ensures the output schema matches the CLI's JSON format exactly.

Adding Speaker Diarization Programmatically

To include speaker segmentation without the CLI, import the diarization utilities and provide a configuration object matching the CLI's argument structure. The diarize() function expects an object with attributes like file_name, device_id, and hf_token:

import torch
from transformers import pipeline
from pyannote.audio import Pipeline
from insanely_fast_whisper.utils.diarization_pipeline import diarize
from insanely_fast_whisper.utils.result import build_result

# 1. ASR pipeline (same as before)

pipe = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large-v3",
    torch_dtype=torch.float16,
    device="cuda:0",
    model_kwargs={"attn_implementation": "flash_attention_2"},
)

# 2. Run ASR

outputs = pipe(
    "my_audio_file.wav",
    chunk_length_s=30,
    batch_size=24,
    return_timestamps="word",
)

# 3. Prepare configuration object for diarization

class Args:
    file_name = "my_audio_file.wav"
    device_id = "0"
    hf_token = "hf_XXXXXXXXXXXXXXXX"          # Your HF token with PyAnnote access

    diarization_model = "pyannote/speaker-diarization-3.1"
    num_speakers = None
    min_speakers = None
    max_speakers = None

args = Args()

# 4. Run diarisation – returns speaker-annotated chunks

speakers_transcript = diarize(args, outputs)

# 5. Combine ASR+diarisation into final JSON

result = build_result(speakers_transcript, outputs)

with open("output_with_speakers.json", "w", encoding="utf-8") as fp:
    import json
    json.dump(result, fp, ensure_ascii=False, indent=2)

The diarize function handles loading the PyAnnote pipeline via Pipeline.from_pretrained, moving it to the specified CUDA or MPS device, and aligning timestamps using the helpers in src/insanely_fast_whisper/utils/diarize.py.

Custom Models and Batch Configuration

You can swap the underlying Whisper model or adjust memory constraints by modifying the pipeline arguments. This example uses distil-whisper/large-v2 with a reduced batch size:

import torch
from transformers import pipeline
from insanely_fast_whisper.utils.result import build_result

pipe = pipeline(
    "automatic-speech-recognition",
    model="distil-whisper/large-v2",
    torch_dtype=torch.float16,
    device="cuda:0",
    model_kwargs={"attn_implementation": "flash_attention_2"},
)

outputs = pipe(
    "meeting.wav",
    chunk_length_s=30,
    batch_size=8,          # Smaller batch to fit GPU memory

    return_timestamps=True,
)

result = build_result([], outputs)

with open("distil_output.json", "w", encoding="utf-8") as fp:
    import json
    json.dump(result, fp, ensure_ascii=False, indent=2)

Changing the model parameter automatically swaps the checkpoint without requiring code changes to the pipeline logic. You can further customize the workflow by importing convert_output.py utilities to generate SRT or VTT files from the resulting JSON.

Summary

  • The insanely-fast-whisper CLI is optional – all functionality resides in importable Python modules under src/insanely_fast_whisper/.
  • Use pipeline() from Transformers with model_kwargs={"attn_implementation": "flash_attention_2"} to enable Flash-Attention 2.
  • Normalize outputs with build_result() from utils/result.py to maintain CLI-compatible JSON schemas.
  • Add diarization by calling diarize() from utils/diarization_pipeline.py with a configuration object containing your hf_token and device settings.
  • Customize freely – swap models (e.g., distil-whisper), adjust batch_size, or loop over file batches without subprocess overhead.

Frequently Asked Questions

How do I enable Flash-Attention 2 when using the Python API?

Pass model_kwargs={"attn_implementation": "flash_attention_2"} when constructing the Transformers pipeline in src/insanely_fast_whisper/cli.py. This mirrors the CLI's --flash flag and requires the flash-attn package installed in your environment.

Can I run speaker diarization without using the CLI token arguments?

Yes. Import diarize from src/insanely_fast_whisper/utils/diarization_pipeline.py and pass a custom class or argparse.Namespace object containing file_name, device_id, hf_token, and speaker count constraints. The function loads the PyAnnote pipeline internally using your token.

What file contains the pipeline logic that the CLI uses?

The core pipeline instantiation logic resides in src/insanely_fast_whisper/cli.py, which creates the pipeline("automatic-speech-recognition", …) object. You can import and reuse this logic directly, or copy the pattern into your own scripts for custom batch processing workflows.

How do I process multiple audio files in a loop using the Python API?

Instantiate the pipeline once outside your loop, then iterate over file paths calling pipe() on each audio file. Collect the outputs and pass them to build_result() individually or aggregate them into a list. Since the pipeline persists in memory, this avoids the initialization overhead of repeated CLI calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →