# Using the Python API Instead of CLI for Custom Pipelines in Insanely‑Fast‑Whisper

> Leverage the insanely-fast-whisper Python API for custom ASR pipelines. Gain full control over model selection, batching, Flash-Attention 2, and diarization. Avoid CLI subprocesses for efficient audio processing.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: how-to-guide
- Published: 2026-03-27

---

**The insanely-fast-whisper CLI is a thin wrapper around Hugging Face Transformers ASR pipelines that you can import directly from the source modules and orchestrate in Python, giving you full control over model selection, batch processing, Flash-Attention 2, and speaker diarization without subprocess calls.**

The `insanely-fast-whisper` repository ships with a command-line interface for rapid Whisper transcription, but the underlying implementation is fully exposed as a Python API. By importing the pipeline builders and utility functions from `src/insanely_fast_whisper/`, you can construct custom transcription workflows that bypass the CLI entirely while retaining access to all performance optimizations including Flash-Attention and PyAnnote speaker segmentation.

## Core Architecture and Source Modules

The project separates concerns across four key modules that the CLI merely wires together. Understanding these components allows you to reconstruct the pipeline programmatically:

- **CLI Entry Point** – Located in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py), this file parses arguments and constructs the base `pipeline("automatic-speech-recognition", …)` object with your selected model, dtype, device, and attention implementation.

- **Result Builder** – The `build_result(transcript, outputs)` function in [`src/insanely_fast_whisper/utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/result.py) normalizes the raw Whisper output into a JSON-compatible dictionary, copying the `chunks` and `text` fields and appending optional diarization data.

- **Diarization Pipeline** – [`src/insanely_fast_whisper/utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarization_pipeline.py) provides the high-level `diarize()` wrapper that loads a `pyannote.audio` pipeline, moves it to the correct device, and aligns speaker segments with ASR timestamps.

- **Low-Level Diarization Helpers** – [`src/insanely_fast_whisper/utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarize.py) contains the actual preprocessing and alignment logic, including `preprocess_inputs`, `diarize_audio`, and `post_process_segments_and_transcripts`.

Because these functions are public, you can swap models, adjust batch sizes on the fly, inject custom preprocessing, or replace the diarization backend entirely.

## Basic ASR Pipeline Implementation

To transcribe audio without invoking the CLI, instantiate the Transformers pipeline directly using the same defaults found in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py). The following snippet demonstrates word-level timestamps with Flash-Attention 2 enabled:

```python
import torch
from transformers import pipeline
from insanely_fast_whisper.utils.result import build_result

# Build the pipeline – mirrors the CLI defaults

pipe = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large-v3",
    torch_dtype=torch.float16,
    device="cuda:0",                     # or "mps" on macOS

    model_kwargs={"attn_implementation": "flash_attention_2"},
)

# Run inference – chunk size = 30s, batch size = 24, word-level timestamps

outputs = pipe(
    "my_audio_file.wav",
    chunk_length_s=30,
    batch_size=24,
    return_timestamps="word",
)

# Build the JSON result (no diarisation)

result = build_result([], outputs)

# Write to disk (optional)

import json
with open("output.json", "w", encoding="utf-8") as fp:
    json.dump(result, fp, ensure_ascii=False, indent=2)

```

The `model_kwargs` dictionary enables Flash-Attention 2 via `"flash_attention_2"`, while `return_timestamps="word"` generates the same granular timestamps produced by the CLI's `--timestamp word` flag. The `build_result` utility ensures the output schema matches the CLI's JSON format exactly.

## Adding Speaker Diarization Programmatically

To include speaker segmentation without the CLI, import the diarization utilities and provide a configuration object matching the CLI's argument structure. The `diarize()` function expects an object with attributes like `file_name`, `device_id`, and `hf_token`:

```python
import torch
from transformers import pipeline
from pyannote.audio import Pipeline
from insanely_fast_whisper.utils.diarization_pipeline import diarize
from insanely_fast_whisper.utils.result import build_result

# 1. ASR pipeline (same as before)

pipe = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large-v3",
    torch_dtype=torch.float16,
    device="cuda:0",
    model_kwargs={"attn_implementation": "flash_attention_2"},
)

# 2. Run ASR

outputs = pipe(
    "my_audio_file.wav",
    chunk_length_s=30,
    batch_size=24,
    return_timestamps="word",
)

# 3. Prepare configuration object for diarization

class Args:
    file_name = "my_audio_file.wav"
    device_id = "0"
    hf_token = "hf_XXXXXXXXXXXXXXXX"          # Your HF token with PyAnnote access

    diarization_model = "pyannote/speaker-diarization-3.1"
    num_speakers = None
    min_speakers = None
    max_speakers = None

args = Args()

# 4. Run diarisation – returns speaker-annotated chunks

speakers_transcript = diarize(args, outputs)

# 5. Combine ASR+diarisation into final JSON

result = build_result(speakers_transcript, outputs)

with open("output_with_speakers.json", "w", encoding="utf-8") as fp:
    import json
    json.dump(result, fp, ensure_ascii=False, indent=2)

```

The `diarize` function handles loading the PyAnnote pipeline via `Pipeline.from_pretrained`, moving it to the specified CUDA or MPS device, and aligning timestamps using the helpers in [`src/insanely_fast_whisper/utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarize.py).

## Custom Models and Batch Configuration

You can swap the underlying Whisper model or adjust memory constraints by modifying the pipeline arguments. This example uses `distil-whisper/large-v2` with a reduced batch size:

```python
import torch
from transformers import pipeline
from insanely_fast_whisper.utils.result import build_result

pipe = pipeline(
    "automatic-speech-recognition",
    model="distil-whisper/large-v2",
    torch_dtype=torch.float16,
    device="cuda:0",
    model_kwargs={"attn_implementation": "flash_attention_2"},
)

outputs = pipe(
    "meeting.wav",
    chunk_length_s=30,
    batch_size=8,          # Smaller batch to fit GPU memory

    return_timestamps=True,
)

result = build_result([], outputs)

with open("distil_output.json", "w", encoding="utf-8") as fp:
    import json
    json.dump(result, fp, ensure_ascii=False, indent=2)

```

Changing the `model` parameter automatically swaps the checkpoint without requiring code changes to the pipeline logic. You can further customize the workflow by importing [`convert_output.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/convert_output.py) utilities to generate SRT or VTT files from the resulting JSON.

## Summary

- **The insanely-fast-whisper CLI is optional** – all functionality resides in importable Python modules under `src/insanely_fast_whisper/`.
- **Use `pipeline()` from Transformers** with `model_kwargs={"attn_implementation": "flash_attention_2"}` to enable Flash-Attention 2.
- **Normalize outputs** with `build_result()` from [`utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/result.py) to maintain CLI-compatible JSON schemas.
- **Add diarization** by calling `diarize()` from [`utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/diarization_pipeline.py) with a configuration object containing your `hf_token` and device settings.
- **Customize freely** – swap models (e.g., `distil-whisper`), adjust `batch_size`, or loop over file batches without subprocess overhead.

## Frequently Asked Questions

### How do I enable Flash-Attention 2 when using the Python API?

Pass `model_kwargs={"attn_implementation": "flash_attention_2"}` when constructing the Transformers pipeline in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py). This mirrors the CLI's `--flash` flag and requires the `flash-attn` package installed in your environment.

### Can I run speaker diarization without using the CLI token arguments?

Yes. Import `diarize` from [`src/insanely_fast_whisper/utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarization_pipeline.py) and pass a custom class or `argparse.Namespace` object containing `file_name`, `device_id`, `hf_token`, and speaker count constraints. The function loads the PyAnnote pipeline internally using your token.

### What file contains the pipeline logic that the CLI uses?

The core pipeline instantiation logic resides in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py), which creates the `pipeline("automatic-speech-recognition", …)` object. You can import and reuse this logic directly, or copy the pattern into your own scripts for custom batch processing workflows.

### How do I process multiple audio files in a loop using the Python API?

Instantiate the pipeline once outside your loop, then iterate over file paths calling `pipe()` on each audio file. Collect the outputs and pass them to `build_result()` individually or aggregate them into a list. Since the pipeline persists in memory, this avoids the initialization overhead of repeated CLI calls.