How to Use `--min-speakers` and `--max-speakers` for Flexible Diarization in Insanely-Fast-Whisper

Use --min-speakers and --max-speakers together to constrain the PyAnnote diarization pipeline to a specific range, allowing automatic discovery of the optimal speaker count within those bounds.

The insanely-fast-whisper repository leverages the PyAnnote speaker diarization pipeline to identify who is speaking when in audio recordings. While the CLI accepts a fixed --num-speakers value for scenarios with known participant counts, it also supports flexible diarization through range-based constraints that let the model adapt to varying numbers of speakers without hardcoding an exact value.

Understanding Speaker Count Constraints

The diarization system operates in three distinct modes depending on which arguments you provide:

  • --num-speakers N – Forces the pipeline to assume exactly N speakers throughout the audio.
  • --min-speakers N – Sets a lower bound; the pipeline will not return fewer than N speakers.
  • --max-speakers N – Sets an upper bound; the pipeline will not exceed N speakers.

When you supply both --min-speakers and --max-speakers while omitting --num-speakers, the PyAnnote pipeline internally executes a clustering step that searches for the best-fitting speaker count within your specified range. This is particularly useful for processing meetings, podcasts, or call center recordings where participant numbers fluctuate.

The Technical Implementation

Audio Preprocessing and Pipeline Setup

Before diarization begins, the preprocess_inputs function in src/insanely_fast_whisper/utils/diarize.py loads the audio file and converts it to a 16 kHz mono torch.FloatTensor. The system then instantiates the diarization pipeline using Pipeline.from_pretrained() with the default pyannote/speaker-diarization model.

Executing Range-Based Diarization

The diarize_audio function receives your speaker constraints and forwards them directly to the PyAnnote pipeline. As implemented in src/insanely_fast_whisper/utils/diarize.py (lines 61-67):

diarization = diarization_pipeline(
    {"waveform": diarizer_inputs, "sample_rate": 16000},
    num_speakers=num_speakers,
    min_speakers=min_speakers,
    max_speakers=max_speakers,
)

The orchestration layer in src/insanely_fast_whisper/utils/diarization_pipeline.py (lines 27-28) calls this function with CLI-provided arguments:

segments = diarize_audio(
    diarizer_inputs,
    diarization_pipeline,
    args.num_speakers,
    args.min_speakers,
    args.max_speakers,
)

After the pipeline returns raw diarization tracks, the system merges contiguous speaker turns and aligns them with ASR chunks to produce the final timestamped transcript.

Command-Line and API Usage

Basic CLI Examples

To let the model automatically determine whether your recording contains 2, 3, 4, or 5 speakers:

insanely-fast-whisper \
    --file-name my_meeting.wav \
    --min-speakers 2 \
    --max-speakers 5 \
    --model-name openai/whisper-large-v3 \
    --device-id 0

To force exactly three speakers when you know the participant count:

insanely-fast-whisper \
    --file-name interview.wav \
    --num-speakers 3

Programmatic Integration

You can access the diarization functionality directly in Python for notebook workflows or custom applications:

from insanely_fast_whisper.utils.diarization_pipeline import diarize

args = type(
    "Args",
    (),
    {
        "file_name": "podcast.wav",
        "diarization_model": "pyannote/speaker-diarization",
        "hf_token": "hf_********",
        "device_id": "0",
        "num_speakers": None,
        "min_speakers": 1,
        "max_speakers": 4,
    },
)()
asr_output = {"chunks": [...]}  # Result from Whisper inference

speaker_segments = diarize(args, asr_output)

Summary

  • --min-speakers and --max-speakers enable flexible diarization by defining a searchable range rather than a fixed count.
  • The arguments are mutually exclusive with --num-speakers; omit the latter to activate range-based clustering.
  • The implementation passes these constraints directly to the PyAnnote pipeline in src/insanely_fast_whisper/utils/diarize.py.
  • Range-based diarization is ideal for dynamic audio where participant numbers are unknown or variable.
  • Both CLI and programmatic APIs support these constraints for batch processing or interactive workflows.

Frequently Asked Questions

Can I use --min-speakers without --max-speakers?

Yes. You can specify only a lower bound (--min-speakers 2) or only an upper bound (--max-speakers 5). The PyAnnote pipeline will use the provided constraint while allowing the speaker count to vary unbounded in the other direction. However, for most production scenarios, defining both bounds yields more consistent results.

What happens if I specify --num-speakers alongside the range options?

If you provide --num-speakers, the pipeline ignores --min-speakers and --max-speakers, enforcing the exact count you specified. These arguments are designed to be mutually exclusive; use range options only when the participant count is unknown.

Which diarization model does insanely-fast-whisper use internally?

By default, the system uses the pyannote/speaker-diarization model from Hugging Face. You can override this with the --diarization-model flag to point to a different PyAnnote-compatible checkpoint, including private models requiring an authentication token via --hf-token.

How does the clustering algorithm determine the final speaker count within the range?

The PyAnnote pipeline employs a hierarchical clustering mechanism that evaluates multiple speaker count hypotheses between your minimum and maximum values. It selects the solution that optimizes internal clustering metrics (typically based on speaker embeddings and temporal continuity), effectively choosing the "best fit" number of speakers that minimizes within-speaker variance while respecting your constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →