# How to Use `--min-speakers` and `--max-speakers` for Flexible Diarization in Insanely-Fast-Whisper

> Control speaker count in insanely-fast-whisper using min-speakers and max-speakers parameters for flexible diarization. Automatically discover optimal speaker numbers within your specified range.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: how-to-guide
- Published: 2026-03-27

---

**Use `--min-speakers` and `--max-speakers` together to constrain the PyAnnote diarization pipeline to a specific range, allowing automatic discovery of the optimal speaker count within those bounds.**

The `insanely-fast-whisper` repository leverages the PyAnnote speaker diarization pipeline to identify who is speaking when in audio recordings. While the CLI accepts a fixed `--num-speakers` value for scenarios with known participant counts, it also supports **flexible diarization** through range-based constraints that let the model adapt to varying numbers of speakers without hardcoding an exact value.

## Understanding Speaker Count Constraints

The diarization system operates in three distinct modes depending on which arguments you provide:

*   **`--num-speakers N`** – Forces the pipeline to assume exactly N speakers throughout the audio.
*   **`--min-speakers N`** – Sets a lower bound; the pipeline will not return fewer than N speakers.
*   **`--max-speakers N`** – Sets an upper bound; the pipeline will not exceed N speakers.

When you supply both `--min-speakers` and `--max-speakers` while omitting `--num-speakers`, the PyAnnote pipeline internally executes a clustering step that searches for the best-fitting speaker count within your specified range. This is particularly useful for processing meetings, podcasts, or call center recordings where participant numbers fluctuate.

## The Technical Implementation

### Audio Preprocessing and Pipeline Setup

Before diarization begins, the `preprocess_inputs` function in [`src/insanely_fast_whisper/utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarize.py) loads the audio file and converts it to a 16 kHz mono `torch.FloatTensor`. The system then instantiates the diarization pipeline using `Pipeline.from_pretrained()` with the default `pyannote/speaker-diarization` model.

### Executing Range-Based Diarization

The `diarize_audio` function receives your speaker constraints and forwards them directly to the PyAnnote pipeline. As implemented in [`src/insanely_fast_whisper/utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarize.py) (lines 61-67):

```python
diarization = diarization_pipeline(
    {"waveform": diarizer_inputs, "sample_rate": 16000},
    num_speakers=num_speakers,
    min_speakers=min_speakers,
    max_speakers=max_speakers,
)

```

The orchestration layer in [`src/insanely_fast_whisper/utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarization_pipeline.py) (lines 27-28) calls this function with CLI-provided arguments:

```python
segments = diarize_audio(
    diarizer_inputs,
    diarization_pipeline,
    args.num_speakers,
    args.min_speakers,
    args.max_speakers,
)

```

After the pipeline returns raw diarization tracks, the system merges contiguous speaker turns and aligns them with ASR chunks to produce the final timestamped transcript.

## Command-Line and API Usage

### Basic CLI Examples

To let the model automatically determine whether your recording contains 2, 3, 4, or 5 speakers:

```bash
insanely-fast-whisper \
    --file-name my_meeting.wav \
    --min-speakers 2 \
    --max-speakers 5 \
    --model-name openai/whisper-large-v3 \
    --device-id 0

```

To force exactly three speakers when you know the participant count:

```bash
insanely-fast-whisper \
    --file-name interview.wav \
    --num-speakers 3

```

### Programmatic Integration

You can access the diarization functionality directly in Python for notebook workflows or custom applications:

```python
from insanely_fast_whisper.utils.diarization_pipeline import diarize

args = type(
    "Args",
    (),
    {
        "file_name": "podcast.wav",
        "diarization_model": "pyannote/speaker-diarization",
        "hf_token": "hf_********",
        "device_id": "0",
        "num_speakers": None,
        "min_speakers": 1,
        "max_speakers": 4,
    },
)()
asr_output = {"chunks": [...]}  # Result from Whisper inference

speaker_segments = diarize(args, asr_output)

```

## Summary

*   **`--min-speakers` and `--max-speakers`** enable flexible diarization by defining a searchable range rather than a fixed count.
*   The arguments are mutually exclusive with `--num-speakers`; omit the latter to activate range-based clustering.
*   The implementation passes these constraints directly to the PyAnnote pipeline in [`src/insanely_fast_whisper/utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarize.py).
*   Range-based diarization is ideal for dynamic audio where participant numbers are unknown or variable.
*   Both CLI and programmatic APIs support these constraints for batch processing or interactive workflows.

## Frequently Asked Questions

### Can I use `--min-speakers` without `--max-speakers`?

Yes. You can specify only a lower bound (`--min-speakers 2`) or only an upper bound (`--max-speakers 5`). The PyAnnote pipeline will use the provided constraint while allowing the speaker count to vary unbounded in the other direction. However, for most production scenarios, defining both bounds yields more consistent results.

### What happens if I specify `--num-speakers` alongside the range options?

If you provide `--num-speakers`, the pipeline ignores `--min-speakers` and `--max-speakers`, enforcing the exact count you specified. These arguments are designed to be mutually exclusive; use range options only when the participant count is unknown.

### Which diarization model does insanely-fast-whisper use internally?

By default, the system uses the `pyannote/speaker-diarization` model from Hugging Face. You can override this with the `--diarization-model` flag to point to a different PyAnnote-compatible checkpoint, including private models requiring an authentication token via `--hf-token`.

### How does the clustering algorithm determine the final speaker count within the range?

The PyAnnote pipeline employs a hierarchical clustering mechanism that evaluates multiple speaker count hypotheses between your minimum and maximum values. It selects the solution that optimizes internal clustering metrics (typically based on speaker embeddings and temporal continuity), effectively choosing the "best fit" number of speakers that minimizes within-speaker variance while respecting your constraints.