# How to Specify the Exact Number of Speakers in Diarization with Insanely-Fast-Whisper

> Control speaker count in audio diarization with insanely-fast-whisper. Easily specify the exact number of speakers using the --num-speakers N flag for precise results.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: how-to-guide
- Published: 2026-03-27

---

**Use the `--num-speakers N` flag to force the PyAnnote diarizer to output exactly N speaker tracks instead of automatically estimating the speaker count.**

When processing multi-speaker audio, you may need to specify the exact number of speakers in diarization with insanely-fast-whisper to obtain deterministic segmentation. This open-source tool integrates OpenAI Whisper with PyAnnote models, exposing a `--num-speakers` parameter that bypasses heuristic speaker counting and locks the output to your specified cardinality.

## Using the `--num-speakers` Flag in the CLI

The command-line interface provides the most straightforward method to fix the speaker count. The argument is defined in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) at lines 90-95, where the parser accepts an integer value passed directly to the underlying pipeline.

```bash
insanely-fast-whisper \
    --file-name podcast.wav \
    --model-name openai/whisper-large-v3 \
    --diarization_model pyannote/speaker-diarization-3.1 \
    --num-speakers 3 \
    --hf-token YOUR_HF_TOKEN \
    --transcript-path output.json

```

In this example, `--num-speakers 3` instructs the diarizer to assign exactly three speaker labels (`speaker_0`, `speaker_1`, `speaker_2`) regardless of the audio content.

### Validation Constraints

The source code enforces strict validation logic in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) to prevent conflicting arguments:

- **Mutual Exclusion**: You cannot combine `--num-speakers` with `--min-speakers` or `--max-speakers`. The validation at lines 14-19 aborts execution with the error: *"--num-speakers cannot be used together with --min-speakers or --max-speakers."*
- **Minimum Value**: The value must be at least 1. The check at lines 11-13 raises: *"--num-speakers must be at least 1."*

## Programmatic Usage in Python

When integrating insanely-fast-whisper into a Python application, the `num_speakers` argument propagates through three critical modules before reaching the PyAnnote pipeline.

The data flow follows this path:
1. [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py) parses the argument from `sys.argv`
2. `diarize()` in [`src/insanely_fast_whisper/utils/diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarization_pipeline.py) receives the value (lines 25-27)
3. `diarize_audio()` in [`src/insanely_fast_whisper/utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarize.py) passes it to the PyAnnote `Pipeline.__call__` method (lines 61-67)

```python
from insanely_fast_whisper.cli import parser, main

# Simulate CLI arguments programmatically

args = parser.parse_args([
    "--file-name", "interview.wav",
    "--model-name", "openai/whisper-large-v3",
    "--diarization_model", "pyannote/speaker-diarization-3.1",
    "--num-speakers", "2",
    "--hf-token", "YOUR_HF_TOKEN",
    "--transcript-path", "result.json"
])

# Execute the pipeline

main()

```

Because this approach reuses the same `ArgumentParser` instance defined in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py), all validation rules apply automatically.

## Inspecting the Diarization Output

After processing, the transcript JSON contains speaker labels corresponding to your specified count. You can verify the assignment by examining the segments generated in [`src/insanely_fast_whisper/utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/result.py):

```python
from insanely_fast_whisper.utils.result import build_result

# Assuming speakers_transcript and outputs exist from the pipeline

result = build_result(speakers_transcript, outputs)

# Verify exactly 2 speakers appear when --num-speakers 2 was used

unique_speakers = set(seg["speaker"] for seg in result["segments"])
print(f"Detected speakers: {unique_speakers}")  # Output: {'speaker_0', 'speaker_1'}

```

Each segment dictionary includes a `"speaker"` key (e.g., `"speaker_0"`), and the distinct label count will match your `--num-speakers` value exactly.

## Summary

- **Fix speaker count**: Pass `--num-speakers N` to force exactly N speaker clusters in the output.
- **Avoid conflicts**: Never combine with `--min-speakers` or `--max-speakers`; the CLI enforces mutual exclusion.
- **Validation**: Values must be integers ≥ 1.
- **Pipeline flow**: The parameter travels from [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py) → [`diarization_pipeline.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/diarization_pipeline.py) → [`diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/diarize.py) before reaching the PyAnnote model.
- **Deterministic output**: Specifying the count produces consistent speaker labels across repeated runs on the same audio.

## Frequently Asked Questions

### Can I use `--num-speakers` with `--min-speakers` or `--max-speakers`?

No. The validation logic in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) explicitly forbids combining these arguments. The program aborts immediately with an error stating that `--num-speakers` cannot be used together with boundary constraints.

### What happens if I set `--num-speakers` to 0?

The CLI rejects values less than 1 and exits with the error: *"--num-speakers must be at least 1."* This check occurs before any model inference begins.

### Does specifying the speaker count improve transcription accuracy?

No. The `--num-speakers` flag affects only speaker attribution (determining which segments belong to which speaker). It does not alter the Whisper ASR model's transcription quality or word error rate.

### Which PyAnnote models support the `num_speakers` parameter?

Any diarization model compatible with insanely-fast-whisper that implements the standard PyAnnote pipeline interface supports this feature. The commonly used `pyannote/speaker-diarization-3.1` model fully supports explicit speaker count specification through the method signature referenced in [`src/insanely_fast_whisper/utils/diarize.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/diarize.py).