How to Specify the Exact Number of Speakers in Diarization with Insanely-Fast-Whisper
Use the --num-speakers N flag to force the PyAnnote diarizer to output exactly N speaker tracks instead of automatically estimating the speaker count.
When processing multi-speaker audio, you may need to specify the exact number of speakers in diarization with insanely-fast-whisper to obtain deterministic segmentation. This open-source tool integrates OpenAI Whisper with PyAnnote models, exposing a --num-speakers parameter that bypasses heuristic speaker counting and locks the output to your specified cardinality.
Using the --num-speakers Flag in the CLI
The command-line interface provides the most straightforward method to fix the speaker count. The argument is defined in src/insanely_fast_whisper/cli.py at lines 90-95, where the parser accepts an integer value passed directly to the underlying pipeline.
insanely-fast-whisper \
--file-name podcast.wav \
--model-name openai/whisper-large-v3 \
--diarization_model pyannote/speaker-diarization-3.1 \
--num-speakers 3 \
--hf-token YOUR_HF_TOKEN \
--transcript-path output.json
In this example, --num-speakers 3 instructs the diarizer to assign exactly three speaker labels (speaker_0, speaker_1, speaker_2) regardless of the audio content.
Validation Constraints
The source code enforces strict validation logic in src/insanely_fast_whisper/cli.py to prevent conflicting arguments:
- Mutual Exclusion: You cannot combine
--num-speakerswith--min-speakersor--max-speakers. The validation at lines 14-19 aborts execution with the error: "--num-speakers cannot be used together with --min-speakers or --max-speakers." - Minimum Value: The value must be at least 1. The check at lines 11-13 raises: "--num-speakers must be at least 1."
Programmatic Usage in Python
When integrating insanely-fast-whisper into a Python application, the num_speakers argument propagates through three critical modules before reaching the PyAnnote pipeline.
The data flow follows this path:
cli.pyparses the argument fromsys.argvdiarize()insrc/insanely_fast_whisper/utils/diarization_pipeline.pyreceives the value (lines 25-27)diarize_audio()insrc/insanely_fast_whisper/utils/diarize.pypasses it to the PyAnnotePipeline.__call__method (lines 61-67)
from insanely_fast_whisper.cli import parser, main
# Simulate CLI arguments programmatically
args = parser.parse_args([
"--file-name", "interview.wav",
"--model-name", "openai/whisper-large-v3",
"--diarization_model", "pyannote/speaker-diarization-3.1",
"--num-speakers", "2",
"--hf-token", "YOUR_HF_TOKEN",
"--transcript-path", "result.json"
])
# Execute the pipeline
main()
Because this approach reuses the same ArgumentParser instance defined in cli.py, all validation rules apply automatically.
Inspecting the Diarization Output
After processing, the transcript JSON contains speaker labels corresponding to your specified count. You can verify the assignment by examining the segments generated in src/insanely_fast_whisper/utils/result.py:
from insanely_fast_whisper.utils.result import build_result
# Assuming speakers_transcript and outputs exist from the pipeline
result = build_result(speakers_transcript, outputs)
# Verify exactly 2 speakers appear when --num-speakers 2 was used
unique_speakers = set(seg["speaker"] for seg in result["segments"])
print(f"Detected speakers: {unique_speakers}") # Output: {'speaker_0', 'speaker_1'}
Each segment dictionary includes a "speaker" key (e.g., "speaker_0"), and the distinct label count will match your --num-speakers value exactly.
Summary
- Fix speaker count: Pass
--num-speakers Nto force exactly N speaker clusters in the output. - Avoid conflicts: Never combine with
--min-speakersor--max-speakers; the CLI enforces mutual exclusion. - Validation: Values must be integers ≥ 1.
- Pipeline flow: The parameter travels from
cli.py→diarization_pipeline.py→diarize.pybefore reaching the PyAnnote model. - Deterministic output: Specifying the count produces consistent speaker labels across repeated runs on the same audio.
Frequently Asked Questions
Can I use --num-speakers with --min-speakers or --max-speakers?
No. The validation logic in src/insanely_fast_whisper/cli.py explicitly forbids combining these arguments. The program aborts immediately with an error stating that --num-speakers cannot be used together with boundary constraints.
What happens if I set --num-speakers to 0?
The CLI rejects values less than 1 and exits with the error: "--num-speakers must be at least 1." This check occurs before any model inference begins.
Does specifying the speaker count improve transcription accuracy?
No. The --num-speakers flag affects only speaker attribution (determining which segments belong to which speaker). It does not alter the Whisper ASR model's transcription quality or word error rate.
Which PyAnnote models support the num_speakers parameter?
Any diarization model compatible with insanely-fast-whisper that implements the standard PyAnnote pipeline interface supports this feature. The commonly used pyannote/speaker-diarization-3.1 model fully supports explicit speaker count specification through the method signature referenced in src/insanely_fast_whisper/utils/diarize.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →