How to Specify the Exact Number of Speakers in Diarization with Insanely-Fast-Whisper

Use the --num-speakers N flag to force the PyAnnote diarizer to output exactly N speaker tracks instead of automatically estimating the speaker count.

When processing multi-speaker audio, you may need to specify the exact number of speakers in diarization with insanely-fast-whisper to obtain deterministic segmentation. This open-source tool integrates OpenAI Whisper with PyAnnote models, exposing a --num-speakers parameter that bypasses heuristic speaker counting and locks the output to your specified cardinality.

Using the --num-speakers Flag in the CLI

The command-line interface provides the most straightforward method to fix the speaker count. The argument is defined in src/insanely_fast_whisper/cli.py at lines 90-95, where the parser accepts an integer value passed directly to the underlying pipeline.

insanely-fast-whisper \
    --file-name podcast.wav \
    --model-name openai/whisper-large-v3 \
    --diarization_model pyannote/speaker-diarization-3.1 \
    --num-speakers 3 \
    --hf-token YOUR_HF_TOKEN \
    --transcript-path output.json

In this example, --num-speakers 3 instructs the diarizer to assign exactly three speaker labels (speaker_0, speaker_1, speaker_2) regardless of the audio content.

Validation Constraints

The source code enforces strict validation logic in src/insanely_fast_whisper/cli.py to prevent conflicting arguments:

  • Mutual Exclusion: You cannot combine --num-speakers with --min-speakers or --max-speakers. The validation at lines 14-19 aborts execution with the error: "--num-speakers cannot be used together with --min-speakers or --max-speakers."
  • Minimum Value: The value must be at least 1. The check at lines 11-13 raises: "--num-speakers must be at least 1."

Programmatic Usage in Python

When integrating insanely-fast-whisper into a Python application, the num_speakers argument propagates through three critical modules before reaching the PyAnnote pipeline.

The data flow follows this path:

  1. cli.py parses the argument from sys.argv
  2. diarize() in src/insanely_fast_whisper/utils/diarization_pipeline.py receives the value (lines 25-27)
  3. diarize_audio() in src/insanely_fast_whisper/utils/diarize.py passes it to the PyAnnote Pipeline.__call__ method (lines 61-67)
from insanely_fast_whisper.cli import parser, main

# Simulate CLI arguments programmatically

args = parser.parse_args([
    "--file-name", "interview.wav",
    "--model-name", "openai/whisper-large-v3",
    "--diarization_model", "pyannote/speaker-diarization-3.1",
    "--num-speakers", "2",
    "--hf-token", "YOUR_HF_TOKEN",
    "--transcript-path", "result.json"
])

# Execute the pipeline

main()

Because this approach reuses the same ArgumentParser instance defined in cli.py, all validation rules apply automatically.

Inspecting the Diarization Output

After processing, the transcript JSON contains speaker labels corresponding to your specified count. You can verify the assignment by examining the segments generated in src/insanely_fast_whisper/utils/result.py:

from insanely_fast_whisper.utils.result import build_result

# Assuming speakers_transcript and outputs exist from the pipeline

result = build_result(speakers_transcript, outputs)

# Verify exactly 2 speakers appear when --num-speakers 2 was used

unique_speakers = set(seg["speaker"] for seg in result["segments"])
print(f"Detected speakers: {unique_speakers}")  # Output: {'speaker_0', 'speaker_1'}

Each segment dictionary includes a "speaker" key (e.g., "speaker_0"), and the distinct label count will match your --num-speakers value exactly.

Summary

  • Fix speaker count: Pass --num-speakers N to force exactly N speaker clusters in the output.
  • Avoid conflicts: Never combine with --min-speakers or --max-speakers; the CLI enforces mutual exclusion.
  • Validation: Values must be integers ≥ 1.
  • Pipeline flow: The parameter travels from cli.py → diarization_pipeline.py → diarize.py before reaching the PyAnnote model.
  • Deterministic output: Specifying the count produces consistent speaker labels across repeated runs on the same audio.

Frequently Asked Questions

Can I use --num-speakers with --min-speakers or --max-speakers?

No. The validation logic in src/insanely_fast_whisper/cli.py explicitly forbids combining these arguments. The program aborts immediately with an error stating that --num-speakers cannot be used together with boundary constraints.

What happens if I set --num-speakers to 0?

The CLI rejects values less than 1 and exits with the error: "--num-speakers must be at least 1." This check occurs before any model inference begins.

Does specifying the speaker count improve transcription accuracy?

No. The --num-speakers flag affects only speaker attribution (determining which segments belong to which speaker). It does not alter the Whisper ASR model's transcription quality or word error rate.

Which PyAnnote models support the num_speakers parameter?

Any diarization model compatible with insanely-fast-whisper that implements the standard PyAnnote pipeline interface supports this feature. The commonly used pyannote/speaker-diarization-3.1 model fully supports explicit speaker count specification through the method signature referenced in src/insanely_fast_whisper/utils/diarize.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →