How to Provide a HuggingFace Token for pyannote.audio Access in insanely-fast-whisper

Pass your HuggingFace access token to insanely-fast-whisper using the --hf-token CLI argument so pyannote.audio can authenticate against restricted model repositories and execute speaker diarization.

The insanely-fast-whisper (IFS) CLI is a thin wrapper around HuggingFace Transformers' Whisper pipeline that optionally adds speaker diarization via pyannote.audio. Because pyannote's speaker-diarization checkpoints are hosted in gated HuggingFace Hub repositories, providing a valid HuggingFace token is mandatory when you want to identify "who spoke when" in your audio files.

Why a HuggingFace Token is Required for pyannote.audio

The pyannote.audio library downloads its pretrained models—such as pyannote/speaker-diarization-3.1—directly from the HuggingFace Hub. These repositories are restricted and require authentication to access.

Without a valid token, the HTTP request initiated by pyannote.audio returns a 401 Unauthorized error, and the diarization pipeline fails to initialize. You must generate a token with read scope from your HuggingFace account settings at hf.co/settings/token to grant the necessary access permissions.

How the Token Flows Through the Source Code

CLI Argument Parsing

In src/insanely_fast_whisper/cli.py, the --hf-token argument is defined with a default sentinel value of "no_token" to distinguish between an omitted flag and an empty string:


# src/insanely_fast_whisper/cli.py

parser.add_argument(
    "--hf-token",
    required=False,
    default="no_token",
    type=str,
    help="Provide a hf.co/settings/token for Pyannote.audio to diarise the audio clips",
)

The main execution logic checks this value before invoking the diarization subsystem:


# src/insanely_fast_whisper/cli.py

if args.hf_token != "no_token":
    speakers_transcript = diarize(args, outputs)
    ...

Pipeline Construction with Authentication

When the token is present, src/insanely_fast_whisper/utils/diarization_pipeline.py passes it to pyannote.audio's Pipeline.from_pretrained method via the use_auth_token parameter:


# src/insanely_fast_whisper/utils/diarization_pipeline.py

diarization_pipeline = Pipeline.from_pretrained(
    checkpoint_path=args.diarization_model,
    use_auth_token=args.hf_token,
)

This single argument propagates your credentials to the HuggingFace Hub downloader, allowing the restricted checkpoint to load into memory.

Providing Your HuggingFace Token

Basic CLI Usage

Run the transcription with diarization enabled by supplying your token via the command line:

insanely-fast-whisper \
    --file-name my_meeting.wav \
    --hf-token hf_XXXXXXXXXXXXXXXXXXXXXXXXXXXX \
    --diarization_model pyannote/speaker-diarization-3.1 \
    --output-path diarised_output.json
  • The --hf-token value must be a valid token string starting with hf_.
  • The --diarization_model flag is optional; if omitted, IFS defaults to pyannote/speaker-diarization-3.1.
  • The resulting JSON contains a speakers field mapping each transcription segment to a speaker label.

Using Environment Variables

While IFS does not automatically read from environment variables, you can forward a token stored in your shell environment to avoid exposing it in your command history:

export HF_TOKEN=hf_XXXXXXXXXXXXXXXXXXXXXXXXXXXX
insanely-fast-whisper \
    --file-name my_meeting.wav \
    --hf-token "$HF_TOKEN"

This approach keeps the token out of .bash_history or process listings while still satisfying the CLI requirement.

Programmatic Integration

If you are embedding IFS functionality within a larger Python application, you can invoke the CLI entry point directly with simulated arguments:

from insanely_fast_whisper.cli import parser, main
import sys

# Configure arguments programmatically

argv = [
    "--file-name", "my_meeting.wav",
    "--hf-token", "hf_XXXXXXXXXXXXXXXXXXXXXXXXXXXX",
    "--diarization_model", "pyannote/speaker-diarization-3.1",
]
sys.argv[1:] = argv

main()  # Executes the full transcription and diarization pipeline

Troubleshooting 401 Unauthorized Errors

If you encounter a 401 error during model loading, verify three specific details:

  1. Token validity – Ensure the token is active and has not been revoked in your HuggingFace settings.
  2. Repository access – Visit the model page (e.g., huggingface.co/pyannote/speaker-diarization-3.1) and accept the user agreement if you haven't already.
  3. Scope permissions – Confirm the token has read access to model repositories, not just inference API access.

Summary

  • The --hf-token argument in src/insanely_fast_whisper/cli.py accepts your HuggingFace credentials and defaults to "no_token" when omitted.
  • The token is passed to Pipeline.from_pretrained in src/insanely_fast_whisper/utils/diarization_pipeline.py via the use_auth_token parameter.
  • pyannote.audio requires this token because its diarization models are hosted in restricted repositories.
  • You can supply the token via direct CLI argument, environment variable forwarding, or programmatic argument injection.

Frequently Asked Questions

What permissions does my HuggingFace token need?

Your token requires the read scope to download model weights from the HuggingFace Hub. When creating the token at hf.co/settings/token, select the "Read" permission under "Repositories". Tokens with only inference API access will fail with a 401 error when attempting to load pyannote.audio checkpoints.

Why do I still get a 401 error after providing a token?

A 401 error typically indicates that your account has not accepted the user license agreement for the specific pyannote model. Navigate to the model's HuggingFace page (e.g., pyannote/speaker-diarization-3.1), click "Access repository," and accept the terms. The token must also belong to the same account that accepted these terms.

Can I store the token in a configuration file instead of the command line?

The current implementation of insanely-fast-whisper does not support configuration files or automatic environment variable detection for the HuggingFace token. You must explicitly pass the value via --hf-token on each invocation, or wrap the CLI call in a shell script or Python subprocess that injects the value.

Does insanely-fast-whisper log or store my HuggingFace token?

According to the source code in src/insanely_fast_whisper/cli.py, the token is used only for the immediate Pipeline.from_pretrained call and is not written to log files or persisted to disk. However, if you provide the token directly in the command line, it may be visible in your shell history or process monitors like ps. Use environment variable forwarding to minimize exposure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →