How to Get Word-Level Timestamps in Insanely-Fast-Whisper for Precise Timing

Use the --timestamp word CLI flag to enable word-level timestamp generation instead of the default chunk-level timestamps.

Insanely-fast-whisper is a high-performance wrapper around Hugging Face's automatic speech recognition pipeline that accelerates Whisper transcription. While the tool defaults to chunk-level timestamps (approximately 30-second segments), you can extract word-level timestamps for frame-accurate subtitle creation and precise audio-text alignment by modifying a single command-line argument.

Understanding the --timestamp Flag in the CLI

The entry point for timestamp control resides in src/insanely_fast_whisper/cli.py. The argument parser explicitly defines the granularity options:

parser.add_argument(
    "--timestamp",
    required=False,
    type=str,
    default="chunk",
    choices=["chunk", "word"],
    help="Whisper supports both chunked as well as word level timestamps. (default: chunk)",
)

Source: cli.py L68-L73

When you pass --timestamp word, the library switches from segment-based timing to token-level timing, giving you start and end times for every individual spoken word.

How Word-Level Timestamps Work Under the Hood

Mapping CLI Arguments to the Pipeline

After parsing, the CLI translates your --timestamp value into the return_timestamps parameter that the underlying Hugging Face pipeline expects. In src/insanely_fast_whisper/cli.py, the logic converts the string flag to the appropriate pipeline argument:

ts = "word" if args.timestamp == "word" else True
…
outputs = pipe(
    args.file_name,
    chunk_length_s=30,
    batch_size=args.batch_size,
    generate_kwargs=generate_kwargs,
    return_timestamps=ts,
)

Source: cli.py L43-L65

  • When --timestamp word is supplied, ts becomes the string "word", triggering the pipeline to return per-token timestamps.
  • When omitted or set to chunk, ts evaluates to boolean True, which produces standard chunk-level timestamps only.

Output Structure in result.py

The pipeline returns a dictionary containing a "chunks" key that holds timestamp metadata. The build_result function in src/insanely_fast_whisper/utils/result.py packages this directly into the final JSON output without transformation:

def build_result(transcript, outputs) -> JsonTranscriptionResult:
    return {
        "speakers": transcript,
        "chunks": outputs["chunks"],
        "text": outputs["text"],
    }

Source: result.py L10-L15

Each element in the "chunks" array contains a "timestamp" key with [start, end] values. When word-level mode is active, these intervals represent individual words rather than 30-second segments.

Converting Word-Level Timestamps to Subtitle Files

The repository includes a convert_output.py utility that transforms the JSON output into SRT, VTT, or plain-text formats. The conversion logic reads the timestamp field generically, meaning it works seamlessly with both chunk and word granularity:

start, end = chunk['timestamp'][0], chunk['timestamp'][1]

Source: convert_output.py L36-L38

Because the converter simply iterates over the "chunks" array and extracts the timestamp pairs, word-level timestamps automatically produce subtitles where each word appears at its precise temporal location.

Complete Usage Examples

Generate a transcription with word-level timestamps using the CLI module:


# Basic word-level transcription

python -m insanely_fast_whisper.cli \
    --file-name audio.wav \
    --timestamp word \
    --output-path output.json

Combine word-level timestamps with speaker diarization by providing a Hugging Face token:

python -m insanely_fast_whisper.cli \
    --file-name audio.wav \
    --timestamp word \
    --hf-token YOUR_HF_TOKEN \
    --diarization_model pyannote/speaker-diarization-3.1

Convert the resulting JSON to SRT subtitles with per-word timing:

python convert_output.py output.json -f srt -o ./subtitles

The resulting output.json contains a "chunks" array where each object includes the precise start and end times for every spoken word, enabling frame-accurate caption editing and downstream temporal analysis.

Summary

  • Flag location: The --timestamp argument is defined in src/insanely_fast_whisper/cli.py with choices ["chunk", "word"].
  • Pipeline integration: The flag maps to return_timestamps="word" in the Hugging Face pipeline call.
  • Output format: Word timestamps appear in the "chunks" array of the JSON output generated by utils/result.py.
  • Subtitle conversion: The convert_output.py utility processes word-level timestamps into SRT/VTT files without requiring additional parameters.
  • Diarization compatibility: Word-level timestamps work concurrently with speaker diarization when using the --hf-token flag.

Frequently Asked Questions

What is the default timestamp granularity in insanely-fast-whisper?

By default, insanely-fast-whisper uses chunk-level timestamps (approximately 30-second segments). This is controlled by the --timestamp parameter defaulting to "chunk" in src/insanely_fast_whisper/cli.py.

Can I use word-level timestamps with speaker diarization simultaneously?

Yes. Pass both --timestamp word and --hf-token YOUR_TOKEN flags when running the CLI. The resulting JSON will contain both the "speakers" array from the diarization pipeline and the "chunks" array with word-level timestamps from the Whisper pipeline.

How do I convert word-level timestamps to SRT format?

Use the included convert_output.py script. The utility reads the "timestamp" field from each chunk in the JSON output, so it automatically handles word-level granularity:

python convert_output.py output.json -f srt -o subtitles.srt

Does enabling word-level timestamps impact transcription speed?

Word-level timestamps require the Whisper model to compute alignment for each token, which adds computational overhead compared to chunk-level mode. However, insanely-fast-whisper's batching optimizations in cli.py (via the batch_size parameter) help mitigate this performance difference when processing long audio files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →