Choosing Between Chunked and Word-Level Timestamps in insanely-fast-whisper

Use chunked timestamps (default) for fastest processing when coarse alignment suffices, or specify --timestamp word to generate per-word timing data required for precise subtitles and editing workflows.

When working with the Vaibhavs10/insanely-fast-whisper repository, you control transcription granularity through the --timestamp CLI flag. This guide examines both timestamp modes, demonstrates their configuration, and traces the implementation from argument parsing through the Hugging Face pipeline invocation.

Chunked vs Word-Level: Key Differences

insanely-fast-whisper offers two timestamp granularities that trade speed against precision:

Chunked timestamps (default) provide a single start and end time for each Whisper output chunk, typically spanning 30 seconds. This mode delivers the fastest inference and works well for chapter markers, rough navigation, or when processing large batches where exact word timing is unnecessary.

Word-level timestamps generate individual start and end times for every transcribed word. This precision is essential for professional subtitle generation, fine-grained video editing, and speaker diarization pipelines that require exact alignment between speakers and transcript content.

How to Select Your Timestamp Mode

Control timestamp behavior using the --timestamp argument in the CLI:


# Default chunked timestamps

python -m insanely_fast_whisper.cli --file-name audio.mp3

# Enable word-level precision

python -m insanely_fast_whisper.cli --file-name audio.mp3 --timestamp word

The --timestamp flag accepts "chunk" or "word" values, with "chunk" set as the default in src/insanely_fast_whisper/cli.py at lines 68-73:

parser.add_argument(
    "--timestamp",
    type=str,
    default="chunk",
    choices=["chunk", "word"],
    help="Choose between chunk and word timestamps",
)

Implementation: From CLI Argument to Whisper Pipeline

The translation from user input to Whisper API parameters happens in three stages within the CLI implementation.

First, the code maps the string argument to the expected pipeline value at line 143 of cli.py:

ts = "word" if args.timestamp == "word" else True

This logic assigns "word" when word-level timestamps are requested, or True for standard chunked timestamps.

Second, this value passes directly to the Hugging Face pipeline call at lines 159-165:

outputs = pipe(
    args.file_name,
    chunk_length_s=30,
    batch_size=args.batch_size,
    generate_kwargs=generate_kwargs,
    return_timestamps=ts,
)

Here, return_timestamps receives either the string "word" or the boolean True, which instructs the underlying Whisper model to return the appropriate timestamp granularity.

Finally, the build_result function in src/insanely_fast_whisper/utils/result.py (lines 10-15) packages these timestamps into the JSON output structure, storing transcription segments in the chunks key regardless of which mode was selected.

Output Format and Practical Usage

Chunked Timestamp Output

When using the default chunked mode, each entry in the chunks array contains one timestamp pair representing the start and end of that segment:

{
  "speakers": [],
  "chunks": [
    {
      "text": "Hello, this is a test.",
      "timestamp": [0.0, 30.0]
    }
  ],
  "text": "Hello, this is a test. ..."
}

Word-Level Timestamp Output

With --timestamp word, the timestamp field becomes an array of [start, end] pairs, one for each word:

{
  "speakers": [],
  "chunks": [
    {
      "text": "Hello, this is a test.",
      "timestamp": [
        [0.0, 0.45],
        [0.45, 0.78],
        [0.78, 1.01],
        [1.01, 1.30],
        [1.30, 1.65]
      ]
    }
  ],
  "text": "Hello, this is a test. ..."
}

Converting to Subtitle Formats

The repository includes convert_output.py to transform JSON transcripts into subtitle files. When word-level timestamps are present, the conversion generates separate subtitle entries for each word, enabling frame-accurate captions:

python convert_output.py transcript_word.json -f srt -o ./subs

The script reads the timestamp structure from the JSON and formats it appropriately for SRT, VTT, or plain-text output, preserving the granularity selected during transcription.

Speaker Diarization Compatibility

When combining timestamps with speaker diarization (using --hf-token), the alignment logic in src/insanely_fast_whisper/utils/diarize.py processes both timestamp formats seamlessly. The diarization pipeline maps speaker segments to Whisper timestamps whether they represent 30-second chunks or individual words, ensuring consistent speaker labeling regardless of your granularity choice.

Summary

  • Chunked timestamps (default) optimize for speed and provide 30-second segment alignment via return_timestamps=True in the Hugging Face pipeline.
  • Word-level timestamps require --timestamp word and set return_timestamps="word" to generate per-word timing arrays.
  • The --timestamp argument is defined in src/insanely_fast_whisper/cli.py with choices ["chunk", "word"] and defaults to "chunk".
  • Both timestamp formats integrate with speaker diarization through utils/diarize.py and convert to subtitles via convert_output.py.
  • Choose chunked mode for batch processing and coarse navigation; select word-level for subtitles, precise editing, and detailed alignment tasks.

Frequently Asked Questions

Does word-level timestamping slow down transcription?

Word-level timestamps require additional processing in the Whisper model to align each token with audio timings, which adds slight overhead compared to chunked mode. However, insanely-fast-whisper's optimized pipeline minimizes this impact, making word-level timestamps viable even for long-form content when precision is required.

Can I convert chunked timestamps to word-level after transcription?

No, the timestamp granularity is determined at inference time through the return_timestamps parameter passed to the Hugging Face pipeline. You must re-run transcription with --timestamp word to obtain per-word timing; the coarse timestamps do not contain sufficient data to reconstruct word boundaries.

How do I use word-level timestamps for subtitle generation?

Pass your JSON output to convert_output.py after transcription with --timestamp word. The conversion script processes the nested timestamp arrays to create SRT or VTT files where each word appears as a separate subtitle entry with precise timing, ideal for frame-accurate captioning workflows.

Are word-level timestamps compatible with speaker diarization?

Yes, the diarization pipeline in src/insanely_fast_whisper/utils/diarize.py aligns speaker segments with Whisper timestamps regardless of granularity. Whether you use chunked or word-level timestamps, the system maps speaker labels to the appropriate text segments, though word-level timing provides more precise speaker change detection at word boundaries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →