How to Get Word-Level Timestamps in Insanely-Fast-Whisper for Precise Timing
Use the --timestamp word CLI flag to enable word-level timestamp generation instead of the default chunk-level timestamps.
Insanely-fast-whisper is a high-performance wrapper around Hugging Face's automatic speech recognition pipeline that accelerates Whisper transcription. While the tool defaults to chunk-level timestamps (approximately 30-second segments), you can extract word-level timestamps for frame-accurate subtitle creation and precise audio-text alignment by modifying a single command-line argument.
Understanding the --timestamp Flag in the CLI
The entry point for timestamp control resides in src/insanely_fast_whisper/cli.py. The argument parser explicitly defines the granularity options:
parser.add_argument(
"--timestamp",
required=False,
type=str,
default="chunk",
choices=["chunk", "word"],
help="Whisper supports both chunked as well as word level timestamps. (default: chunk)",
)
Source: cli.py L68-L73
When you pass --timestamp word, the library switches from segment-based timing to token-level timing, giving you start and end times for every individual spoken word.
How Word-Level Timestamps Work Under the Hood
Mapping CLI Arguments to the Pipeline
After parsing, the CLI translates your --timestamp value into the return_timestamps parameter that the underlying Hugging Face pipeline expects. In src/insanely_fast_whisper/cli.py, the logic converts the string flag to the appropriate pipeline argument:
ts = "word" if args.timestamp == "word" else True
…
outputs = pipe(
args.file_name,
chunk_length_s=30,
batch_size=args.batch_size,
generate_kwargs=generate_kwargs,
return_timestamps=ts,
)
Source: cli.py L43-L65
- When
--timestamp wordis supplied,tsbecomes the string"word", triggering the pipeline to return per-token timestamps. - When omitted or set to
chunk,tsevaluates to booleanTrue, which produces standard chunk-level timestamps only.
Output Structure in result.py
The pipeline returns a dictionary containing a "chunks" key that holds timestamp metadata. The build_result function in src/insanely_fast_whisper/utils/result.py packages this directly into the final JSON output without transformation:
def build_result(transcript, outputs) -> JsonTranscriptionResult:
return {
"speakers": transcript,
"chunks": outputs["chunks"],
"text": outputs["text"],
}
Source: result.py L10-L15
Each element in the "chunks" array contains a "timestamp" key with [start, end] values. When word-level mode is active, these intervals represent individual words rather than 30-second segments.
Converting Word-Level Timestamps to Subtitle Files
The repository includes a convert_output.py utility that transforms the JSON output into SRT, VTT, or plain-text formats. The conversion logic reads the timestamp field generically, meaning it works seamlessly with both chunk and word granularity:
start, end = chunk['timestamp'][0], chunk['timestamp'][1]
Source: convert_output.py L36-L38
Because the converter simply iterates over the "chunks" array and extracts the timestamp pairs, word-level timestamps automatically produce subtitles where each word appears at its precise temporal location.
Complete Usage Examples
Generate a transcription with word-level timestamps using the CLI module:
# Basic word-level transcription
python -m insanely_fast_whisper.cli \
--file-name audio.wav \
--timestamp word \
--output-path output.json
Combine word-level timestamps with speaker diarization by providing a Hugging Face token:
python -m insanely_fast_whisper.cli \
--file-name audio.wav \
--timestamp word \
--hf-token YOUR_HF_TOKEN \
--diarization_model pyannote/speaker-diarization-3.1
Convert the resulting JSON to SRT subtitles with per-word timing:
python convert_output.py output.json -f srt -o ./subtitles
The resulting output.json contains a "chunks" array where each object includes the precise start and end times for every spoken word, enabling frame-accurate caption editing and downstream temporal analysis.
Summary
- Flag location: The
--timestampargument is defined insrc/insanely_fast_whisper/cli.pywith choices["chunk", "word"]. - Pipeline integration: The flag maps to
return_timestamps="word"in the Hugging Face pipeline call. - Output format: Word timestamps appear in the
"chunks"array of the JSON output generated byutils/result.py. - Subtitle conversion: The
convert_output.pyutility processes word-level timestamps into SRT/VTT files without requiring additional parameters. - Diarization compatibility: Word-level timestamps work concurrently with speaker diarization when using the
--hf-tokenflag.
Frequently Asked Questions
What is the default timestamp granularity in insanely-fast-whisper?
By default, insanely-fast-whisper uses chunk-level timestamps (approximately 30-second segments). This is controlled by the --timestamp parameter defaulting to "chunk" in src/insanely_fast_whisper/cli.py.
Can I use word-level timestamps with speaker diarization simultaneously?
Yes. Pass both --timestamp word and --hf-token YOUR_TOKEN flags when running the CLI. The resulting JSON will contain both the "speakers" array from the diarization pipeline and the "chunks" array with word-level timestamps from the Whisper pipeline.
How do I convert word-level timestamps to SRT format?
Use the included convert_output.py script. The utility reads the "timestamp" field from each chunk in the JSON output, so it automatically handles word-level granularity:
python convert_output.py output.json -f srt -o subtitles.srt
Does enabling word-level timestamps impact transcription speed?
Word-level timestamps require the Whisper model to compute alignment for each token, which adds computational overhead compared to chunk-level mode. However, insanely-fast-whisper's batching optimizations in cli.py (via the batch_size parameter) help mitigate this performance difference when processing long audio files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →