# How to Get Word-Level Timestamps in Insanely-Fast-Whisper for Precise Timing

> Unlock precise timing with word-level timestamps in insanely-fast-whisper. Simply use the --timestamp word CLI flag for accurate audio analysis. Get started now!

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: how-to-guide
- Published: 2026-03-27

---

**Use the `--timestamp word` CLI flag to enable word-level timestamp generation instead of the default chunk-level timestamps.**

Insanely-fast-whisper is a high-performance wrapper around Hugging Face's automatic speech recognition pipeline that accelerates Whisper transcription. While the tool defaults to chunk-level timestamps (approximately 30-second segments), you can extract **word-level timestamps** for frame-accurate subtitle creation and precise audio-text alignment by modifying a single command-line argument.

## Understanding the `--timestamp` Flag in the CLI

The entry point for timestamp control resides in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py). The argument parser explicitly defines the granularity options:

```python
parser.add_argument(
    "--timestamp",
    required=False,
    type=str,
    default="chunk",
    choices=["chunk", "word"],
    help="Whisper supports both chunked as well as word level timestamps. (default: chunk)",
)

```

*Source: [cli.py L68-L73](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L68)*

When you pass `--timestamp word`, the library switches from segment-based timing to token-level timing, giving you start and end times for every individual spoken word.

## How Word-Level Timestamps Work Under the Hood

### Mapping CLI Arguments to the Pipeline

After parsing, the CLI translates your `--timestamp` value into the `return_timestamps` parameter that the underlying Hugging Face pipeline expects. In [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py), the logic converts the string flag to the appropriate pipeline argument:

```python
ts = "word" if args.timestamp == "word" else True
…
outputs = pipe(
    args.file_name,
    chunk_length_s=30,
    batch_size=args.batch_size,
    generate_kwargs=generate_kwargs,
    return_timestamps=ts,
)

```

*Source: [cli.py L43-L65](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py#L43)*

- When `--timestamp word` is supplied, `ts` becomes the string `"word"`, triggering the pipeline to return per-token timestamps.
- When omitted or set to `chunk`, `ts` evaluates to boolean `True`, which produces standard chunk-level timestamps only.

### Output Structure in result.py

The pipeline returns a dictionary containing a `"chunks"` key that holds timestamp metadata. The `build_result` function in [`src/insanely_fast_whisper/utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/result.py) packages this directly into the final JSON output without transformation:

```python
def build_result(transcript, outputs) -> JsonTranscriptionResult:
    return {
        "speakers": transcript,
        "chunks": outputs["chunks"],
        "text": outputs["text"],
    }

```

*Source: [result.py L10-L15](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/result.py#L10)*

Each element in the `"chunks"` array contains a `"timestamp"` key with `[start, end]` values. When word-level mode is active, these intervals represent individual words rather than 30-second segments.

## Converting Word-Level Timestamps to Subtitle Files

The repository includes a [`convert_output.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/convert_output.py) utility that transforms the JSON output into SRT, VTT, or plain-text formats. The conversion logic reads the timestamp field generically, meaning it works seamlessly with both chunk and word granularity:

```python
start, end = chunk['timestamp'][0], chunk['timestamp'][1]

```

*Source: [convert_output.py L36-L38](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/convert_output.py#L36)*

Because the converter simply iterates over the `"chunks"` array and extracts the timestamp pairs, word-level timestamps automatically produce subtitles where each word appears at its precise temporal location.

## Complete Usage Examples

Generate a transcription with word-level timestamps using the CLI module:

```bash

# Basic word-level transcription

python -m insanely_fast_whisper.cli \
    --file-name audio.wav \
    --timestamp word \
    --output-path output.json

```

Combine word-level timestamps with speaker diarization by providing a Hugging Face token:

```bash
python -m insanely_fast_whisper.cli \
    --file-name audio.wav \
    --timestamp word \
    --hf-token YOUR_HF_TOKEN \
    --diarization_model pyannote/speaker-diarization-3.1

```

Convert the resulting JSON to SRT subtitles with per-word timing:

```bash
python convert_output.py output.json -f srt -o ./subtitles

```

The resulting [`output.json`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/output.json) contains a `"chunks"` array where each object includes the precise start and end times for every spoken word, enabling frame-accurate caption editing and downstream temporal analysis.

## Summary

- **Flag location**: The `--timestamp` argument is defined in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) with choices `["chunk", "word"]`.
- **Pipeline integration**: The flag maps to `return_timestamps="word"` in the Hugging Face pipeline call.
- **Output format**: Word timestamps appear in the `"chunks"` array of the JSON output generated by [`utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/result.py).
- **Subtitle conversion**: The [`convert_output.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/convert_output.py) utility processes word-level timestamps into SRT/VTT files without requiring additional parameters.
- **Diarization compatibility**: Word-level timestamps work concurrently with speaker diarization when using the `--hf-token` flag.

## Frequently Asked Questions

### What is the default timestamp granularity in insanely-fast-whisper?

By default, insanely-fast-whisper uses **chunk-level timestamps** (approximately 30-second segments). This is controlled by the `--timestamp` parameter defaulting to `"chunk"` in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py).

### Can I use word-level timestamps with speaker diarization simultaneously?

Yes. Pass both `--timestamp word` and `--hf-token YOUR_TOKEN` flags when running the CLI. The resulting JSON will contain both the `"speakers"` array from the diarization pipeline and the `"chunks"` array with word-level timestamps from the Whisper pipeline.

### How do I convert word-level timestamps to SRT format?

Use the included [`convert_output.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/convert_output.py) script. The utility reads the `"timestamp"` field from each chunk in the JSON output, so it automatically handles word-level granularity:

```bash
python convert_output.py output.json -f srt -o subtitles.srt

```

### Does enabling word-level timestamps impact transcription speed?

Word-level timestamps require the Whisper model to compute alignment for each token, which adds computational overhead compared to chunk-level mode. However, insanely-fast-whisper's batching optimizations in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py) (via the `batch_size` parameter) help mitigate this performance difference when processing long audio files.