# Using the Translate Task Instead of Transcribe in Insanely-Fast-Whisper: A Complete Guide

> Learn how to use the translate task in insanely-fast-whisper with the --task translate flag. Get translated text instead of transcripts while keeping the same JSON output. Optimize your audio translation workflow.

- Repository: [vb/insanely-fast-whisper](https://github.com/Vaibhavs10/insanely-fast-whisper)
- Tags: how-to-guide
- Published: 2026-03-27

---

**Use the `--task translate` CLI flag to generate translated text instead of raw transcripts, leveraging Whisper's built-in translation capabilities while maintaining the same JSON output format.**

Insanely-Fast-Whisper accelerates OpenAI's Whisper model through Hugging Face's `transformers` pipeline. Understanding how to switch between the **`transcribe`** and **`translate`** tasks allows you to process multilingual audio without external translation tools. This guide examines the actual CLI implementation to show you exactly how the translation pipeline works.

## How the `--task` Parameter Controls Generation

The command-line interface defines the `--task` argument in [`src/insanely_fast_whisper/cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/cli.py) (lines 38-44) with strict validation:

```python
parser.add_argument(
    "--task",
    default="transcribe",
    choices=["transcribe", "translate"],
    help="Task to perform: transcribe or translate.",
)

```

**`transcribe`** (the default) returns text in the original spoken language. **`translate`** converts the speech into the target language specified by the `--language` parameter, or English if omitted.

## The `generate_kwargs` Mechanism

When you execute the CLI, the script constructs a `generate_kwargs` dictionary at lines 47-51 in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py) that passes your selected task directly to the Hugging Face pipeline:

```python
generate_kwargs = {"task": args.task, "language": language}
if args.model_name.split(".")[-1] == "en":
    generate_kwargs.pop("task")

```

This dictionary is injected into the pipeline call at lines 59-64, instructing Whisper's generation method to either transcribe verbatim or translate the audio content. The rest of the processing pipeline—including batching, chunking, and optional speaker diarization—remains identical regardless of which task you select.

## Automatic Handling of English-Only Models

The source code contains specific logic for English-only checkpoints. If your model name ends with `.en` (e.g., `openai/whisper-small.en`), the code automatically removes the task key from `generate_kwargs`:

```python
if args.model_name.split(".")[-1] == "en":
    generate_kwargs.pop("task")

```

**English-only models cannot perform translation** because they lack multilingual weights. Attempting to force a translate task on these checkpoints would cause runtime errors, so the wrapper safely defaults to transcription-only mode.

## Practical CLI Examples

The following commands demonstrate real-world usage patterns. Each produces a JSON output file formatted by `build_result` in [`src/insanely_fast_whisper/utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/result.py) (lines 10-15), which preserves the standard chunk and text structure regardless of task type.

### Standard Transcription

```bash
python -m insanely_fast_whisper.cli \
    --file-name audio.wav \
    --model-name openai/whisper-large-v3 \
    --task transcribe \
    --output-path transcript.json

```

### Translate to English (Auto-detect Source)

```bash
python -m insanely_fast_whisper.cli \
    --file-name french_audio.wav \
    --model-name openai/whisper-large-v3 \
    --task translate \
    --language en \
    --output-path translation.json

```

### Translate to Spanish

```bash
python -m insanely_fast_whisper.cli \
    --file-name english_audio.wav \
    --model-name openai/whisper-large-v3 \
    --task translate \
    --language es \
    --output-path spanish_translation.json

```

## Summary

- **Task selection** happens via `--task transcribe` or `--task translate` in the CLI, validated against a strict choices list in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py).
- **English-only models** (names ending in `.en`) automatically have the task parameter stripped to prevent compatibility errors.
- **Output consistency** is maintained through [`utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/utils/result.py), ensuring JSON structure remains identical whether translating or transcribing.
- **Language specification** uses the `--language` flag alongside `--task translate` to set the target tongue, or defaults to English if omitted.

## Frequently Asked Questions

### What is the difference between transcribe and translate in Whisper?

**Transcribe** outputs text in the same language as the spoken audio. **Translate** converts the speech content into a different target language using Whisper's built-in cross-lingual capabilities. The distinction is handled entirely by the `task` parameter passed to the model's generation method.

### Does the translate task require specifying a target language?

No. If you omit the `--language` parameter, Whisper defaults to translating into English. However, you can specify any supported language code (e.g., `es`, `fr`, `de`) to translate into that specific language instead.

### Why does the task parameter get removed for .en models?

English-only model checkpoints lack the multilingual training necessary for translation. As implemented in [`cli.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/cli.py) lines 49-51, the code detects the `.en` suffix and pops the task key from `generate_kwargs` to prevent the pipeline from attempting unsupported operations that would raise errors.

### What output format does insanely-fast-whisper produce?

The tool generates JSON files containing both full text and timestamped chunks. According to [`src/insanely_fast_whisper/utils/result.py`](https://github.com/Vaibhavs10/insanely-fast-whisper/blob/main/src/insanely_fast_whisper/utils/result.py), the output includes `outputs["chunks"]` and `outputs["text"]` from the pipeline, optionally combined with speaker diarization data if a Hugging Face token is provided.