Using the Translate Task Instead of Transcribe in Insanely-Fast-Whisper: A Complete Guide

Use the --task translate CLI flag to generate translated text instead of raw transcripts, leveraging Whisper's built-in translation capabilities while maintaining the same JSON output format.

Insanely-Fast-Whisper accelerates OpenAI's Whisper model through Hugging Face's transformers pipeline. Understanding how to switch between the transcribe and translate tasks allows you to process multilingual audio without external translation tools. This guide examines the actual CLI implementation to show you exactly how the translation pipeline works.

How the --task Parameter Controls Generation

The command-line interface defines the --task argument in src/insanely_fast_whisper/cli.py (lines 38-44) with strict validation:

parser.add_argument(
    "--task",
    default="transcribe",
    choices=["transcribe", "translate"],
    help="Task to perform: transcribe or translate.",
)

transcribe (the default) returns text in the original spoken language. translate converts the speech into the target language specified by the --language parameter, or English if omitted.

The generate_kwargs Mechanism

When you execute the CLI, the script constructs a generate_kwargs dictionary at lines 47-51 in cli.py that passes your selected task directly to the Hugging Face pipeline:

generate_kwargs = {"task": args.task, "language": language}
if args.model_name.split(".")[-1] == "en":
    generate_kwargs.pop("task")

This dictionary is injected into the pipeline call at lines 59-64, instructing Whisper's generation method to either transcribe verbatim or translate the audio content. The rest of the processing pipeline—including batching, chunking, and optional speaker diarization—remains identical regardless of which task you select.

Automatic Handling of English-Only Models

The source code contains specific logic for English-only checkpoints. If your model name ends with .en (e.g., openai/whisper-small.en), the code automatically removes the task key from generate_kwargs:

if args.model_name.split(".")[-1] == "en":
    generate_kwargs.pop("task")

English-only models cannot perform translation because they lack multilingual weights. Attempting to force a translate task on these checkpoints would cause runtime errors, so the wrapper safely defaults to transcription-only mode.

Practical CLI Examples

The following commands demonstrate real-world usage patterns. Each produces a JSON output file formatted by build_result in src/insanely_fast_whisper/utils/result.py (lines 10-15), which preserves the standard chunk and text structure regardless of task type.

Standard Transcription

python -m insanely_fast_whisper.cli \
    --file-name audio.wav \
    --model-name openai/whisper-large-v3 \
    --task transcribe \
    --output-path transcript.json

Translate to English (Auto-detect Source)

python -m insanely_fast_whisper.cli \
    --file-name french_audio.wav \
    --model-name openai/whisper-large-v3 \
    --task translate \
    --language en \
    --output-path translation.json

Translate to Spanish

python -m insanely_fast_whisper.cli \
    --file-name english_audio.wav \
    --model-name openai/whisper-large-v3 \
    --task translate \
    --language es \
    --output-path spanish_translation.json

Summary

  • Task selection happens via --task transcribe or --task translate in the CLI, validated against a strict choices list in cli.py.
  • English-only models (names ending in .en) automatically have the task parameter stripped to prevent compatibility errors.
  • Output consistency is maintained through utils/result.py, ensuring JSON structure remains identical whether translating or transcribing.
  • Language specification uses the --language flag alongside --task translate to set the target tongue, or defaults to English if omitted.

Frequently Asked Questions

What is the difference between transcribe and translate in Whisper?

Transcribe outputs text in the same language as the spoken audio. Translate converts the speech content into a different target language using Whisper's built-in cross-lingual capabilities. The distinction is handled entirely by the task parameter passed to the model's generation method.

Does the translate task require specifying a target language?

No. If you omit the --language parameter, Whisper defaults to translating into English. However, you can specify any supported language code (e.g., es, fr, de) to translate into that specific language instead.

Why does the task parameter get removed for .en models?

English-only model checkpoints lack the multilingual training necessary for translation. As implemented in cli.py lines 49-51, the code detects the .en suffix and pops the task key from generate_kwargs to prevent the pipeline from attempting unsupported operations that would raise errors.

What output format does insanely-fast-whisper produce?

The tool generates JSON files containing both full text and timestamped chunks. According to src/insanely_fast_whisper/utils/result.py, the output includes outputs["chunks"] and outputs["text"] from the pipeline, optionally combined with speaker diarization data if a Hugging Face token is provided.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →