How to Start a Video Dubbing Job Using the VoiceStudio API

You can start a video dubbing job by sending a POST request to /api/dub/generate with the video file and target languages, optionally specifying a voice ID and background audio preservation.

VoiceStudio is an open-source dubbing pipeline that automates video localization through a RESTful API. To initiate the full dubbing workflow—including transcription, translation, speech synthesis, and audio mixing—you interact with the generate endpoint defined in backend/api/routers/dub_core.py. This guide explains the exact API contract and provides ready-to-run examples for submitting jobs and tracking their progress.

The Dubbing Endpoint and Request Format

The primary entry point for creating dubbing jobs is POST /api/dub/generate, implemented in backend/api/routers/dub_core.py at lines 44-66. This endpoint accepts multipart/form-data requests and immediately queues the video for processing through the internal pipeline.

Required Parameters

Every request must include:

  • video – The source video file uploaded as binary data.
  • langs – A comma-separated string or JSON list of target language codes (e.g., "en,es" or ["en", "es"]).

Optional Parameters

You can fine-tune the output with these optional fields:

  • voice_id – A string identifier for a cloned voice to use instead of the default.
  • preserve_bg – Boolean indicating whether to retain the original background audio; defaults to true if omitted.

Submitting a Dubbing Request

Below are production-ready examples for calling the VoiceStudio API.

Using cURL

curl -X POST "http://localhost:8000/api/dub/generate" \
     -F "video=@/path/to/movie.mp4" \
     -F "langs=en,es" \
     -F "voice_id=12345" \
     -F "preserve_bg=true"

Using Python with httpx

import httpx

url = "http://localhost:8000/api/dub/generate"
files = {"video": open("movie.mp4", "rb")}
data = {
    "langs": ["en", "es"],
    "voice_id": "12345",
    "preserve_bg": "true",
}

resp = httpx.post(url, files=files, data=data)
print(resp.json())

Monitoring Job Status

After submission, the API returns a job identifier that you can use to poll for progress. The status model BatchJobStatus defined in backend/api/routers/batch.py (lines 36-48) provides the current state and completion percentage.

Polling for Completion

import time
import httpx

job_id = resp.json()["id"]
status_url = f"http://localhost:8000/api/dub/status/{job_id}"

while True:
    st = httpx.get(status_url).json()
    print(f"{st['status']} – {st.get('progress')}")
    if st["status"] in ("done", "failed", "cancelled"):
        break
    time.sleep(2)

How the Pipeline Processes Your Video

When you start a video dubbing job using the VoiceStudio API, the request delegates to services/dub_pipeline.py. This internal orchestrator performs the following sequence:

  1. Audio extraction from the uploaded video.
  2. ASR (Automatic Speech Recognition) to generate transcripts.
  3. Translation via backend/services/translator.py using professional dubbing prompts.
  4. TTS generation for each translated segment, optionally using the specified voice_id.
  5. Alignment of new speech to the original timing.
  6. Mixing of the generated voice track with the background audio (if preserve_bg is true).
  7. Export to the job's output directory.

The backend/services/video_context.py module analyzes visual context to optimize dubbing parameters during this process.

Summary

  • Submit jobs to POST /api/dub/generate with video and langs parameters.
  • Optionally specify voice_id for custom voices and preserve_bg to control background audio.
  • Poll GET /api/dub/status/{job_id} to track progress using the BatchJobStatus model.
  • The pipeline handles transcription, translation, synthesis, and mixing automatically via services/dub_pipeline.py.

Frequently Asked Questions

What endpoint do I use to start a video dubbing job?

Send a POST request to /api/dub/generate as defined in backend/api/routers/dub_core.py. This endpoint accepts the video file and target language specifications, then returns a job ID for tracking.

How do I check the status of a dubbing job?

Query GET /api/dub/status/{job_id} to retrieve the current state. The response follows the BatchJobStatus schema from backend/api/routers/batch.py, indicating whether the job is pending, processing, done, failed, or cancelled.

Can I preserve the original background audio while dubbing?

Yes. Include the preserve_bg parameter set to true (the default) in your request to POST /api/dub/generate. When enabled, the pipeline mixes the new synthesized speech over the original background track rather than replacing it entirely.

What parameters are required to start a dubbing job?

You must provide the video file and langs list. The voice_id and preserve_bg fields are optional. The langs parameter accepts either a comma-separated string or a JSON array of ISO language codes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →