VoiceStudio Video Dubbing Pipeline: 10 Stages from Ingest to Export

The VoiceStudio video dubbing pipeline consists of 10 distinct stages—from media ingestion and speaker diarization to final video composition— orchestrated by the dub_video coroutine in the SoniTranslate sidecar.

VoiceStudio (from the debpalash/VoiceStudio repository) implements a full-stack video dubbing workflow that runs entirely inside the SoniTranslate sidecar. The VoiceStudio video dubbing pipeline processes source media through ten specialized backend stages to produce lip-synced, translated videos with cloned voices and soft subtitles.

The 10 Stages of the Dubbing Pipeline

The pipeline follows a strict sequential flow, with each stage handled by a dedicated service in the backend/services/ directory.

1. Server Initialization

Before processing begins, the system ensures the Gradio server is active. If the server is not running, sonitranslate.start() launches it as a subprocess. This initialization logic resides in backend/services/sonitranslate.py at lines 42-48.

2. Media Ingestion

The source video file or URL is uploaded to the sidecar via the batch_multilingual_media_conversion endpoint. The dub_video() function submits the job to the Gradio interface, as implemented in backend/services/sonitranslate.py (lines 44-53).

3. Audio Extraction and Diarization

FFmpeg extracts the audio track from the source video, then a diarization model (pyannote) separates individual speakers using segment-after-each-speaker logic. This creates distinct audio chunks for each speaker, handled by backend/services/segmentation.py.

4. Transcription

The extracted audio segments are sent to the transcription backend for speech-to-text conversion. Using a Whisper-style model, this stage generates the source script and is implemented in backend/services/asr_backend.py.

5. Translation

Transcribed text passes through the translation service, which applies dubbing-specific prompts to ensure natural timing and phrasing for spoken dialogue. The logic resides in backend/services/translator.py.

6. Speaker-Clone Pre-processing

If voice cloning is requested, the system analyzes the source voice and prepares a clone model using RVC or FreeVC. This stage is managed by backend/services/speaker_clone.py.

7. TTS Generation

The translated script feeds into the TTS engine (supporting Edge TTS, Piper, or XTTS). A duration planner ensures the synthesized speech matches the original timing for proper lip-sync. This critical stage is found in backend/services/tts_backend.py at lines 2959-2975.

8. Audio Mixing

Synthesized voice tracks are mixed with the original audio, applying volume balancing and optional vocal refinement. The mixing service is located in backend/services/audio_mix.py.

9. Subtitle Generation

The pipeline generates SRT or VTT subtitle files, applying soft-subtitle rules for timing and line-length constraints. This stage is implemented in backend/services/subtitle_generator.py.

10. Final Composition

Using FFmpeg, the mixed audio and subtitles are muxed back onto the original video container. The video export service completes the workflow in backend/services/video_export.py.

Pipeline Orchestration

Each stage is coordinated by the dub_video coroutine defined in backend/services/sonitranslate.py. This function sends required parameters to the Gradio endpoint, validates the returned file name upon completion (lines 24-33), copies the result to the user-specified output directory, and returns a concise result dictionary.

Running the Complete Pipeline

To execute a full dubbing job, ensure the sidecar is running and invoke dub_video() with your parameters:

import asyncio
from backend.services.sonitranslate import dub_video, start, stop

async def run():
    # Ensure the side‑car is up

    await start()

    # Perform dubbing

    result = await dub_video(
        video_path="~/my_video.mp4",
        target_language="Japanese (ja)",
        source_language="Automatic detection",
        tts_voice="ja-JP-NaokiNeural-Male",
        max_speakers=2,
        output_dir="~/dubbed_videos",
    )
    print("Dubbed file:", result["output_file"])

    # Shut down the side‑car when done

    await stop()

asyncio.run(run())

Invoking Individual Stages

You can also interact with specific services directly. For example, to translate text using the dubbing-specific translator:

from backend.services.translator import translate_text

translated = translate_text(
    source_text="Hello, world!",
    source_lang="en",
    target_lang="fr",
    style="professional dubbing translator",
)
print(translated)

Summary

  • The VoiceStudio video dubbing pipeline runs inside the SoniTranslate sidecar and consists of 10 distinct processing stages.
  • Pipeline orchestration is handled by the dub_video() coroutine in backend/services/sonitranslate.py, which manages the Gradio endpoint communication.
  • Key services include diarization (segmentation.py), transcription (asr_backend.py), translation (translator.py), voice cloning (speaker_clone.py), and TTS generation (tts_backend.py).
  • The duration planner in the TTS stage ensures synthesized speech matches original timing for lip-sync.
  • Final output is produced by video_export.py, which muxes audio, video, and subtitles using FFmpeg.

Frequently Asked Questions

What is the SoniTranslate sidecar in VoiceStudio?

The SoniTranslate sidecar is a Gradio-based service that hosts the complete dubbing pipeline. VoiceStudio wraps this service, using sonitranslate.start() to launch the subprocess and dub_video() to submit jobs. It handles all heavy lifting including transcription, translation, and synthesis.

How does VoiceStudio handle multiple speakers in a video?

The pipeline uses speaker diarization via the pyannote model in backend/services/segmentation.py. This stage segments the audio after extracting it with FFmpeg, creating distinct chunks for each speaker that are processed separately through transcription and voice cloning.

Which text-to-speech engines does VoiceStudio support?

According to backend/services/tts_backend.py, VoiceStudio supports multiple TTS engines including Edge TTS, Piper, and XTTS, selectable based on quality requirements and resource constraints. The system applies a duration planner (lines 2959-2975) to maintain lip-sync regardless of the engine chosen.

How does the pipeline maintain lip-sync during dubbing?

Lip-sync is preserved through a duration planner integrated into the TTS generation stage (tts_backend.py). This component adjusts the synthesized speech timing to match the original audio duration, ensuring that the dubbed voice aligns with the speaker's mouth movements in the final video composition.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →