# VoiceStudio Video Dubbing Pipeline: 10 Stages from Ingest to Export

> Explore the 10 stages of the VoiceStudio video dubbing pipeline, from media ingest and diarization to final export. Understand how VoiceStudio streamlines your video dubbing workflow.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: architecture
- Published: 2026-09-09

---

**The VoiceStudio video dubbing pipeline consists of 10 distinct stages—from media ingestion and speaker diarization to final video composition— orchestrated by the `dub_video` coroutine in the SoniTranslate sidecar.**

VoiceStudio (from the `debpalash/VoiceStudio` repository) implements a full-stack video dubbing workflow that runs entirely inside the SoniTranslate sidecar. The **VoiceStudio video dubbing pipeline** processes source media through ten specialized backend stages to produce lip-synced, translated videos with cloned voices and soft subtitles.

## The 10 Stages of the Dubbing Pipeline

The pipeline follows a strict sequential flow, with each stage handled by a dedicated service in the `backend/services/` directory.

### 1. Server Initialization

Before processing begins, the system ensures the Gradio server is active. If the server is not running, `sonitranslate.start()` launches it as a subprocess. This initialization logic resides in [`backend/services/sonitranslate.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/sonitranslate.py) at lines 42-48.

### 2. Media Ingestion

The source video file or URL is uploaded to the sidecar via the `batch_multilingual_media_conversion` endpoint. The `dub_video()` function submits the job to the Gradio interface, as implemented in [`backend/services/sonitranslate.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/sonitranslate.py) (lines 44-53).

### 3. Audio Extraction and Diarization

FFmpeg extracts the audio track from the source video, then a diarization model (`pyannote`) separates individual speakers using segment-after-each-speaker logic. This creates distinct audio chunks for each speaker, handled by [`backend/services/segmentation.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/segmentation.py).

### 4. Transcription

The extracted audio segments are sent to the transcription backend for speech-to-text conversion. Using a Whisper-style model, this stage generates the source script and is implemented in [`backend/services/asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/asr_backend.py).

### 5. Translation

Transcribed text passes through the translation service, which applies dubbing-specific prompts to ensure natural timing and phrasing for spoken dialogue. The logic resides in [`backend/services/translator.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/translator.py).

### 6. Speaker-Clone Pre-processing

If voice cloning is requested, the system analyzes the source voice and prepares a clone model using **RVC** or **FreeVC**. This stage is managed by [`backend/services/speaker_clone.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/speaker_clone.py).

### 7. TTS Generation

The translated script feeds into the TTS engine (supporting **Edge TTS**, **Piper**, or **XTTS**). A **duration planner** ensures the synthesized speech matches the original timing for proper lip-sync. This critical stage is found in [`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py) at lines 2959-2975.

### 8. Audio Mixing

Synthesized voice tracks are mixed with the original audio, applying volume balancing and optional vocal refinement. The mixing service is located in [`backend/services/audio_mix.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/audio_mix.py).

### 9. Subtitle Generation

The pipeline generates SRT or VTT subtitle files, applying soft-subtitle rules for timing and line-length constraints. This stage is implemented in [`backend/services/subtitle_generator.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/subtitle_generator.py).

### 10. Final Composition

Using FFmpeg, the mixed audio and subtitles are muxed back onto the original video container. The video export service completes the workflow in [`backend/services/video_export.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/video_export.py).

## Pipeline Orchestration

Each stage is coordinated by the **`dub_video`** coroutine defined in [`backend/services/sonitranslate.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/sonitranslate.py). This function sends required parameters to the Gradio endpoint, validates the returned file name upon completion (lines 24-33), copies the result to the user-specified output directory, and returns a concise result dictionary.

## Running the Complete Pipeline

To execute a full dubbing job, ensure the sidecar is running and invoke `dub_video()` with your parameters:

```python
import asyncio
from backend.services.sonitranslate import dub_video, start, stop

async def run():
    # Ensure the side‑car is up

    await start()

    # Perform dubbing

    result = await dub_video(
        video_path="~/my_video.mp4",
        target_language="Japanese (ja)",
        source_language="Automatic detection",
        tts_voice="ja-JP-NaokiNeural-Male",
        max_speakers=2,
        output_dir="~/dubbed_videos",
    )
    print("Dubbed file:", result["output_file"])

    # Shut down the side‑car when done

    await stop()

asyncio.run(run())

```

## Invoking Individual Stages

You can also interact with specific services directly. For example, to translate text using the dubbing-specific translator:

```python
from backend.services.translator import translate_text

translated = translate_text(
    source_text="Hello, world!",
    source_lang="en",
    target_lang="fr",
    style="professional dubbing translator",
)
print(translated)

```

## Summary

- The **VoiceStudio video dubbing pipeline** runs inside the SoniTranslate sidecar and consists of 10 distinct processing stages.
- Pipeline orchestration is handled by the `dub_video()` coroutine in [`backend/services/sonitranslate.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/sonitranslate.py), which manages the Gradio endpoint communication.
- Key services include **diarization** ([`segmentation.py`](https://github.com/debpalash/VoiceStudio/blob/main/segmentation.py)), **transcription** ([`asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/asr_backend.py)), **translation** ([`translator.py`](https://github.com/debpalash/VoiceStudio/blob/main/translator.py)), **voice cloning** ([`speaker_clone.py`](https://github.com/debpalash/VoiceStudio/blob/main/speaker_clone.py)), and **TTS generation** ([`tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/tts_backend.py)).
- The **duration planner** in the TTS stage ensures synthesized speech matches original timing for lip-sync.
- Final output is produced by [`video_export.py`](https://github.com/debpalash/VoiceStudio/blob/main/video_export.py), which muxes audio, video, and subtitles using FFmpeg.

## Frequently Asked Questions

### What is the SoniTranslate sidecar in VoiceStudio?

The SoniTranslate sidecar is a Gradio-based service that hosts the complete dubbing pipeline. VoiceStudio wraps this service, using `sonitranslate.start()` to launch the subprocess and `dub_video()` to submit jobs. It handles all heavy lifting including transcription, translation, and synthesis.

### How does VoiceStudio handle multiple speakers in a video?

The pipeline uses **speaker diarization** via the `pyannote` model in [`backend/services/segmentation.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/segmentation.py). This stage segments the audio after extracting it with FFmpeg, creating distinct chunks for each speaker that are processed separately through transcription and voice cloning.

### Which text-to-speech engines does VoiceStudio support?

According to [`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py), VoiceStudio supports multiple TTS engines including **Edge TTS**, **Piper**, and **XTTS**, selectable based on quality requirements and resource constraints. The system applies a duration planner (lines 2959-2975) to maintain lip-sync regardless of the engine chosen.

### How does the pipeline maintain lip-sync during dubbing?

Lip-sync is preserved through a **duration planner** integrated into the TTS generation stage ([`tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/tts_backend.py)). This component adjusts the synthesized speech timing to match the original audio duration, ensuring that the dubbed voice aligns with the speaker's mouth movements in the final video composition.