VoiceStudio Video Dubbing Pipeline: 10 Stages from Ingest to Export
The VoiceStudio video dubbing pipeline consists of 10 distinct stages—from media ingestion and speaker diarization to final video composition— orchestrated by the dub_video coroutine in the SoniTranslate sidecar.
VoiceStudio (from the debpalash/VoiceStudio repository) implements a full-stack video dubbing workflow that runs entirely inside the SoniTranslate sidecar. The VoiceStudio video dubbing pipeline processes source media through ten specialized backend stages to produce lip-synced, translated videos with cloned voices and soft subtitles.
The 10 Stages of the Dubbing Pipeline
The pipeline follows a strict sequential flow, with each stage handled by a dedicated service in the backend/services/ directory.
1. Server Initialization
Before processing begins, the system ensures the Gradio server is active. If the server is not running, sonitranslate.start() launches it as a subprocess. This initialization logic resides in backend/services/sonitranslate.py at lines 42-48.
2. Media Ingestion
The source video file or URL is uploaded to the sidecar via the batch_multilingual_media_conversion endpoint. The dub_video() function submits the job to the Gradio interface, as implemented in backend/services/sonitranslate.py (lines 44-53).
3. Audio Extraction and Diarization
FFmpeg extracts the audio track from the source video, then a diarization model (pyannote) separates individual speakers using segment-after-each-speaker logic. This creates distinct audio chunks for each speaker, handled by backend/services/segmentation.py.
4. Transcription
The extracted audio segments are sent to the transcription backend for speech-to-text conversion. Using a Whisper-style model, this stage generates the source script and is implemented in backend/services/asr_backend.py.
5. Translation
Transcribed text passes through the translation service, which applies dubbing-specific prompts to ensure natural timing and phrasing for spoken dialogue. The logic resides in backend/services/translator.py.
6. Speaker-Clone Pre-processing
If voice cloning is requested, the system analyzes the source voice and prepares a clone model using RVC or FreeVC. This stage is managed by backend/services/speaker_clone.py.
7. TTS Generation
The translated script feeds into the TTS engine (supporting Edge TTS, Piper, or XTTS). A duration planner ensures the synthesized speech matches the original timing for proper lip-sync. This critical stage is found in backend/services/tts_backend.py at lines 2959-2975.
8. Audio Mixing
Synthesized voice tracks are mixed with the original audio, applying volume balancing and optional vocal refinement. The mixing service is located in backend/services/audio_mix.py.
9. Subtitle Generation
The pipeline generates SRT or VTT subtitle files, applying soft-subtitle rules for timing and line-length constraints. This stage is implemented in backend/services/subtitle_generator.py.
10. Final Composition
Using FFmpeg, the mixed audio and subtitles are muxed back onto the original video container. The video export service completes the workflow in backend/services/video_export.py.
Pipeline Orchestration
Each stage is coordinated by the dub_video coroutine defined in backend/services/sonitranslate.py. This function sends required parameters to the Gradio endpoint, validates the returned file name upon completion (lines 24-33), copies the result to the user-specified output directory, and returns a concise result dictionary.
Running the Complete Pipeline
To execute a full dubbing job, ensure the sidecar is running and invoke dub_video() with your parameters:
import asyncio
from backend.services.sonitranslate import dub_video, start, stop
async def run():
# Ensure the side‑car is up
await start()
# Perform dubbing
result = await dub_video(
video_path="~/my_video.mp4",
target_language="Japanese (ja)",
source_language="Automatic detection",
tts_voice="ja-JP-NaokiNeural-Male",
max_speakers=2,
output_dir="~/dubbed_videos",
)
print("Dubbed file:", result["output_file"])
# Shut down the side‑car when done
await stop()
asyncio.run(run())
Invoking Individual Stages
You can also interact with specific services directly. For example, to translate text using the dubbing-specific translator:
from backend.services.translator import translate_text
translated = translate_text(
source_text="Hello, world!",
source_lang="en",
target_lang="fr",
style="professional dubbing translator",
)
print(translated)
Summary
- The VoiceStudio video dubbing pipeline runs inside the SoniTranslate sidecar and consists of 10 distinct processing stages.
- Pipeline orchestration is handled by the
dub_video()coroutine inbackend/services/sonitranslate.py, which manages the Gradio endpoint communication. - Key services include diarization (
segmentation.py), transcription (asr_backend.py), translation (translator.py), voice cloning (speaker_clone.py), and TTS generation (tts_backend.py). - The duration planner in the TTS stage ensures synthesized speech matches original timing for lip-sync.
- Final output is produced by
video_export.py, which muxes audio, video, and subtitles using FFmpeg.
Frequently Asked Questions
What is the SoniTranslate sidecar in VoiceStudio?
The SoniTranslate sidecar is a Gradio-based service that hosts the complete dubbing pipeline. VoiceStudio wraps this service, using sonitranslate.start() to launch the subprocess and dub_video() to submit jobs. It handles all heavy lifting including transcription, translation, and synthesis.
How does VoiceStudio handle multiple speakers in a video?
The pipeline uses speaker diarization via the pyannote model in backend/services/segmentation.py. This stage segments the audio after extracting it with FFmpeg, creating distinct chunks for each speaker that are processed separately through transcription and voice cloning.
Which text-to-speech engines does VoiceStudio support?
According to backend/services/tts_backend.py, VoiceStudio supports multiple TTS engines including Edge TTS, Piper, and XTTS, selectable based on quality requirements and resource constraints. The system applies a duration planner (lines 2959-2975) to maintain lip-sync regardless of the engine chosen.
How does the pipeline maintain lip-sync during dubbing?
Lip-sync is preserved through a duration planner integrated into the TTS generation stage (tts_backend.py). This component adjusts the synthesized speech timing to match the original audio duration, ensuring that the dubbed voice aligns with the speaker's mouth movements in the final video composition.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →