# Voice-Pro Audio Processing Tools: FFmpeg, Demucs, and the Machine Learning Pipeline

> Voice-Pro uses FFmpeg and Demucs for advanced audio processing including source separation and format conversion. Discover the machine learning pipeline behind transcription and dubbing.

- Repository: [ABUS/voice-pro](https://github.com/abus-aikorea/voice-pro)
- Tags: deep-dive
- Published: 2026-08-03

---

**Voice-Pro leverages FFmpeg, Demucs, Librosa, Torchaudio, and Pydub to execute format conversion, source separation, deep-learning preprocessing, and waveform manipulation across its transcription and dubbing pipeline.**

The open-source Voice-Pro repository by abus-aikorea integrates industry-standard audio processing libraries to handle complex media workflows from download to final render. These tools form a cohesive pipeline that supports everything from command-line media conversion to neural network-ready tensor preparation.

## FFmpeg: The Core Media Engine

**FFmpeg** serves as the primary command-line engine for media manipulation in Voice-Pro. Implemented in [`app/abus_ffmpeg.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_ffmpeg.py), it handles extraction, conversion, compression, and metadata probing through both the `ffmpeg-python` wrapper and direct CLI calls.

Key functions include `ffmpeg_extract_audio` for ripping audio tracks from video containers, `ffmpeg_replace_audio` for overdubbing operations, and `ffmpeg_compress_video` for output optimization. The module also provides utilities for channel conversion (`ffmpeg_to_mono`, `ffmpeg_to_stereo`), volume adjustment (`ffmpeg_volume_control`), and precise trimming (`ffmpeg_trim_seconds`).

### Media Extraction and Probing

The FFmpeg integration probes codec information, resolution, FPS, and duration before processing. This ensures compatible audio extraction into formats like WAV or AAC for downstream machine learning models.

### Audio Replacement and Mixing

For dubbing workflows, `ffmpeg_replace_audio` synchronizes new voice tracks with existing video streams, handling format standardization and codec compatibility automatically.

## Deep Learning Preprocessing Libraries

Voice-Pro employs **Torchaudio** and **Librosa** for preparing audio data compatible with neural network inference.

### Torchaudio for Tensor I/O

Located primarily in [`app/abus_tts_cosyvoice.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_cosyvoice.py), Torchaudio provides tensor-based audio I/O essential for PyTorch models. The `torchaudio.save` function exports generated speech waveforms, while loading utilities convert raw audio into tensors for GPU processing.

### Librosa for Spectral Analysis

Referenced in [`app/abus_tts_cosyvoice.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_cosyvoice.py) and [`app/abus_aicover.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_aicover.py), Librosa offers advanced analysis capabilities including silence trimming via `librosa.effects.trim`, pitch-shifting, and multi-channel waveform loading as NumPy arrays. This enables precise audio cleaning before TTS inference.

## AI-Powered Source Separation with Demucs

For karaoke and instrumental generation features, Voice-Pro integrates **Demucs** (MDX-Net) through [`app/abus_demucs.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_demucs.py). The `demucs_split_file` function performs deep-learning-based source separation, isolating vocals from instrumentals.

Called from controller modules like [`gradio_gulliver.py`](https://github.com/abus-aikorea/voice-pro/blob/main/gradio_gulliver.py) and [`gradio_kara.py`](https://github.com/abus-aikorea/voice-pro/blob/main/gradio_kara.py), Demucs supports multiple model variants including `htdemucs` and outputs split stems in formats such as WAV or MP3.

## High-Level Audio Manipulation via Pydub

**Pydub** provides a Pythonic interface for simple audio operations without direct FFmpeg complexity. Used in [`app/abus_audio.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_audio.py) and various TTS modules including [`app/abus_tts_edge.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_edge.py), Pydub handles quick segment slicing, silence detection, and decibel-based volume adjustments.

## Practical Implementation Examples

The following patterns demonstrate how Voice-Pro orchestrates these tools in production:

```python

# Extract WAV audio from video using FFmpeg

from app.abus_ffmpeg import ffmpeg_extract_audio
audio_path = ffmpeg_extract_audio("input_video.mp4", "output.wav")

```

```python

# Replace video audio track with new voiceover

from app.abus_ffmpeg import ffmpeg_replace_audio
ffmpeg_replace_audio("input_video.mp4", "new_voice.aac", "output_dubbed.mp4")

```

```python

# Separate vocals and instruments using Demucs

from app.abus_demucs import demucs_split_file
instrumental, vocal = demucs_split_file(
    input_path="song.mp3",
    output_dir="demucs_output",
    demucs_model="htdemucs",
    audio_format="wav"
)

```

```python

# Preprocess audio for ML: trim silence with Librosa, save with Torchaudio

import librosa
import torch
import torchaudio

y, sr = librosa.load("raw_input.wav", sr=None)
y_trim, _ = librosa.effects.trim(y)
torchaudio.save("trimmed.wav", torch.tensor(y_trim).unsqueeze(0), sr)

```

```python

# Volume adjustment using Pydub

from pydub import AudioSegment
segment = AudioSegment.from_file("speech.wav")
louder = segment + 6  # Boost by 6dB

louder.export("speech_louder.wav", format="wav")

```

## Summary

- **FFmpeg** in [`app/abus_ffmpeg.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_ffmpeg.py) handles all command-line media conversion, extraction, and compression tasks including `ffmpeg_extract_audio` and `ffmpeg_replace_audio`.
- **Demucs** via [`app/abusc_demucs.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abusc_demucs.py) provides deep-learning source separation for vocal/instrumental isolation using `demucs_split_file`.
- **Torchaudio** and **Librosa** in TTS modules like [`app/abus_tts_cosyvoice.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_cosyvoice.py) prepare tensor-ready audio for neural networks with functions like `torchaudio.save` and `librosa.effects.trim`.
- **Pydub** in [`app/abus_audio.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_audio.py) enables rapid high-level editing such as slicing and volume control without complex CLI calls.
- These tools collectively support Voice-Pro's cross-platform transcription, translation, and dubbing workflows across Windows and Linux environments.

## Frequently Asked Questions

### What audio formats does Voice-Pro support for input and output?

Voice-Pro supports virtually all audio and video formats through its FFmpeg integration in [`app/abus_ffmpeg.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_ffmpeg.py), including MP3, WAV, AAC, FLAC, and OGG. The `ffmpeg_extract_audio` function automatically handles codec detection and conversion to model-compatible formats like 16-bit WAV.

### How does Voice-Pro separate vocals from background music?

The repository implements **Demucs** (MDX-Net) through [`app/abus_demucs.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_demucs.py), specifically the `demucs_split_file` function. This deep-learning model isolates vocal and instrumental stems, enabling karaoke features and clean voice extraction for dubbing projects.

### Which libraries handle audio preprocessing for Voice-Pro's AI models?

**Librosa** and **Torchaudio** manage preprocessing in modules like [`app/abus_tts_cosyvoice.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_cosyvoice.py) and [`app/abus_aicover.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_aicover.py). Librosa performs silence trimming and spectral analysis, while Torchaudio converts waveforms to PyTorch tensors and saves generated speech outputs to disk.

### Can Voice-Pro adjust audio volume programmatically without FFmpeg?

Yes, Voice-Pro uses **Pydub** in [`app/abus_audio.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_audio.py) for simple decibel-based adjustments. This provides a lighter alternative to FFmpeg for basic volume boosting, fading, and segment concatenation tasks before feeding data to TTS engines like those in [`app/abus_tts_edge.py`](https://github.com/abus-aikorea/voice-pro/blob/main/app/abus_tts_edge.py).