# Main Dependencies for the VoiceStudio Backend: Complete Dependency Guide

> Discover the core dependencies for the VoiceStudio backend, including PyTorch, FastAPI, WhisperX, and KittentTS. Explore the complete dependency guide for seamless integration.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: dependency-guide
- Published: 2026-09-11

---

**The VoiceStudio backend is built on PyTorch 2.4+, FastAPI, and specialized audio libraries like WhisperX and KittentTS, with all dependencies strictly pinned in [`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml).**

VoiceStudio is an open-source Python voice synthesis platform that combines deep learning models, real-time audio processing, and a web-based API. The **main dependencies for the VoiceStudio backend** are declared centrally in the project's [`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml) file, which governs everything from GPU-accelerated inference to WebSocket communication.

## Deep Learning and Model Infrastructure

The foundation of VoiceStudio's AI capabilities rests on the PyTorch ecosystem and HuggingFace tools. According to the source code in [`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml), the core machine learning stack includes:

- **`torch`** ≥ 2.4 — The primary deep learning framework
- **`torchaudio`** ≥ 2.4 — Audio-specific PyTorch operations
- **`torchvision`** ≥ 0.19 — Computer vision utilities for multimodal processing
- **`transformers`** ≥ 5.10.0 — Model loading and inference for TTS and NLP tasks
- **`accelerate`** — Hardware optimization for distributed training and inference

These packages enable the backend to load and execute large voice models. In [`services/tts.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/tts.py), the backend uses Transformers to initialize text-to-speech pipelines:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "OmniStudio/omnivoice-tts"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto")

```

## Audio Processing and Generation

VoiceStudio handles complex audio manipulation through a suite of specialized libraries that process raw waveforms, apply effects, and separate audio sources:

- **`pydub`** — High-level audio manipulation and format conversion
- **`numpy`** — Numerical operations on audio tensors
- **`soundfile`** — Reading and writing audio files
- **`pedalboard`** ≥ 0.9.14 — Audio effects and signal processing
- **`demucs`** ≥ 4.0.1 — Source separation for isolating vocals

The backend leverages `pydub` in processing pipelines to concatenate audio segments. This pattern appears in utilities that stitch together generated voice chunks:

```python
from pydub import AudioSegment

def concat_wavs(paths: list[str]) -> AudioSegment:
    combined = AudioSegment.empty()
    for p in paths:
        combined += AudioSegment.from_wav(p)
    return combined

```

## Speech-to-Text and Alignment

For automatic speech recognition (ASR) and speaker diarization, VoiceStudio integrates multiple backend options with different performance characteristics:

- **`whisperx`** ≥ 3.1.0 — Fast transcription with word-level alignment
- **`faster-whisper`** ≥ 1.0.0 — Optimized OpenAI Whisper implementation
- **`mlx-whisper`** ≥ 0.2.1 — Apple Silicon optimized ASR (macOS ARM only)
- **`pyannote-audio`** ≥ 3.3.2 — Speaker identification and segmentation

The [`services/asr.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/asr.py) module wraps these engines. A typical implementation loads WhisperX for transcription tasks:

```python
import whisperx

model = whisperx.load_model("large-v2")
audio = whisperx.load_audio("sample.wav")
result = model.transcribe(audio)

```

## Text-to-Speech (TTS) Engines

VoiceStudio supports multiple TTS backends depending on the deployment platform, with specific optimizations for Apple Silicon:

- **`kittentts`** — Primary TTS engine (installed from external URL)
- **`mlx-audio`** ≥ 0.3.0 — Optimized audio generation for macOS ARM
- **`parakeet-mlx`** ≥ 0.5.2 — Additional macOS ARM voice synthesis support

Platform-specific entries like `mlx-audio` include environment markers in [`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml), ensuring they install only on macOS ARM devices.

## Web API and Server Infrastructure

The backend exposes functionality through a FastAPI application defined in [`backend/spec.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/spec.py). The web stack dependencies include:

- **`fastapi`** < 0.137 — The ASGI web framework for REST endpoints
- **`scalar-fastapi`** — OpenAPI documentation UI
- **`uvicorn`** — ASGI server implementation
- **`python-multipart`** ≥ 0.0.31 — Handling file uploads
- **`websockets`** — Real-time WebSocket communication

The entry point in [`backend/spec.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/spec.py) initializes the FastAPI application with WebSocket support:

```python
from fastapi import FastAPI, UploadFile, WebSocket
import uvicorn

app = FastAPI()

@app.get("/health")
def health() -> dict:
    return {"status": "ok"}

if __name__ == "__main__":
    uvicorn.run(app, host="0.0.0.0", port=8000)

```

## Frontend Interface and UI

While primarily a backend concern, the dependency tree includes **`gradio`** ≥ 6.15.1 for the web interface. The [`frontend/src/app.jsx`](https://github.com/debpalash/VoiceStudio/blob/main/frontend/src/app.jsx) module uses Gradio to create interactive voice synthesis interfaces that communicate with the FastAPI backend.

## Data Handling, Security, and Utilities

Supporting infrastructure includes tools for content ingestion, security, and system monitoring:

- **`yt-dlp`** ≥ 2026.7.4 — Downloading audio from video platforms
- **`pypdf`** ≥ 4.0 — Extracting text from PDF documents
- **`cryptography`** ≥ 41 — Encryption and secure token handling
- **`psutil`** ≥ 7.2.2 — System resource monitoring
- **`tensorboardX`** — Training visualization and logging
- **`alembic`** ≥ 1.16 — Database migration management

## Installation Source File

All dependency specifications originate from **[`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml)** at the repository root. Key version constraints include:
- `torch>=2.4` (line 36)
- `fastapi<0.137` (line 111)
- `gradio>=6.15.1` (line 43)

The file also pins development tools and optional analytics packages like `posthog` ≥ 3.7 for usage telemetry.

## Summary

- The **VoiceStudio backend dependencies** are centralized in [`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml) with strict semantic versioning
- **PyTorch 2.4+** and **Transformers 5.10+** provide the deep learning foundation
- **Audio processing** combines `pydub` for manipulation, `whisperx` for transcription, and `kittentts` for synthesis
- **FastAPI** and **Uvicorn** power the REST API and WebSocket endpoints defined in [`backend/spec.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/spec.py)
- Platform-specific packages like `mlx-whisper` and `mlx-audio` optimize performance for Apple Silicon Macs

## Frequently Asked Questions

### What file contains the VoiceStudio backend dependencies?

All production dependencies are declared in **[`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml)** at the repository root. This file pins exact versions for the PyTorch ecosystem, FastAPI web stack, and audio processing libraries referenced throughout the `backend/` and `services/` directories.

### Which deep learning frameworks does VoiceStudio use?

The backend relies on **PyTorch** ≥ 2.4 as its primary tensor computation engine, supplemented by **HuggingFace Transformers** ≥ 5.10.0 for model architecture implementations and **Accelerate** for hardware optimization. These support both CPU and GPU inference pipelines.

### Are there platform-specific dependencies for VoiceStudio?

Yes. The [`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml) includes environment markers for Apple Silicon (macOS ARM) that install **`mlx-whisper`**, **`mlx-audio`**, and **`parakeet-mlx`** only on compatible devices. These packages leverage Apple's MLX framework for accelerated inference on M1/M2/M3 chips.

### How does VoiceStudio handle speech-to-text processing?

The backend supports multiple ASR backends through **[`services/asr.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/asr.py)**, including **WhisperX** for alignment-accurate transcription, **faster-whisper** for speed-optimized processing, and **sherpa-onnx** for offline recognition. The system defaults to WhisperX ≥ 3.1.0 for most production transcoding tasks.