Main Dependencies for the VoiceStudio Backend: Complete Dependency Guide
The VoiceStudio backend is built on PyTorch 2.4+, FastAPI, and specialized audio libraries like WhisperX and KittentTS, with all dependencies strictly pinned in pyproject.toml.
VoiceStudio is an open-source Python voice synthesis platform that combines deep learning models, real-time audio processing, and a web-based API. The main dependencies for the VoiceStudio backend are declared centrally in the project's pyproject.toml file, which governs everything from GPU-accelerated inference to WebSocket communication.
Deep Learning and Model Infrastructure
The foundation of VoiceStudio's AI capabilities rests on the PyTorch ecosystem and HuggingFace tools. According to the source code in pyproject.toml, the core machine learning stack includes:
torch≥ 2.4 — The primary deep learning frameworktorchaudio≥ 2.4 — Audio-specific PyTorch operationstorchvision≥ 0.19 — Computer vision utilities for multimodal processingtransformers≥ 5.10.0 — Model loading and inference for TTS and NLP tasksaccelerate— Hardware optimization for distributed training and inference
These packages enable the backend to load and execute large voice models. In services/tts.py, the backend uses Transformers to initialize text-to-speech pipelines:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "OmniStudio/omnivoice-tts"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto")
Audio Processing and Generation
VoiceStudio handles complex audio manipulation through a suite of specialized libraries that process raw waveforms, apply effects, and separate audio sources:
pydub— High-level audio manipulation and format conversionnumpy— Numerical operations on audio tensorssoundfile— Reading and writing audio filespedalboard≥ 0.9.14 — Audio effects and signal processingdemucs≥ 4.0.1 — Source separation for isolating vocals
The backend leverages pydub in processing pipelines to concatenate audio segments. This pattern appears in utilities that stitch together generated voice chunks:
from pydub import AudioSegment
def concat_wavs(paths: list[str]) -> AudioSegment:
combined = AudioSegment.empty()
for p in paths:
combined += AudioSegment.from_wav(p)
return combined
Speech-to-Text and Alignment
For automatic speech recognition (ASR) and speaker diarization, VoiceStudio integrates multiple backend options with different performance characteristics:
whisperx≥ 3.1.0 — Fast transcription with word-level alignmentfaster-whisper≥ 1.0.0 — Optimized OpenAI Whisper implementationmlx-whisper≥ 0.2.1 — Apple Silicon optimized ASR (macOS ARM only)pyannote-audio≥ 3.3.2 — Speaker identification and segmentation
The services/asr.py module wraps these engines. A typical implementation loads WhisperX for transcription tasks:
import whisperx
model = whisperx.load_model("large-v2")
audio = whisperx.load_audio("sample.wav")
result = model.transcribe(audio)
Text-to-Speech (TTS) Engines
VoiceStudio supports multiple TTS backends depending on the deployment platform, with specific optimizations for Apple Silicon:
kittentts— Primary TTS engine (installed from external URL)mlx-audio≥ 0.3.0 — Optimized audio generation for macOS ARMparakeet-mlx≥ 0.5.2 — Additional macOS ARM voice synthesis support
Platform-specific entries like mlx-audio include environment markers in pyproject.toml, ensuring they install only on macOS ARM devices.
Web API and Server Infrastructure
The backend exposes functionality through a FastAPI application defined in backend/spec.py. The web stack dependencies include:
fastapi< 0.137 — The ASGI web framework for REST endpointsscalar-fastapi— OpenAPI documentation UIuvicorn— ASGI server implementationpython-multipart≥ 0.0.31 — Handling file uploadswebsockets— Real-time WebSocket communication
The entry point in backend/spec.py initializes the FastAPI application with WebSocket support:
from fastapi import FastAPI, UploadFile, WebSocket
import uvicorn
app = FastAPI()
@app.get("/health")
def health() -> dict:
return {"status": "ok"}
if __name__ == "__main__":
uvicorn.run(app, host="0.0.0.0", port=8000)
Frontend Interface and UI
While primarily a backend concern, the dependency tree includes gradio ≥ 6.15.1 for the web interface. The frontend/src/app.jsx module uses Gradio to create interactive voice synthesis interfaces that communicate with the FastAPI backend.
Data Handling, Security, and Utilities
Supporting infrastructure includes tools for content ingestion, security, and system monitoring:
yt-dlp≥ 2026.7.4 — Downloading audio from video platformspypdf≥ 4.0 — Extracting text from PDF documentscryptography≥ 41 — Encryption and secure token handlingpsutil≥ 7.2.2 — System resource monitoringtensorboardX— Training visualization and loggingalembic≥ 1.16 — Database migration management
Installation Source File
All dependency specifications originate from pyproject.toml at the repository root. Key version constraints include:
torch>=2.4(line 36)fastapi<0.137(line 111)gradio>=6.15.1(line 43)
The file also pins development tools and optional analytics packages like posthog ≥ 3.7 for usage telemetry.
Summary
- The VoiceStudio backend dependencies are centralized in
pyproject.tomlwith strict semantic versioning - PyTorch 2.4+ and Transformers 5.10+ provide the deep learning foundation
- Audio processing combines
pydubfor manipulation,whisperxfor transcription, andkittenttsfor synthesis - FastAPI and Uvicorn power the REST API and WebSocket endpoints defined in
backend/spec.py - Platform-specific packages like
mlx-whisperandmlx-audiooptimize performance for Apple Silicon Macs
Frequently Asked Questions
What file contains the VoiceStudio backend dependencies?
All production dependencies are declared in pyproject.toml at the repository root. This file pins exact versions for the PyTorch ecosystem, FastAPI web stack, and audio processing libraries referenced throughout the backend/ and services/ directories.
Which deep learning frameworks does VoiceStudio use?
The backend relies on PyTorch ≥ 2.4 as its primary tensor computation engine, supplemented by HuggingFace Transformers ≥ 5.10.0 for model architecture implementations and Accelerate for hardware optimization. These support both CPU and GPU inference pipelines.
Are there platform-specific dependencies for VoiceStudio?
Yes. The pyproject.toml includes environment markers for Apple Silicon (macOS ARM) that install mlx-whisper, mlx-audio, and parakeet-mlx only on compatible devices. These packages leverage Apple's MLX framework for accelerated inference on M1/M2/M3 chips.
How does VoiceStudio handle speech-to-text processing?
The backend supports multiple ASR backends through services/asr.py, including WhisperX for alignment-accurate transcription, faster-whisper for speed-optimized processing, and sherpa-onnx for offline recognition. The system defaults to WhisperX ≥ 3.1.0 for most production transcoding tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →