Main Dependencies for the VoiceStudio Backend: Complete Dependency Guide

The VoiceStudio backend is built on PyTorch 2.4+, FastAPI, and specialized audio libraries like WhisperX and KittentTS, with all dependencies strictly pinned in pyproject.toml.

VoiceStudio is an open-source Python voice synthesis platform that combines deep learning models, real-time audio processing, and a web-based API. The main dependencies for the VoiceStudio backend are declared centrally in the project's pyproject.toml file, which governs everything from GPU-accelerated inference to WebSocket communication.

Deep Learning and Model Infrastructure

The foundation of VoiceStudio's AI capabilities rests on the PyTorch ecosystem and HuggingFace tools. According to the source code in pyproject.toml, the core machine learning stack includes:

  • torch ≥ 2.4 — The primary deep learning framework
  • torchaudio ≥ 2.4 — Audio-specific PyTorch operations
  • torchvision ≥ 0.19 — Computer vision utilities for multimodal processing
  • transformers ≥ 5.10.0 — Model loading and inference for TTS and NLP tasks
  • accelerate — Hardware optimization for distributed training and inference

These packages enable the backend to load and execute large voice models. In services/tts.py, the backend uses Transformers to initialize text-to-speech pipelines:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "OmniStudio/omnivoice-tts"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto")

Audio Processing and Generation

VoiceStudio handles complex audio manipulation through a suite of specialized libraries that process raw waveforms, apply effects, and separate audio sources:

  • pydub — High-level audio manipulation and format conversion
  • numpy — Numerical operations on audio tensors
  • soundfile — Reading and writing audio files
  • pedalboard ≥ 0.9.14 — Audio effects and signal processing
  • demucs ≥ 4.0.1 — Source separation for isolating vocals

The backend leverages pydub in processing pipelines to concatenate audio segments. This pattern appears in utilities that stitch together generated voice chunks:

from pydub import AudioSegment

def concat_wavs(paths: list[str]) -> AudioSegment:
    combined = AudioSegment.empty()
    for p in paths:
        combined += AudioSegment.from_wav(p)
    return combined

Speech-to-Text and Alignment

For automatic speech recognition (ASR) and speaker diarization, VoiceStudio integrates multiple backend options with different performance characteristics:

  • whisperx ≥ 3.1.0 — Fast transcription with word-level alignment
  • faster-whisper ≥ 1.0.0 — Optimized OpenAI Whisper implementation
  • mlx-whisper ≥ 0.2.1 — Apple Silicon optimized ASR (macOS ARM only)
  • pyannote-audio ≥ 3.3.2 — Speaker identification and segmentation

The services/asr.py module wraps these engines. A typical implementation loads WhisperX for transcription tasks:

import whisperx

model = whisperx.load_model("large-v2")
audio = whisperx.load_audio("sample.wav")
result = model.transcribe(audio)

Text-to-Speech (TTS) Engines

VoiceStudio supports multiple TTS backends depending on the deployment platform, with specific optimizations for Apple Silicon:

  • kittentts — Primary TTS engine (installed from external URL)
  • mlx-audio ≥ 0.3.0 — Optimized audio generation for macOS ARM
  • parakeet-mlx ≥ 0.5.2 — Additional macOS ARM voice synthesis support

Platform-specific entries like mlx-audio include environment markers in pyproject.toml, ensuring they install only on macOS ARM devices.

Web API and Server Infrastructure

The backend exposes functionality through a FastAPI application defined in backend/spec.py. The web stack dependencies include:

  • fastapi < 0.137 — The ASGI web framework for REST endpoints
  • scalar-fastapi — OpenAPI documentation UI
  • uvicorn — ASGI server implementation
  • python-multipart ≥ 0.0.31 — Handling file uploads
  • websockets — Real-time WebSocket communication

The entry point in backend/spec.py initializes the FastAPI application with WebSocket support:

from fastapi import FastAPI, UploadFile, WebSocket
import uvicorn

app = FastAPI()

@app.get("/health")
def health() -> dict:
    return {"status": "ok"}

if __name__ == "__main__":
    uvicorn.run(app, host="0.0.0.0", port=8000)

Frontend Interface and UI

While primarily a backend concern, the dependency tree includes gradio ≥ 6.15.1 for the web interface. The frontend/src/app.jsx module uses Gradio to create interactive voice synthesis interfaces that communicate with the FastAPI backend.

Data Handling, Security, and Utilities

Supporting infrastructure includes tools for content ingestion, security, and system monitoring:

  • yt-dlp ≥ 2026.7.4 — Downloading audio from video platforms
  • pypdf ≥ 4.0 — Extracting text from PDF documents
  • cryptography ≥ 41 — Encryption and secure token handling
  • psutil ≥ 7.2.2 — System resource monitoring
  • tensorboardX — Training visualization and logging
  • alembic ≥ 1.16 — Database migration management

Installation Source File

All dependency specifications originate from pyproject.toml at the repository root. Key version constraints include:

  • torch>=2.4 (line 36)
  • fastapi<0.137 (line 111)
  • gradio>=6.15.1 (line 43)

The file also pins development tools and optional analytics packages like posthog ≥ 3.7 for usage telemetry.

Summary

  • The VoiceStudio backend dependencies are centralized in pyproject.toml with strict semantic versioning
  • PyTorch 2.4+ and Transformers 5.10+ provide the deep learning foundation
  • Audio processing combines pydub for manipulation, whisperx for transcription, and kittentts for synthesis
  • FastAPI and Uvicorn power the REST API and WebSocket endpoints defined in backend/spec.py
  • Platform-specific packages like mlx-whisper and mlx-audio optimize performance for Apple Silicon Macs

Frequently Asked Questions

What file contains the VoiceStudio backend dependencies?

All production dependencies are declared in pyproject.toml at the repository root. This file pins exact versions for the PyTorch ecosystem, FastAPI web stack, and audio processing libraries referenced throughout the backend/ and services/ directories.

Which deep learning frameworks does VoiceStudio use?

The backend relies on PyTorch ≥ 2.4 as its primary tensor computation engine, supplemented by HuggingFace Transformers ≥ 5.10.0 for model architecture implementations and Accelerate for hardware optimization. These support both CPU and GPU inference pipelines.

Are there platform-specific dependencies for VoiceStudio?

Yes. The pyproject.toml includes environment markers for Apple Silicon (macOS ARM) that install mlx-whisper, mlx-audio, and parakeet-mlx only on compatible devices. These packages leverage Apple's MLX framework for accelerated inference on M1/M2/M3 chips.

How does VoiceStudio handle speech-to-text processing?

The backend supports multiple ASR backends through services/asr.py, including WhisperX for alignment-accurate transcription, faster-whisper for speed-optimized processing, and sherpa-onnx for offline recognition. The system defaults to WhisperX ≥ 3.1.0 for most production transcoding tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →