Dependencies for Speech-to-Speech: Complete Guide to Installation Requirements
The speech-to-speech package declares its dependencies in pyproject.toml, organized into core runtime requirements (22 packages), optional extras for specific model backends (11 groups), and development tools.
The Hugging Face speech-to-speech repository is a real-time voice-to-voice conversation system that requires specific Python dependencies for audio processing, model inference, and web server functionality. Understanding these dependencies is essential for successful installation and deployment, whether you are running the minimal pipeline or integrating specialized backends like ChatTTS or Faster-Whisper.
Core Runtime Dependencies
The essential packages required to run the base pipeline are defined in pyproject.toml starting at line 27. These include web framework components, audio libraries, and machine learning frameworks:
- Web server and API:
fastapi,uvicorn, andwebsocketsprovide the asynchronous HTTP/WebSocket server infrastructure - Networking:
httpxfor async HTTP client operations - Audio I/O:
sounddeviceandsoundfilehandle microphone capture and audio playback - Scientific computing:
numpy,scipy,torch, andtorchaudioenable tensor manipulation and neural network inference - ML/NLP:
transformersfor model loading,nltk==3.10.0for text tokenization, andopenai==2.28.0for API compatibility - Data validation:
pydantic>=2.0for configuration management - Utilities:
richfor formatted console output andpillowfor image processing
# Install only core dependencies
pip install speech-to-speech
Optional Extras for Specific Backends
The repository supports modular installation via optional extras defined at line 64 in pyproject.toml. These allow you to install only the dependencies needed for your specific speech-to-text (STT) or text-to-speech (TTS) backend:
chattts – ChatTTS>=0.1.1 for the ChatTTS neural TTS model
facebook-mms – transformers>=4.57.0 for Facebook's Massively Multilingual Speech models
faster-whisper – faster-whisper>=1.0.3 for optimized Whisper inference
kokoro – kokoro>=0.9.2 (skipped on macOS) for the Kokoro TTS engine
language-detection – lingua-language-detector>=2.0.2 for automatic language identification
mlx-lm – mlx-lm, mlx-vlm, and mlx-metal for Apple Silicon optimized inference (macOS only)
paraformer – funasr, modelscope, and onnxruntime (Python < 3.11 only) for Alibaba's Paraformer models
pocket – pocket-tts>=0.1.0 for Pocket TTS integration
webrtc – aiortc>=1.9.0 for WebRTC peer-to-peer communication support
websocket – websockets>=12.0 (redundant with core but pinned version)
whisper-mlx – lightning-whisper-mlx>=0.0.10 for optimized Whisper on Apple Silicon (macOS only)
# Example: Install with ChatTTS and Faster-Whisper support
pip install "speech-to-speech[chattts,faster-whisper]"
Development Dependencies
For contributors and those running the test suite, the dev dependency group at line 111 includes linting and testing tools:
- Linting:
rufffor code formatting and import sorting - Type checking:
mypyfor static type analysis - Testing:
pytestandpytest-asynciofor unit and integration tests - Network testing:
websocketsandaiortcfor testing real-time communication components
# Install with development dependencies
pip install "speech-to-speech[dev]"
Platform-Specific Requirements
Several dependencies have platform-specific constraints, particularly for macOS (Darwin). According to the source code in pyproject.toml, certain packages require specific binary builds on macOS for compatibility:
numpy: Pinned to specific builds on macOS for binary compatibilitytorch: Uses macOS-specific wheels to ensure Metal Performance Shaders (MPS) supportkokoro: Excluded from macOS installations due to compatibility issues- MLX packages: The
mlx-lm,mlx-vlm,mlx-metal, andwhisper-mlxextras are exclusively for macOS with Apple Silicon
Linux and Windows users should install the standard PyTorch wheels with CUDA support where applicable, while macOS users should leverage the MLX extras for optimal performance on M1/M2/M3 chips.
How Dependencies Map to Architecture
The dependency structure directly supports the pipeline architecture implemented in src/speech_to_speech/s2s_pipeline.py and related modules:
Web Interface Layer – fastapi and uvicorn instantiate the HTTP server that hosts the speech-to-speech endpoint, while websockets and optional aiortc enable real-time bidirectional streaming.
Audio Processing Layer – sounddevice handles microphone input streams, soundfile manages audio file I/O, and torch/torchaudio provide the tensor operations necessary for audio preprocessing and neural vocoding.
Model Inference Layer – transformers loads the underlying speech recognition and language models, while specific extras like faster-whisper or ChatTTS provide optimized implementations of particular model classes found in src/speech_to_speech/STT/ and src/speech_to_speech/TTS/.
Orchestration Layer – pydantic validates the argument classes defined in src/speech_to_speech/arguments_classes/, ensuring type safety when configuring handlers like WhisperSTTHandlerArguments or ChatTTSArguments.
Installation Examples
Install the core package for basic functionality:
# Basic installation (includes torch, transformers, fastapi, etc.)
pip install "speech-to-speech==0.2.11"
Install with specific model support:
# For ChatTTS TTS backend
pip install "speech-to-speech[chattts]==0.2.11"
# For Paraformer STT (requires Python < 3.11)
pip install "speech-to-speech[paraformer]==0.2.11"
# For macOS with MLX acceleration
pip install "speech-to-speech[mlx-lm,whisper-mlx]==0.2.11"
Launch the pipeline after installation:
python -m speech_to_speech.s2s_pipeline
Summary
- Core dependencies (22 packages) are defined in
pyproject.tomllines 27-62 and includefastapi,torch,transformers, and audio processing libraries. - Optional extras (11 groups) at lines 64-99 allow modular installation of specific backends like ChatTTS, Faster-Whisper, and MLX-optimized models without bloating the base install.
- Development tools at lines 111-118 include
ruff,mypy, andpytestfor code quality and testing. - Platform constraints apply particularly to macOS users, who should use MLX-specific extras for optimal performance while avoiding Linux-only packages like
kokoro. - The dependency structure supports the modular architecture in
src/speech_to_speech/, separating concerns between web serving, audio I/O, and model inference.
Frequently Asked Questions
What are the minimum dependencies required to run speech-to-speech?
The minimum installation requires the 22 core runtime packages defined in pyproject.toml including fastapi, uvicorn, torch, transformers, sounddevice, and websockets. These provide sufficient functionality to run the base pipeline with default STT and TTS handlers, though specific model backends may require additional optional extras.
How do I install support for ChatTTS or Faster-Whisper specifically?
Use the optional extras syntax: pip install "speech-to-speech[chattts]" for ChatTTS support or pip install "speech-to-speech[faster-whisper]" for optimized Whisper inference. You can combine multiple extras using commas, such as "speech-to-speech[chattts,faster-whisper]", to install dependencies for multiple backends simultaneously.
Are there any macOS-specific dependency requirements?
Yes, macOS users with Apple Silicon should install the mlx-lm and whisper-mlx extras for optimal performance, while avoiding the kokoro extra which is skipped on Darwin platforms. The pyproject.toml specifies platform-specific markers that automatically handle version pinning for numpy and torch on macOS to ensure Metal Performance Shaders compatibility.
What Python version is required for Paraformer support?
The Paraformer STT backend requires Python versions less than 3.11 due to dependencies on funasr, modelscope, and onnxruntime that are not compatible with Python 3.11+. This constraint is explicitly defined in the paraformer optional extra group in pyproject.toml, whereas the core package supports standard Python versions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →