24-Layer Models vs Standard Models in Pocket-TTS: Architecture and Performance Differences
The 24-layer models in Pocket-TTS feature 24 transformer layers compared to the 6 layers in standard models, offering higher audio quality as undistilled preview versions but requiring approximately 4× the inference time and memory.
Pocket-TTS from Kyutai Labs provides text-to-speech capabilities through transformer-based architectures that vary in depth and optimization. When choosing between the lightweight standard models and the deeper 24-layer variants, you must balance audio fidelity against computational requirements. Understanding the specific architectural differences between these 24-layer models and standard models ensures you select the appropriate configuration for your latency and quality constraints.
Transformer Depth and Model Capacity
The primary distinction lies in the transformer layer count defined in the YAML configuration files. Standard models configure num_layers: 6, while the 24-layer variants set num_layers: 24, creating a significantly larger parameter space for modeling complex prosodic patterns.
In pocket_tts/config/french.yaml, the standard French model defines 6 transformer layers. Conversely, pocket_tts/config/french_24l.yaml specifies 24 layers, increasing the network's capacity to capture subtle linguistic nuances. This architectural difference directly impacts the model's ability to generate natural-sounding speech, with the deeper architecture producing more accurate prosody and intonation.
Distillation Status and Inference Performance
Standard models undergo knowledge distillation to create efficient, production-ready checkpoints, while the 24-layer models remain undistilled preview releases. This distinction affects both output quality and runtime characteristics.
Standard models (6-layer):
- Distilled for efficient inference and smaller footprint
- Approximately 4× faster generation speed
- Suitable for real-time and resource-constrained environments
24-layer models:
- Undistilled original checkpoints with full parameter count
- Higher audio fidelity and more natural prosody
- Significantly slower inference and higher memory usage
The pocket_tts/models/tts_model.py implementation loads the appropriate architecture based on the configuration, selecting either the distilled 6-layer weights or the full 24-layer weights accordingly.
Configuration and Usage Examples
The repository organizes these variants through distinct configuration files that point to separate weight checkpoints. The 24-layer models use paths like languages/french_24l/model.safetensors rather than the standard languages/french/model.safetensors.
Loading models in Python:
from pocket_tts import TTSModel
# Load standard 6-layer distilled model
tts_standard = TTSModel.load_model(language="french")
# Load 24-layer undistilled preview model
tts_24l = TTSModel.load_model(language="french_24l")
# Generate audio - same API, different performance characteristics
audio_standard = tts_standard.generate("Bonjour le monde")
audio_24l = tts_24l.generate("Bonjour le monde")
Using the CLI interface:
# Standard model (fast, distilled)
pocket-tts generate --language french "Bonjour le monde"
# 24-layer model (slow, high quality)
pocket-tts generate --language french_24l "Bonjour le monde"
According to the documentation in docs/CLI Commands/serve.md, the 24l suffix identifies these preview models that have not undergone the distillation process applied to standard variants.
Summary
- 24-layer models contain 24 transformer layers versus 6 in standard models, providing 4× the architectural depth for improved audio modeling
- Standard models are distilled for production use, while 24-layer models remain undistilled preview versions with higher fidelity
- Performance trade-off: 24-layer variants require significantly more compute time and memory, making them unsuitable for real-time applications
- Configuration differences include distinct
num_layersvalues and separate weight paths (french_24l/model.safetensorsvsfrench/model.safetensors) - Selection method: Use
language="french_24l"in Python or--language french_24lin CLI to load the deeper architecture
Frequently Asked Questions
Can 24-layer models be used for real-time text-to-speech applications?
No, the 24-layer models are not recommended for real-time use due to their inference latency being approximately four times higher than standard models. These variants are designed for offline batch processing or quality-critical scenarios where the highest fidelity audio outweighs speed requirements.
Are 24-layer models available for all languages in Pocket-TTS?
The 24-layer variants are currently limited to specific languages such as French (french_24l) and German (german_24l) as preview releases. Availability varies by language, so you should check the pocket_tts/config/ directory for _24l.yaml files corresponding to your target language.
Will the 24-layer models eventually be distilled into standard models?
The standard models represent distilled versions of larger architectures, and the 24-layer models are currently provided as undistilled preview checkpoints. While the repository documentation indicates these are "preview" versions, future releases may include distilled 24-layer variants once the compression process completes, though this depends on the Kyutai Labs development roadmap.
How much additional memory do 24-layer models require compared to standard models?
The 24-layer models require substantially more VRAM or system RAM due to the 4× increase in transformer layers and lack of distillation compression. Memory usage scales roughly proportionally with layer count, so expect significantly higher resource requirements when loading configurations from pocket_tts/config/french_24l.yaml compared to the standard 6-layer configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →