# 24-Layer Models vs Standard Models in Pocket-TTS: Architecture and Performance Differences

> Explore 24-layer vs standard Pocket-TTS models. Discover architectural differences, performance gains, and trade-offs in inference time and memory for superior audio quality.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: architecture
- Published: 2026-07-11

---

**The 24-layer models in Pocket-TTS feature 24 transformer layers compared to the 6 layers in standard models, offering higher audio quality as undistilled preview versions but requiring approximately 4× the inference time and memory.**

Pocket-TTS from Kyutai Labs provides text-to-speech capabilities through transformer-based architectures that vary in depth and optimization. When choosing between the lightweight standard models and the deeper 24-layer variants, you must balance audio fidelity against computational requirements. Understanding the specific architectural differences between these 24-layer models and standard models ensures you select the appropriate configuration for your latency and quality constraints.

## Transformer Depth and Model Capacity

The primary distinction lies in the transformer layer count defined in the YAML configuration files. Standard models configure `num_layers: 6`, while the 24-layer variants set `num_layers: 24`, creating a significantly larger parameter space for modeling complex prosodic patterns.

In [pocket_tts/config/french.yaml](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/config/french.yaml), the standard French model defines 6 transformer layers. Conversely, [pocket_tts/config/french_24l.yaml](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/config/french_24l.yaml) specifies 24 layers, increasing the network's capacity to capture subtle linguistic nuances. This architectural difference directly impacts the model's ability to generate natural-sounding speech, with the deeper architecture producing more accurate prosody and intonation.

## Distillation Status and Inference Performance

Standard models undergo knowledge distillation to create efficient, production-ready checkpoints, while the 24-layer models remain undistilled preview releases. This distinction affects both output quality and runtime characteristics.

**Standard models (6-layer):**
- Distilled for efficient inference and smaller footprint
- Approximately 4× faster generation speed
- Suitable for real-time and resource-constrained environments

**24-layer models:**
- Undistilled original checkpoints with full parameter count
- Higher audio fidelity and more natural prosody
- Significantly slower inference and higher memory usage

The [pocket_tts/models/tts_model.py](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) implementation loads the appropriate architecture based on the configuration, selecting either the distilled 6-layer weights or the full 24-layer weights accordingly.

## Configuration and Usage Examples

The repository organizes these variants through distinct configuration files that point to separate weight checkpoints. The 24-layer models use paths like `languages/french_24l/model.safetensors` rather than the standard `languages/french/model.safetensors`.

Loading models in Python:

```python
from pocket_tts import TTSModel

# Load standard 6-layer distilled model

tts_standard = TTSModel.load_model(language="french")

# Load 24-layer undistilled preview model

tts_24l = TTSModel.load_model(language="french_24l")

# Generate audio - same API, different performance characteristics

audio_standard = tts_standard.generate("Bonjour le monde")
audio_24l = tts_24l.generate("Bonjour le monde")

```

Using the CLI interface:

```bash

# Standard model (fast, distilled)

pocket-tts generate --language french "Bonjour le monde"

# 24-layer model (slow, high quality)

pocket-tts generate --language french_24l "Bonjour le monde"

```

According to the documentation in [docs/CLI Commands/serve.md](https://github.com/kyutai-labs/pocket-tts/blob/main/docs/CLI%20Commands/serve.md), the `24l` suffix identifies these preview models that have not undergone the distillation process applied to standard variants.

## Summary

- **24-layer models** contain 24 transformer layers versus 6 in standard models, providing 4× the architectural depth for improved audio modeling
- **Standard models** are distilled for production use, while **24-layer models** remain undistilled preview versions with higher fidelity
- **Performance trade-off**: 24-layer variants require significantly more compute time and memory, making them unsuitable for real-time applications
- **Configuration differences** include distinct `num_layers` values and separate weight paths (`french_24l/model.safetensors` vs `french/model.safetensors`)
- **Selection method**: Use `language="french_24l"` in Python or `--language french_24l` in CLI to load the deeper architecture

## Frequently Asked Questions

### Can 24-layer models be used for real-time text-to-speech applications?

No, the 24-layer models are not recommended for real-time use due to their inference latency being approximately four times higher than standard models. These variants are designed for offline batch processing or quality-critical scenarios where the highest fidelity audio outweighs speed requirements.

### Are 24-layer models available for all languages in Pocket-TTS?

The 24-layer variants are currently limited to specific languages such as French (`french_24l`) and German (`german_24l`) as preview releases. Availability varies by language, so you should check the [pocket_tts/config/](https://github.com/kyutai-labs/pocket-tts/tree/main/pocket_tts/config) directory for [`_24l.yaml`](https://github.com/kyutai-labs/pocket-tts/blob/main/_24l.yaml) files corresponding to your target language.

### Will the 24-layer models eventually be distilled into standard models?

The standard models represent distilled versions of larger architectures, and the 24-layer models are currently provided as undistilled preview checkpoints. While the repository documentation indicates these are "preview" versions, future releases may include distilled 24-layer variants once the compression process completes, though this depends on the Kyutai Labs development roadmap.

### How much additional memory do 24-layer models require compared to standard models?

The 24-layer models require substantially more VRAM or system RAM due to the 4× increase in transformer layers and lack of distillation compression. Memory usage scales roughly proportionally with layer count, so expect significantly higher resource requirements when loading configurations from [pocket_tts/config/french_24l.yaml](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/config/french_24l.yaml) compared to the standard 6-layer configuration.