# Pocket-TTS Sample Rate and Audio Resampling: A Technical Guide

> Learn how Pocket-TTS generates audio at 24 kHz and uses lightweight convolutional blocks for efficient audio resampling. Understand the Pocket-TTS sample rate and resampling process.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: technical-guide
- Published: 2026-07-11

---

**Pocket-TTS generates audio at a fixed 24 kHz sample rate and uses lightweight convolutional blocks in [`pocket_tts/modules/resample.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/resample.py) to automatically resample input and output audio.**

The **kyutai-labs/pocket-tts** repository implements a neural text-to-speech system that standardizes on a **24 kHz sample rate** for all audio generation. This sample rate is hardcoded in the model configuration and enforced throughout the inference pipeline, from the Mimi codec to the final WAV output. When working with external audio files—whether for voice cloning prompts or output storage—the library automatically handles conversion using specialized convolutional resampling modules.

## Pocket-TTS Output Sample Rate

The model outputs audio at **24 kHz** (24,000 Hz), as defined in the language-specific configuration files. For the default English model, this setting is located in [`pocket_tts/config/english.yaml`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/config/english.yaml) at line 28:

```yaml
sample_rate: 24000

```

This value is exposed programmatically through the `TTSModel` class in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py). All internal processing—including the Mimi codec, flow-LM, and streaming modules—operates at this fixed rate to ensure consistency across the generative pipeline.

## How Audio Resampling Works in Pocket-TTS

When the input sample rate differs from the target 24 kHz, Pocket-TTS performs resampling using two custom PyTorch modules located in [`pocket_tts/modules/resample.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/resample.py). These classes implement integer-factor resampling using learned convolutional layers rather than traditional interpolation methods.

### ConvDownsample1d: Integer Downsampling

The `ConvDownsample1d` class handles reduction of sample rates by integer factors. According to the source code at lines 13-26, it constructs a `StreamingConv1d` layer with specific parameters:

- **Kernel size**: Set to `2 × stride` to provide an adequate receptive field
- **Stride**: The integer downsampling factor (e.g., 2, 4)

This design uses a regular convolution operation to compress the temporal dimension while preserving channel information, making it suitable for converting high-rate conditioning audio down to the 24 kHz target.

### ConvTrUpsample1d: Integer Upsampling

For increasing sample rates, the `ConvTrUpsample1d` class (lines 37-48) implements upsampling via transposed convolution. It wraps a `StreamingConvTranspose1d` with identical kernel sizing:

- **Kernel size**: `2 × stride`
- **Stride**: The integer upsampling factor

This transposed convolution approach allows the model to generate higher-resolution waveforms from internal latent representations when preparing final output or intermediate conditioning signals.

## Resampling During the Inference Pipeline

The resampling modules integrate seamlessly into the TTS pipeline. When you provide an external audio file for voice cloning, the system automatically detects and converts the sample rate.

The process works as follows:

1. **Audio Loading**: [`pocket_tts/data/audio.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/data/audio.py) reads the input file and returns both the waveform tensor and its original sample rate.
2. **Rate Checking**: The pipeline compares the input rate against the model's configured 24 kHz target.
3. **Conditional Resampling**: If rates differ, the code applies `ConvDownsample1d` (for rates above 24 kHz) or `ConvTrUpsample1d` (for rates below 24 kHz) to normalize the signal before Mimi encoding.
4. **Output Generation**: Generated latent audio is automatically upsampled back to 24 kHz before being written to disk or streamed, ensuring all output files maintain consistent specifications.

## Code Examples

The following examples demonstrate how Pocket-TTS handles sample rates automatically:

```python
from pocket_tts import TTSModel

# Load the default English model (sample_rate = 24000)

tts = TTSModel.from_pretrained("english")

# Generate speech - output is always 24 kHz

audio, sr = tts.generate("Hello world!")
print(f"Sample rate: {sr}")  # Output: 24000

```

When using voice cloning with an arbitrary input file:

```python

# Load an external audio file at any sample rate

cond_audio, cond_sr = tts.read_audio("voice_prompt.wav")

# The model internally resamples cond_audio to 24000 Hz

# using ConvDownsample1d or ConvTrUpsample1d as needed

audio, sr = tts.generate("Hello world!", conditioning=cond_audio)

```

## Summary

- Pocket-TTS outputs audio at a fixed **24 kHz sample rate**, configured in [`pocket_tts/config/english.yaml`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/config/english.yaml).
- The **resampling** mechanism uses two convolutional classes: `ConvDownsample1d` and `ConvTrUpsample1d` in [`pocket_tts/modules/resample.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/resample.py).
- Both modules use a **kernel size of 2× stride** with integer stride factors for efficient downsampling and upsampling.
- The pipeline automatically resamples conditioning audio and generated output to maintain the 24 kHz standard throughout inference.

## Frequently Asked Questions

### Does Pocket-TTS support output sample rates other than 24 kHz?

No, the model is designed to operate exclusively at 24 kHz. This value is hardcoded in the configuration files and enforced throughout the internal processing pipeline. While you could theoretically resample the output after generation using external tools, the model itself does not support configurable output rates.

### What happens if I provide a 44.1 kHz or 48 kHz audio file as a voice prompt?

The library automatically detects the higher sample rate and applies `ConvDownsample1d` to convert the audio to 24 kHz before processing it through the Mimi encoder. This conversion happens transparently during the `read_audio` or conditioning preparation stage, requiring no manual intervention.

### Why does Pocket-TTS use convolutional resampling instead of standard methods?

The convolutional approach (`StreamingConv1d` and `StreamingConvTranspose1d`) integrates natively with the PyTorch-based pipeline and supports streaming inference. The learned convolution kernels provide better control over aliasing and frequency response compared to simple interpolation methods, which is critical for maintaining audio quality in neural codec-based TTS systems.

### How can I verify the sample rate of generated audio?

The `TTSModel.generate()` method returns a tuple of `(audio, sample_rate)`, where `sample_rate` is always 24000. You can also access the expected rate programmatically via `tts.sample_rate` before generation, as exposed in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py).