Pocket-TTS Sample Rate and Audio Resampling: A Technical Guide

Pocket-TTS generates audio at a fixed 24 kHz sample rate and uses lightweight convolutional blocks in pocket_tts/modules/resample.py to automatically resample input and output audio.

The kyutai-labs/pocket-tts repository implements a neural text-to-speech system that standardizes on a 24 kHz sample rate for all audio generation. This sample rate is hardcoded in the model configuration and enforced throughout the inference pipeline, from the Mimi codec to the final WAV output. When working with external audio files—whether for voice cloning prompts or output storage—the library automatically handles conversion using specialized convolutional resampling modules.

Pocket-TTS Output Sample Rate

The model outputs audio at 24 kHz (24,000 Hz), as defined in the language-specific configuration files. For the default English model, this setting is located in pocket_tts/config/english.yaml at line 28:

sample_rate: 24000

This value is exposed programmatically through the TTSModel class in pocket_tts/models/tts_model.py. All internal processing—including the Mimi codec, flow-LM, and streaming modules—operates at this fixed rate to ensure consistency across the generative pipeline.

How Audio Resampling Works in Pocket-TTS

When the input sample rate differs from the target 24 kHz, Pocket-TTS performs resampling using two custom PyTorch modules located in pocket_tts/modules/resample.py. These classes implement integer-factor resampling using learned convolutional layers rather than traditional interpolation methods.

ConvDownsample1d: Integer Downsampling

The ConvDownsample1d class handles reduction of sample rates by integer factors. According to the source code at lines 13-26, it constructs a StreamingConv1d layer with specific parameters:

  • Kernel size: Set to 2 × stride to provide an adequate receptive field
  • Stride: The integer downsampling factor (e.g., 2, 4)

This design uses a regular convolution operation to compress the temporal dimension while preserving channel information, making it suitable for converting high-rate conditioning audio down to the 24 kHz target.

ConvTrUpsample1d: Integer Upsampling

For increasing sample rates, the ConvTrUpsample1d class (lines 37-48) implements upsampling via transposed convolution. It wraps a StreamingConvTranspose1d with identical kernel sizing:

  • Kernel size: 2 × stride
  • Stride: The integer upsampling factor

This transposed convolution approach allows the model to generate higher-resolution waveforms from internal latent representations when preparing final output or intermediate conditioning signals.

Resampling During the Inference Pipeline

The resampling modules integrate seamlessly into the TTS pipeline. When you provide an external audio file for voice cloning, the system automatically detects and converts the sample rate.

The process works as follows:

  1. Audio Loading: pocket_tts/data/audio.py reads the input file and returns both the waveform tensor and its original sample rate.
  2. Rate Checking: The pipeline compares the input rate against the model's configured 24 kHz target.
  3. Conditional Resampling: If rates differ, the code applies ConvDownsample1d (for rates above 24 kHz) or ConvTrUpsample1d (for rates below 24 kHz) to normalize the signal before Mimi encoding.
  4. Output Generation: Generated latent audio is automatically upsampled back to 24 kHz before being written to disk or streamed, ensuring all output files maintain consistent specifications.

Code Examples

The following examples demonstrate how Pocket-TTS handles sample rates automatically:

from pocket_tts import TTSModel

# Load the default English model (sample_rate = 24000)

tts = TTSModel.from_pretrained("english")

# Generate speech - output is always 24 kHz

audio, sr = tts.generate("Hello world!")
print(f"Sample rate: {sr}")  # Output: 24000

When using voice cloning with an arbitrary input file:


# Load an external audio file at any sample rate

cond_audio, cond_sr = tts.read_audio("voice_prompt.wav")

# The model internally resamples cond_audio to 24000 Hz

# using ConvDownsample1d or ConvTrUpsample1d as needed

audio, sr = tts.generate("Hello world!", conditioning=cond_audio)

Summary

  • Pocket-TTS outputs audio at a fixed 24 kHz sample rate, configured in pocket_tts/config/english.yaml.
  • The resampling mechanism uses two convolutional classes: ConvDownsample1d and ConvTrUpsample1d in pocket_tts/modules/resample.py.
  • Both modules use a kernel size of 2× stride with integer stride factors for efficient downsampling and upsampling.
  • The pipeline automatically resamples conditioning audio and generated output to maintain the 24 kHz standard throughout inference.

Frequently Asked Questions

Does Pocket-TTS support output sample rates other than 24 kHz?

No, the model is designed to operate exclusively at 24 kHz. This value is hardcoded in the configuration files and enforced throughout the internal processing pipeline. While you could theoretically resample the output after generation using external tools, the model itself does not support configurable output rates.

What happens if I provide a 44.1 kHz or 48 kHz audio file as a voice prompt?

The library automatically detects the higher sample rate and applies ConvDownsample1d to convert the audio to 24 kHz before processing it through the Mimi encoder. This conversion happens transparently during the read_audio or conditioning preparation stage, requiring no manual intervention.

Why does Pocket-TTS use convolutional resampling instead of standard methods?

The convolutional approach (StreamingConv1d and StreamingConvTranspose1d) integrates natively with the PyTorch-based pipeline and supports streaming inference. The learned convolution kernels provide better control over aliasing and frequency response compared to simple interpolation methods, which is critical for maintaining audio quality in neural codec-based TTS systems.

How can I verify the sample rate of generated audio?

The TTSModel.generate() method returns a tuple of (audio, sample_rate), where sample_rate is always 24000. You can also access the expected rate programmatically via tts.sample_rate before generation, as exposed in pocket_tts/models/tts_model.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →