# How to Use apply_dynamic_int8 for Int8 Quantization in Pocket-TTS

> Reduce Pocket-TTS memory usage with apply_dynamic_int8. Convert transformer weights to int8 for about 50% memory savings with minimal audio quality loss. Learn how now.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: how-to-guide
- Published: 2026-07-11

---

**Call `apply_dynamic_int8(model.flow_lm, {"attn", "ffn"})` after loading a `TTSModel` to convert transformer weights to int8, cutting memory usage by roughly 50% while preserving over 95% of audio quality.**

Pocket-TTS from Kyutai Labs ships a lightweight int8 dynamic quantization pipeline that shrinks model memory footprints without sacrificing audio fidelity. The core routine, `apply_dynamic_int8`, lives in [`pocket_tts/quantization.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/quantization.py) and performs in-place quantization of the internal **FlowLM** transformer using group-wise selection and automatic backend detection.

## How apply_dynamic_int8 Works Internally

### Backend Detection via _get_backend()

When invoked, `apply_dynamic_int8` first calls `_get_backend()` to select the optimal quantization engine. If **TorchAO** (version ≥ 0.16) is installed, it routes to the high-performance `torchao.quantization` API. Otherwise, it falls back to PyTorch’s native `torch.ao.quantization`, automatically selecting QNNPACK on ARM devices and FBGEMM on x86 architectures.

### Group-Wise Quantization Strategy

Users specify layer-group keys—such as `"attn"` for attention blocks or `"ffn"` for feed-forward networks—to target specific modules within the FlowLM architecture. This selective approach lets you keep sensitive layers in full-precision float32 while compressing compute-heavy paths, balancing quality against memory savings.

### In-Place Weight Transformation

The chosen backend quantizes weights to int8 in-place, replacing 32-bit tensors with 8-bit integers while maintaining the original computation graph. Activations remain in float32 during inference, creating a dynamic quantization scheme that requires no calibration dataset.

## Applying apply_dynamic_int8 Programmatically

Because quantization mutates the model in-place, you can invoke it manually after loading to experiment with different group configurations:

```python
from pocket_tts import TTSModel
from pocket_tts.quantization import apply_dynamic_int8

# Load full-precision model

model = TTSModel.load_model(quantize=False)

# Quantize attention and feed-forward groups

apply_dynamic_int8(model.flow_lm, {"attn", "ffn"})

# Generate with int8 weights

audio = model.generate_audio(text="Hello world!")

```

## CLI Integration with --quantize Flag

The command-line interface wraps the same logic. Passing `--quantize` triggers automatic backend detection and applies default groups defined in `pocket_tts/config/*.yaml`:

```bash
uv run pocket-tts generate \
    --text "Fast and tiny speech" \
    --output-path out.wav \
    --quantize

```

The flag also works with the `serve` command, enabling low-memory quantized servers.

## Memory Savings vs. Quality Trade-offs

According to [`scripts/evaluate_quantization.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/scripts/evaluate_quantization.py), quantizing the standard "attn" and "ffn" groups yields a **~50% reduction** in model size while retaining **>95% of baseline quality** metrics (SNR, PESQ, WER). These benchmarks are documented in [`docs/quantization.md`](https://github.com/kyutai-labs/pocket-tts/blob/main/docs/quantization.md).

## Summary

- **`apply_dynamic_int8`** in [`pocket_tts/quantization.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/quantization.py) provides dynamic int8 quantization for the FlowLM transformer.
- The routine auto-detects the best backend (TorchAO ≥0.16 preferred) and applies group-wise quantization based on user-specified keys like `"attn"` and `"ffn"`.
- Quantization happens in-place, allowing manual application after model loading or automatic application via the `--quantize` CLI flag.
- Typical configurations reduce memory by 50% with minimal quality degradation (>95% retention).

## Frequently Asked Questions

### What is the difference between TorchAO and the legacy PyTorch backend?

TorchAO (version 0.16 or newer) provides optimized quantization kernels that are faster than the legacy `torch.ao.quantization` path. If TorchAO is not installed, Pocket-TTS automatically falls back to QNNPACK on ARM or FBGEMM on x86, ensuring cross-platform compatibility without code changes.

### Can I quantize only specific parts of the model?

Yes. The `quantize_groups` parameter accepts a set of strings mapping to layer collections inside FlowLM. Common choices include `{"attn"}` for attention blocks only or `{"attn", "ffn"}` for both attention and feed-forward layers. This lets you exclude layers where precision is critical.

### Do I need a calibration dataset for dynamic quantization?

No. `apply_dynamic_int8` uses dynamic quantization, which keeps activations in float32 and converts only weights to int8. This eliminates the need for calibration data or post-training quantization passes, making it ideal for quick deployment.

### How do I verify that quantization actually reduced memory usage?

You can check the model size before and after calling `apply_dynamic_int8` by inspecting `model.flow_lm` parameters. Additionally, the evaluation script [`scripts/evaluate_quantization.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/scripts/evaluate_quantization.py) runs SNR, PESQ, and WER benchmarks to confirm that audio quality remains within acceptable thresholds after quantization.