# How Temperature and Noise_Clamp Parameters Affect Pocket-TTS Generation Quality

> Discover how temperature and noise_clamp parameters impact Pocket-TTS generation quality. Understand the trade-offs between diverse speech and stable output for optimal results.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: deep-dive
- Published: 2026-07-11

---

**Temperature scales the variance of injected Gaussian noise to trade off between diverse, expressive speech and stable, high-fidelity output, while noise_clamp truncates extreme values to prevent out-of-distribution artifacts during flow-based decoding.**

The kyutai-labs/pocket-tts repository implements a flow-based language model for text-to-speech synthesis that relies on stochastic sampling during latent code generation. Understanding how **temperature** and **noise_clamp** parameters influence this process is essential for tuning the balance between creative variation and audio clarity.

## Temperature Controls Noise Variance in FlowLMModel

In [`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py), the `temperature` parameter directly scales the standard deviation of Gaussian noise injected before Lagrangian Self-Distillation (LSD) decoding. According to the source code, the implementation calculates `std = temp**0.5` within the `FlowLMModel.forward` method, meaning higher temperatures exponentially increase the variance of the noise distribution.

- **Higher temperature (> 1.0)**: Increases `std` above 1.0, encouraging the model to explore more of the latent space. This produces **more varied speech** with diverse prosody and occasional creative phrasing, but raises the risk of artifacts and reduced intelligibility.

- **Lower temperature (< 1.0)**: Compresses the noise variance, yielding **more deterministic, higher-fidelity** output. Values around 0.7–0.8 tighten the distribution around the model's most likely predictions, sacrificing variation for consistency.

## Noise_Clamp Prevents Extreme Artifacts

The `noise_clamp` parameter, defined in the `load_model` signature in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py), determines whether noise sampling uses a truncated normal distribution. When specified, it limits the amplitude of random perturbations to prevent out-of-distribution spikes.

### Implementation Details

In `FlowLMModel.forward` (lines 33–38 of [`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py)), the code branches based on the `noise_clamp` value:

- **When `noise_clamp` is `None`**: The model uses `torch.nn.init.normal_` to draw from an unrestricted standard normal distribution, allowing **more expressive** output but potentially introducing high-amplitude artifacts in difficult regions like fast phoneme transitions.

- **When `noise_clamp` is set** (e.g., `0.5`): The model uses `torch.nn.init.trunc_normal_` bounded by `±noise_clamp`, cutting off extreme noise values. This stabilizes generation by preventing the latent trajectory from entering regions where the pretrained flow has not been calibrated.

## Practical Configuration Examples

You can configure these parameters via the Python API or CLI. Here are practical recipes for different quality trade-offs:

**High-quality, clear output:**

- `temp` ≈ 0.7–1.0
- `noise_clamp` ≈ 0.5

**Creative, expressive output:**

- `temp` ≈ 1.2–1.5
- `noise_clamp` ≈ `None` or 1.0

### Python API Usage

```python
from pocket_tts import TTSModel

# Conservative settings for clarity

model = TTSModel.load_model(temp=0.8, noise_clamp=0.5)

# Expressive settings with more variation

model = TTSModel.load_model(temp=1.3, noise_clamp=None)

audio = model.generate("Hello, this is a test.", speaker="en_female")

```

### CLI Usage

```bash

# Default generation (temp=1.0, noise_clamp=0.5)

uv run pocket-tts generate "Hello world"

# Higher fidelity, less noise

uv run pocket-tts generate "Hello world" --temperature 0.8 --noise-clamp 0.4

# More diverse output

uv run pocket-tts generate "Hello world" --temperature 1.4 --noise-clamp 1.0

```

## Summary

- **Temperature** directly scales noise standard deviation (`std = temp**0.5`) in `FlowLMModel.forward`, controlling the exploration-exploitation trade-off between diverse and deterministic speech.
- **Noise_Clamp** truncates the noise distribution at `±noise_clamp` using `torch.nn.init.trunc_normal_`, preventing extreme latent values that cause audible glitches.
- For production use, combine `temp` values of 0.7–1.0 with `noise_clamp` of 0.5 to balance fidelity and stability.
- For creative applications, raise `temp` to 1.2–1.5 and optionally disable clamping to maximize expressive variation.

## Frequently Asked Questions

### What is the recommended temperature for production speech synthesis?

For production environments requiring clear, consistent speech, set `temperature` between **0.7 and 1.0**. This range keeps the noise variance close to the model's training distribution while allowing slight prosodic variation. Values below 0.7 may sound robotic, while values above 1.2 increase the risk of pronunciation artifacts.

### How does noise_clamp differ from temperature scaling?

While **temperature** scales the overall variance of the noise distribution, **noise_clamp** imposes hard limits on individual sample values. Temperature affects the "spread" of the Gaussian curve, whereas clamping cuts off the distribution's tails. This means clamping specifically targets extreme outliers that cause isolated audio glitches, without affecting the general stochasticity of the generation process.

### Can I disable noise_clamp entirely for maximum expressiveness?

Yes, by setting `noise_clamp=None` in `TTSModel.load_model()` or omitting the `--noise-clamp` flag in the CLI, you allow unbounded normal sampling. This maximizes expressive potential but may introduce **higher-amplitude artifacts** during complex phoneme transitions or rare word pronunciations. For critical applications, maintain a clamp value of at least 0.5.

### Where are these parameters implemented in the Pocket-TTS source code?

The parameters are defined in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) within the `load_model` signature, consumed in [`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py) in the `FlowLMModel.forward` method (lines 31–38), and exposed via the CLI in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py). The actual noise sampling logic switches between `torch.nn.init.normal_` and `torch.nn.init.trunc_normal_` based on the clamp setting.