How Temperature and Noise_Clamp Parameters Affect Pocket-TTS Generation Quality

Temperature scales the variance of injected Gaussian noise to trade off between diverse, expressive speech and stable, high-fidelity output, while noise_clamp truncates extreme values to prevent out-of-distribution artifacts during flow-based decoding.

The kyutai-labs/pocket-tts repository implements a flow-based language model for text-to-speech synthesis that relies on stochastic sampling during latent code generation. Understanding how temperature and noise_clamp parameters influence this process is essential for tuning the balance between creative variation and audio clarity.

Temperature Controls Noise Variance in FlowLMModel

In pocket_tts/models/flow_lm.py, the temperature parameter directly scales the standard deviation of Gaussian noise injected before Lagrangian Self-Distillation (LSD) decoding. According to the source code, the implementation calculates std = temp**0.5 within the FlowLMModel.forward method, meaning higher temperatures exponentially increase the variance of the noise distribution.

  • Higher temperature (> 1.0): Increases std above 1.0, encouraging the model to explore more of the latent space. This produces more varied speech with diverse prosody and occasional creative phrasing, but raises the risk of artifacts and reduced intelligibility.

  • Lower temperature (< 1.0): Compresses the noise variance, yielding more deterministic, higher-fidelity output. Values around 0.7–0.8 tighten the distribution around the model's most likely predictions, sacrificing variation for consistency.

Noise_Clamp Prevents Extreme Artifacts

The noise_clamp parameter, defined in the load_model signature in pocket_tts/models/tts_model.py, determines whether noise sampling uses a truncated normal distribution. When specified, it limits the amplitude of random perturbations to prevent out-of-distribution spikes.

Implementation Details

In FlowLMModel.forward (lines 33–38 of pocket_tts/models/flow_lm.py), the code branches based on the noise_clamp value:

  • When noise_clamp is None: The model uses torch.nn.init.normal_ to draw from an unrestricted standard normal distribution, allowing more expressive output but potentially introducing high-amplitude artifacts in difficult regions like fast phoneme transitions.

  • When noise_clamp is set (e.g., 0.5): The model uses torch.nn.init.trunc_normal_ bounded by ±noise_clamp, cutting off extreme noise values. This stabilizes generation by preventing the latent trajectory from entering regions where the pretrained flow has not been calibrated.

Practical Configuration Examples

You can configure these parameters via the Python API or CLI. Here are practical recipes for different quality trade-offs:

High-quality, clear output:

  • temp ≈ 0.7–1.0
  • noise_clamp ≈ 0.5

Creative, expressive output:

  • temp ≈ 1.2–1.5
  • noise_clamp ≈ None or 1.0

Python API Usage

from pocket_tts import TTSModel

# Conservative settings for clarity

model = TTSModel.load_model(temp=0.8, noise_clamp=0.5)

# Expressive settings with more variation

model = TTSModel.load_model(temp=1.3, noise_clamp=None)

audio = model.generate("Hello, this is a test.", speaker="en_female")

CLI Usage


# Default generation (temp=1.0, noise_clamp=0.5)

uv run pocket-tts generate "Hello world"

# Higher fidelity, less noise

uv run pocket-tts generate "Hello world" --temperature 0.8 --noise-clamp 0.4

# More diverse output

uv run pocket-tts generate "Hello world" --temperature 1.4 --noise-clamp 1.0

Summary

  • Temperature directly scales noise standard deviation (std = temp**0.5) in FlowLMModel.forward, controlling the exploration-exploitation trade-off between diverse and deterministic speech.
  • Noise_Clamp truncates the noise distribution at ±noise_clamp using torch.nn.init.trunc_normal_, preventing extreme latent values that cause audible glitches.
  • For production use, combine temp values of 0.7–1.0 with noise_clamp of 0.5 to balance fidelity and stability.
  • For creative applications, raise temp to 1.2–1.5 and optionally disable clamping to maximize expressive variation.

Frequently Asked Questions

For production environments requiring clear, consistent speech, set temperature between 0.7 and 1.0. This range keeps the noise variance close to the model's training distribution while allowing slight prosodic variation. Values below 0.7 may sound robotic, while values above 1.2 increase the risk of pronunciation artifacts.

How does noise_clamp differ from temperature scaling?

While temperature scales the overall variance of the noise distribution, noise_clamp imposes hard limits on individual sample values. Temperature affects the "spread" of the Gaussian curve, whereas clamping cuts off the distribution's tails. This means clamping specifically targets extreme outliers that cause isolated audio glitches, without affecting the general stochasticity of the generation process.

Can I disable noise_clamp entirely for maximum expressiveness?

Yes, by setting noise_clamp=None in TTSModel.load_model() or omitting the --noise-clamp flag in the CLI, you allow unbounded normal sampling. This maximizes expressive potential but may introduce higher-amplitude artifacts during complex phoneme transitions or rare word pronunciations. For critical applications, maintain a clamp value of at least 0.5.

Where are these parameters implemented in the Pocket-TTS source code?

The parameters are defined in pocket_tts/models/tts_model.py within the load_model signature, consumed in pocket_tts/models/flow_lm.py in the FlowLMModel.forward method (lines 31–38), and exposed via the CLI in pocket_tts/main.py. The actual noise sampling logic switches between torch.nn.init.normal_ and torch.nn.init.trunc_normal_ based on the clamp setting.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →