How TTSModel.load_model() Handles Language Configuration and Model Selection in Pocket TTS

TTSModel.load_model() selects language configurations by mapping string identifiers to YAML files in pocket_tts/config/, with automatic fallback to English and strict validation rules that reject conflicting arguments and unsupported plain "french" identifiers.

The TTSModel.load_model() method in kyutai-labs/pocket-tts serves as the primary factory for initializing text-to-speech models. It provides a streamlined interface that automatically resolves language-specific architectures and weights while offering flexibility for custom configurations. Understanding how this method handles the language parameter is essential for correctly instantiating models across the supported multilingual capabilities of the Pocket TTS library.

Language Configuration Resolution Logic

Mutual Exclusion and Argument Validation

In pocket_tts/models/tts_model.py, the method first enforces mutual exclusion between custom config files and language identifiers. Lines 89-92 raise a ValueError if both the config and language arguments are provided simultaneously, preventing ambiguous configuration states.

Default Language Fallback

When neither argument is supplied, the method falls back to the library-wide default defined in pocket_tts/default_parameters.py. Lines 93-95 import DEFAULT_LANGUAGE, which resolves to "english" (see lines 1-2 in default_parameters.py), ensuring consistent behavior across different environments.

Special Case: French Model Restrictions

The implementation includes specific validation for French language requests. Lines 96-99 explicitly reject the plain "french" identifier with a descriptive error message, as only the 24-layer variant is available. Users must specify "french_24l" to proceed.

Available Language Options and Config Mappings

Supported Language Identifiers

As documented in the method's docstring (lines 50-53 in tts_model.py), the following identifiers map to configuration files in pocket_tts/config/:

Configuration Path Resolution

For valid language strings (excluding the rejected "french"), lines 100-101 construct the configuration path using CONFIGS_DIR / f"{language}.yaml". Lines 102-106 validate the .yaml suffix and parse the file via load_config from pocket_tts/utils/config.py, returning a validated Pydantic Config object.

Model Instantiation and Optional Quantization

Architecture Loading Pipeline

With the parsed configuration, lines 108-110 invoke the private helper _from_pydantic_config_with_weights. This creates the FlowLMModel and MimiModel instances, loads all associated weight files, and returns a fully initialized TTSModel instance (line 115).

Dynamic Int-8 Quantization

If quantize=True is passed, lines 112-113 apply dynamic int-8 quantization to the flow language model after loading, optimizing for CPU inference performance regardless of the selected language.

Practical Usage Examples

from pocket_tts import TTSModel

# Load default English model (uses DEFAULT_LANGUAGE fallback)

model_default = TTSModel.load_model()
print(model_default.sample_rate)  # Output: 24000

# Load specific language variant

model_french = TTSModel.load_model(language="french_24l")

# Load custom fine-tuned configuration (mutually exclusive with language arg)

model_custom = TTSModel.load_model(config="my_fine_tuned.yaml")

# Enable int-8 quantization for faster CPU inference

model_quant = TTSModel.load_model(language="german_24l", quantize=True)

# This raises ValueError: plain "french" is not supported

# TTSModel.load_model(language="french")

Summary

  • TTSModel.load_model() enforces mutual exclusion between language and config arguments, raising ValueError if both are provided according to lines 89-92.
  • The method defaults to "english" when no language or config is specified, utilizing DEFAULT_LANGUAGE from pocket_tts/default_parameters.py (lines 93-95).
  • Plain "french" is explicitly rejected; users must specify "french_24l" to access the only available French variant (lines 96-99).
  • Language identifiers map directly to YAML files in pocket_tts/config/, parsed via load_config and validated before model instantiation (lines 100-106).
  • The helper _from_pydantic_config_with_weights handles actual model creation and weight loading (lines 108-110).
  • Optional quantize=True applies dynamic int-8 quantization to the flow model for CPU optimization (lines 112-113).

Frequently Asked Questions

What happens if I provide both a config file and a language identifier?

The method raises a ValueError immediately. According to lines 89-92 in pocket_tts/models/tts_model.py, TTSModel.load_model() treats these arguments as mutually exclusive to prevent configuration conflicts.

Why can't I use "french" as a language option?

The plain "french" identifier is blocked by design. Lines 96-99 in tts_model.py explicitly reject it because only a 24-layer French model exists in the repository. You must specify "french_24l" instead, which maps to french_24l.yaml in the configs directory.

How does the method determine which configuration file to load?

It constructs a path using CONFIGS_DIR / f"{language}.yaml" (lines 100-101). The CONFIGS_DIR points to pocket_tts/config/ within the repository, and the method validates the file exists and ends with .yaml before parsing via load_config.

Can I use quantization with any language model?

Yes. When quantize=True is passed (lines 112-113), dynamic int-8 quantization is applied to the flow language model after instantiation, regardless of which specific language configuration was loaded.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →