How to Use the pocket-tts CLI for Speech Generation with Custom Options
The pocket-tts CLI provides a Typer-based command-line interface in pocket_tts/main.py that supports custom voices, generation parameters, and device selection for high-quality text-to-speech synthesis.
The pocket-tts command-line tool from the kyutai-labs/pocket-tts repository offers fine-grained control over speech generation through a comprehensive set of options. Built on the Typer framework, the CLI exposes the generate sub-command that orchestrates model loading, voice conditioning, and streaming audio output. This guide explains how to leverage custom options for tailoring voice characteristics, generation quality, and hardware utilization.
CLI Architecture and Entry Point
The CLI entry point resides in pocket_tts/main.py, where the cli_app.command() decorator registers the generate sub-command at lines 22-84. Typer automatically generates help text and parses arguments, mapping command-line flags directly to the underlying TTSModel methods.
When you execute pocket-tts generate, the tool chains together several discrete operations:
- Logging configuration via
enable_logging, which respects the--quietflag to suppress informational output - Default text resolution using
get_default_text_for_languagewhen--textis omitted - Model initialization through
TTSModel.load_modelwith optional INT8 quantization - Voice state preparation via
TTSModel.get_state_for_audio_promptwith LRU caching for repeated prompts - Streaming generation using
generate_audio_streamand the Mimi codec decoder - Audio output through
stream_audio_chunksto either a file path or standard output
Core Generation Pipeline
Understanding the internal pipeline helps optimize your CLI usage with custom options.
Logging and Default Text Handling
By default, the CLI enables informational logging. Use --quiet to disable verbose output. If you omit the --text argument, the system calls get_default_text_for_language from pocket_tts/default_parameters.py to retrieve a language-specific demonstration prompt.
Model Loading and Quantization
The TTSModel.load_model function initializes the pre-trained weights and can apply INT8 quantization when you pass the --quantize flag. This reduces memory footprint and accelerates CPU inference without requiring GPU acceleration.
Voice Conditioning and Caching
Voice customization happens through two pathways:
- Default voices: If
--voiceis omitted,get_default_voice_for_languageselects a built-in voice (such as alba for English) from the default parameters file - Custom voices: You may specify local
.wavfiles or HuggingFace URLs via--voice
The TTSModel.get_state_for_audio_prompt method extracts a model state encoding the speaker's timbre. The CLI caches these states using Python's lru_cache mechanism, ensuring efficient reuse when processing multiple texts with the same voice.
Streaming Audio Output
The generate_audio_stream generator yields audio chunks while a background thread decodes them through the Mimi codec. The stream_audio_chunks function handles the final write operation, directing output to either a specified file path (--output-path) or standard output when set to -.
Customization Options and Examples
The pocket-tts CLI exposes parameters that map directly to the generation routine, allowing precise control over temperature, decoding steps, noise clamping, and device placement.
Basic Usage and Language Selection
Generate speech using English defaults:
pocket-tts generate
Specify custom text and language:
pocket-tts generate --text "The quick brown fox jumps over the lazy dog."
pocket-tts generate --language italian --text "Ciao mondo"
Voice Customization
Provide a local voice sample or reference a HuggingFace-hosted voice:
# Local voice file
pocket-tts generate \
--text "Testing custom voice conditioning." \
--voice ./my_voice.wav
# HuggingFace URL
pocket-tts generate \
--text "Bonjour, je parle avec une voix différente." \
--voice hf://kyutai/tts-voices/alba-mackenna/casual.wav
Generation Quality Parameters
Control the stochasticity and depth of the generation process:
pocket-tts generate \
--text "Testing higher temperature." \
--temperature 1.2 \
--lsd-decode-steps 3 \
--noise-clamp 0.8
Tune end-of-speech detection to control truncation behavior:
pocket-tts generate \
--text "Short phrase." \
--eos-threshold -3.0 \
--frames-after-eos 5
Limit token consumption per chunk (default is 50):
pocket-tts generate \
--text "A long paragraph that will be split into multiple chunks." \
--max-tokens 80
Output and Device Control
Target specific hardware and output destinations:
# Save to specific path
pocket-tts generate \
--text "Saving to a custom location." \
--output-path /tmp/output.wav
# Stream to stdout for piping
pocket-tts generate --text "Direct output" --output-path -
# GPU acceleration (defaults to CPU if unavailable)
pocket-tts generate --text "Running on GPU." --device cuda
# INT8 quantization for faster CPU inference
pocket-tts generate --text "Quantised model." --quantize
Summary
- The pocket-tts CLI in
pocket_tts/main.pyuses Typer to expose thegeneratecommand with extensive customization options - Voice conditioning supports local files, HuggingFace URLs, or built-in defaults via
get_default_voice_for_language - Model loading through
TTSModel.load_modelsupports INT8 quantization via the--quantizeflag - Generation parameters include
--temperature,--lsd-decode-steps,--noise-clamp,--eos-threshold, and--max-tokensfor fine-tuning output quality - Hardware control allows device selection via
--device(cuda/cpu) and output redirection via--output-path - The streaming pipeline uses
generate_audio_streamand the Mimi codec for efficient audio synthesis
Frequently Asked Questions
How do I use a custom voice file with the pocket-tts CLI?
Pass the path to your .wav file using the --voice argument. The CLI supports local file paths (e.g., --voice ./my_voice.wav) and HuggingFace URLs (e.g., --voice hf://kyutai/tts-voices/alba-mackenna/casual.wav). The TTSModel.get_state_for_audio_prompt function processes these files and caches the resulting voice state for efficient reuse.
What parameters control the quality and speed of speech generation?
Quality is controlled by --temperature (stochasticity, default varies), --lsd-decode-steps (decoding iterations), and --noise-clamp (noise limitation). For faster inference on CPU, enable INT8 quantization with the --quantize flag. Use --max-tokens to adjust the token budget per chunk, which affects both quality and generation speed.
Can I stream the audio output to another program instead of saving to a file?
Yes. Set --output-path to - (hyphen) to stream raw audio chunks to standard output. This allows you to pipe the output to other command-line tools or media players. The stream_audio_chunks function handles both file writes and stdout streaming.
How does the CLI handle language selection and default text?
When you specify --language (e.g., italian or english), the CLI calls get_default_text_for_language from pocket_tts/default_parameters.py to retrieve an appropriate demonstration text if you omit the --text argument. It also selects a default voice for that language using get_default_voice_for_language, falling back to built-in voices like alba for English.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →