Difference Between generate_audio() and generate_audio_stream() in Pocket-TTS
The generate_audio() method returns the complete audio waveform as a single torch.Tensor, while generate_audio_stream() yields audio chunks progressively through a generator, enabling real-time playback with minimal latency.
Both methods are implemented in the TTSModel class within the kyutai-labs/pocket-tts repository in pocket_tts/models/tts_model.py. While they produce identical audio output from the same text input, their internal architecture differs significantly depending on whether you need batch synthesis or streaming synthesis.
Internal Architecture and Implementation
The relationship between these methods is hierarchical: generate_audio() acts as a convenience wrapper that consumes the streaming output of generate_audio_stream() and concatenates the results before returning.
The Synchronous Wrapper: generate_audio()
Located at lines 776‑785 in pocket_tts/models/tts_model.py, this method implements a simple loop:
audio_chunks = []
for chunk in self.generate_audio_stream(...):
audio_chunks.append(chunk)
return torch.cat(audio_chunks, dim=0)
This means you wait for the entire text to be processed before receiving any audio data. The method returns a tensor with shape [channels, samples], making it suitable for immediate file writing or post-processing. Because it aggregates all chunks in memory before returning, it requires sufficient RAM to hold the complete audio waveform.
The Streaming Pipeline: generate_audio_stream()
The actual synthesis logic resides in generate_audio_stream() at lines 445‑456. This method orchestrates a multi-threaded pipeline that includes:
- Text Segmentation: Uses
split_into_best_sentencesto divide input into optimal processing units - Worker Threading: Spawns a decoder worker thread (
_decode_audio_worker) that consumes latent vectors from a queue and decodes them into audio chunks - Concurrent Generation: Runs
_generate()in a separate thread to produce latents and push them into the same queue - Progressive Yield: Returns each decoded chunk immediately via a generator as it becomes available
Each yielded chunk is a 1-D torch.Tensor with shape [samples] (single-channel, batch dimension removed), allowing downstream audio players to consume data while generation continues.
Performance and Latency Characteristics
generate_audio() introduces higher perceived latency because it blocks until the entire text sequence is synthesized. This approach is ideal when you need the complete audio for non-real-time applications such as file export or batch processing.
generate_audio_stream() provides lower latency by yielding audio chunks as soon as they are decoded. This enables real-time playback scenarios where audio can begin playing while the model is still processing subsequent text segments. The streaming approach also maintains a lower memory footprint for long texts since it processes chunks incrementally rather than storing the entire waveform.
Both methods accept a copy_state parameter (default True). When enabled, the original model_state is deep-copied before generation, preserving it for reuse. In the streaming variant implemented via _generate_audio_stream_short_text, this copying occurs inside the pipeline for each short-text segment.
Thread Safety Considerations
Neither method is thread-safe. Each TTSModel instance should be accessed by only one thread at a time. The streaming implementation spawns internal decoder threads, but these are managed within the method scope and joined before completion. Attempting to call either method concurrently from multiple threads on the same instance will result in race conditions.
Practical Code Examples
Batch Processing with generate_audio()
Use this approach when you need the complete waveform for file output or offline processing:
from pocket_tts import TTSModel
# Initialize the model
model = TTSModel.load_model()
# Configure voice state from an audio sample
voice_state = model.get_state_for_audio_prompt(
"hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)
# Generate complete audio tensor
audio = model.generate_audio(
model_state=voice_state,
text_to_generate="Hello world! This is a fully rendered audio example.",
frames_after_eos=2,
)
print(f"Audio shape: {audio.shape}") # → [channels, samples]
# Ready for torchaudio.save() or similar
Real-Time Streaming with generate_audio_stream()
Use this for interactive applications where audio should play during generation:
from pocket_tts import TTSModel
model = TTSModel.load_model()
voice_state = model.get_state_for_audio_prompt(
"hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)
# Stream audio chunk-by-chunk
for i, chunk in enumerate(
model.generate_audio_stream(
model_state=voice_state,
text_to_generate="This is a long paragraph that will be streamed in real time.",
frames_after_eos=None,
)
):
# chunk is a 1-D Tensor: [samples]
print(f"Chunk {i}: {chunk.shape[0]} samples")
# Send to audio player or write to circular buffer
Summary
generate_audio()is a synchronous wrapper located at lines 776‑785 inpocket_tts/models/tts_model.pythat returns a complete[channels, samples]tensor after processing all textgenerate_audio_stream()implements the core streaming logic at lines 445‑456, yielding[samples]chunks immediately via a generator for low-latency applications- Both methods support the
copy_stateparameter to preserve model state and are not thread-safe - The streaming method uses internal worker threads (
_decode_audio_worker) and text segmentation (split_into_best_sentences) to enable real-time audio production
Frequently Asked Questions
Can I use generate_audio_stream() for file output?
Yes, though it requires manual concatenation. You must iterate through the generator and accumulate chunks into a list before calling torch.cat(), essentially replicating what generate_audio() does internally. For direct file output, generate_audio() is more convenient as it returns the complete tensor ready for torchaudio.save().
Why does generate_audio_stream() return 1-D tensors instead of 2-D?
The streaming method yields single-channel mono audio with shape [samples] rather than [1, samples] to simplify integration with real-time audio pipelines. Most streaming audio consumers expect 1-D buffers. You can reshape the chunks to [1, samples] if your downstream processor requires explicit channel dimensions.
Is generate_audio() slower than generate_audio_stream()?
The total processing time is nearly identical because generate_audio() simply wraps the streaming logic. However, generate_audio() has higher perceived latency because it blocks until completion. For long texts, generate_audio_stream() allows playback to begin immediately while generation continues in background threads.
How does error handling differ between the two methods?
With generate_audio(), errors propagate after the full process completes. In contrast, generate_audio_stream() propagates errors immediately as they occur in the generation or decoding threads. The implementation ensures the decoder thread is cleanly joined via error handling logic at lines 462‑470, preventing resource leaks even when exceptions occur mid-stream.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →