How EOS Detection Works in Pocket-TTS: Controlling Speech Termination with eos_threshold and frames_after_eos
Pocket-TTS determines when to stop generating audio by using a learned EOS head that compares logits against eos_threshold, then appends exactly frames_after_eos additional frames to ensure natural speech termination.
End-of-speech (EOS) detection in the kyutai-labs/pocket-tts repository relies on a two-stage mechanism that balances immediate stopping criteria with a configurable safety buffer. The eos_threshold parameter controls how aggressively the model flags the end of an utterance, while frames_after_eos reserves trailing frames to prevent abrupt audio cutoffs.
The EOS Head in Flow-LM
Inside the FlowLMModel class, a dedicated linear projection layer named out_eos transforms the transformer’s final hidden state into a single logit representing the end-of-speech probability. During the forward pass in pocket_tts/models/flow_lm.py, this logit is compared directly against the user-specified eos_threshold. If the logit exceeds the threshold, the model sets the Boolean tensor is_eos to True, signaling that the current frame represents the end of the generated utterance.
# Conceptual flow inside pocket_tts/models/flow_lm.py (lines 29-31)
eos_logit = self.out_eos(hidden_state) # Linear projection
is_eos = eos_logit > eos_threshold # Boolean comparison
The method returns both the sampled latent audio representation and this is_eos flag, which propagate back to the generation controller.
Propagation to the Autoregressive Loop
The TTSModel class orchestrates generation by calling self.flow_lm._sample_next_latent, forwarding the eos_threshold argument deep into the Flow-LM sampling routine. In pocket_tts/models/tts_model.py (lines 58-66), the returned is_eos tensor feeds into the _autoregressive_generation loop, where the system tracks whether the model has begun terminating the sequence.
Detecting the First EOS Step
During generation, the autoregressive loop monitors for the initial EOS flag to establish a termination anchor point. When is_eos.item() returns True and no previous EOS step has been recorded, the current generation step index is captured in the eos_step variable. This checkpoint allows the system to distinguish between the first detection event and subsequent frames.
# Inside the generation loop in pocket_tts/models/tts_model.py (lines 761-763)
if is_eos.item() and eos_step is None:
eos_step = generation_step
Extending Generation with frames_after_eos
After capturing the first EOS step, the model does not immediately halt. Instead, it continues synthesizing for frames_after_eos additional steps to capture natural decay and prevent truncated speech. The loop breaks only when the condition generation_step >= eos_step + frames_after_eos is satisfied.
If the user does not explicitly provide frames_after_eos, the TTSModel.generate_audio method infers a small default value—typically between 1 and 3 frames—based on the input text length. This automatic calculation ensures brief utterances receive minimal padding while longer sentences retain appropriate trailing audio.
Default Values and Configuration
The pocket_tts/utils/config.py file defines DEFAULT_EOS_THRESHOLD at approximately -3.0, creating a relatively conservative threshold that requires strong confidence before triggering EOS. Users can override this default to make the model more or less sensitive to termination signals.
- Lowering the threshold (e.g., to -4.0) makes EOS detection harder, potentially generating longer audio.
- Raising the threshold (e.g., to -2.0) makes detection easier, risking premature cutoff without sufficient
frames_after_eospadding.
Practical Usage Examples
Configure EOS detection programmatically to fine-tune termination behavior:
from pocket_tts import TTSModel
# Load with a more permissive EOS threshold
model = TTSModel.load_model(eos_threshold=-2.0)
# Generate speech with 5 trailing frames after EOS detection
audio = model.generate_audio(
voice_state,
"The quick brown fox jumps over the lazy dog.",
frames_after_eos=5,
copy_state=True,
)
From the command line, adjust both parameters when generating audio:
pocket-tts generate \
--eos-threshold -1.5 \
--frames-after-eos 3 \
"Hello from Pocket-TTS!"
Summary
- EOS Head Architecture: The
FlowLMModeluses anout_eoslinear layer inpocket_tts/models/flow_lm.pyto project hidden states into logits compared againsteos_threshold. - First Detection: The autoregressive loop in
pocket_tts/models/tts_model.pycaptures the initial EOS step index whenis_eosbecomes True. - Frame Buffering: The
frames_after_eosparameter dictates how many additional frames generate after the first EOS signal, preventing audio truncation. - Defaults:
DEFAULT_EOS_THRESHOLD(≈ -3.0) and dynamicframes_after_eos(1-3 frames) provide sensible out-of-the-box behavior. - Tuning: Adjust
eos_thresholdto control sensitivity andframes_after_eosto manage trailing audio length.
Frequently Asked Questions
What happens if eos_threshold is set too high?
If eos_threshold is set too high (e.g., -1.0), the model becomes overly sensitive and may trigger EOS prematurely on ambiguous frames, causing the speech to cut off mid-word unless compensated by a large frames_after_eos value.
How does frames_after_eos affect audio quality?
The frames_after_eos parameter preserves natural speech decay and reverb tails. Setting this to 0 risks abrupt, robotic terminations, while excessively high values add unnecessary silence or breath noise at the end of utterances.
Can I disable EOS detection entirely?
While you cannot fully disable the EOS head in pocket_tts/models/flow_lm.py, setting an extremely low eos_threshold (e.g., -10.0) effectively prevents the is_eos flag from ever triggering True, causing generation to continue until hitting the maximum step limit.
Where are the default EOS values defined?
Default values reside in pocket_tts/utils/config.py, where DEFAULT_EOS_THRESHOLD is defined as approximately -3.0, and the TTSModel.generate_audio signature handles the dynamic default assignment for frames_after_eos when the parameter is omitted.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →