How the Speech-to-Speech Pipeline Handles Soft-Ended Turns and Turn Reopening

The pipeline uses a VADHandler with configurable grace windows and a SpeculativeTurnTracker to keep turns reopenable for short periods after speech stops, allowing users to continue speaking without starting new conversational turns.

In the huggingface/speech-to-speech repository, a turn represents a discrete chunk of audio processed as a single unit by the STT or LLM components. Rather than treating every voice activity detection (VAD) endpoint as a hard boundary, the system implements soft-ended turns that remain reopenable for configurable durations. This design prevents fragmentation of continuous user utterances while managing the latency constraints of speculative processing.

Understanding Soft-Ended Turns in Real-Time Speech Processing

When the VAD detects that speech has ended, the pipeline does not immediately commit the turn. Instead, it enters a soft-ended state where the turn stays "reopenable" for a grace period. During this window, if the user resumes speaking, the pipeline appends the new audio to the existing turn rather than initializing a new one.

The system logs soft-ended transitions in VADHandler._process_normal via:

logger.info(
    "Speech soft-ended (segment=%.0fms, active=%.0fms, turn=%s rev=%s)",
    duration_ms,
    active_speech_duration_ms,
    turn_id,
    turn_revision,
)

This behavior depends on three cooperating components defined in the source: VADHandlerArguments for configuration, VADHandler for detection logic, and SpeculativeTurnTracker for revision management.

Configuration Parameters for Turn Reopening

The VADHandlerArguments class in src/speech_to_speech/arguments_classes/vad_arguments.py exposes three critical parameters that control soft-ended turn behavior:

  • min_speech_continuation_ms (default 192 ms): Hysteresis duration that defines how long the system waits for speech continuation before considering a turn potentially complete.
  • speculative_reopen_ms (default 1000 ms): Short grace window applied while a response is actively being generated.
  • unanswered_reopen_ms (default 7000 ms): Extended cap for turns that have not yet received any assistant output, preventing premature closure of long pauses.
from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments

vad_args = VADHandlerArguments(
    min_speech_ms=384,
    min_speech_continuation_ms=192,
    speculative_reopen_ms=1000,
    unanswered_reopen_ms=7000,
)

The Decision Logic: When to Reopen a Turn

The VADHandler in src/speech_to_speech/VAD/vad_handler.py implements the core reopening logic through the _should_reopen_current_turn method. This function determines whether a soft-ended turn can accept new audio based on timing constraints and commit status:

def _should_reopen_current_turn(self, audio_start_ms: int) -> bool:
    if not self._uses_realtime_turn_handling():
        return False
    if self._current_turn_id is None or self._current_turn_revision is None or self._last_final_audio_ms is None:
        return False

    # Do not reopen a turn that has already been committed by the LLM

    is_committed = self.speculative_turns is not None and self.speculative_turns.is_committed(
        self._current_turn_id,
        self._current_turn_revision,
    )
    if is_committed:
        return False

    elapsed_ms = max(0, audio_start_ms - self._last_final_audio_ms)

    # Two windows control reopenability

    reopen_limit_ms = self.speculative_reopen_ms
    if self.speculative_turns is not None:
        reopen_limit_ms = self.unanswered_reopen_ms
    return elapsed_ms <= reopen_limit_ms

The method returns True only when three conditions are met: real-time turn handling is enabled, the turn has not been committed by the LLM, and the elapsed time since the last final audio chunk falls within the applicable grace window.

The Speculative Turn Tracker

The SpeculativeTurnTracker in src/speech_to_speech/pipeline/speculative_turns.py maintains thread-safe state for turn revisions and manages the reopen grace period. When a turn finally closes, the tracker initiates a grace window that allows the client to resume speaking without breaking conversational continuity:

self.speculative_turns.start_reopen_grace(
    turn_id,
    turn_revision,
    self.speculative_reopen_ms / 1000.0,
)

Key methods include:

  • is_latest(turn_id, revision): Verifies if the provided revision represents the current turn state.
  • begin_reopen_candidate(turn_id, revision): Registers a candidate revision when the VAD detects potential speech continuation.
  • has_pending_reopen_or_grace(turn_id, revision): Checks whether the turn remains in a reopenable state.

End-to-End Flow of Turn Reopening

The pipeline implements soft-ended turns and turn reopening through a coordinated sequence:

  1. Detection: VAD identifies speech cessation and emits a SpeechStoppedEvent while logging the soft-ended status.
  2. Evaluation: _should_reopen_current_turn checks commit status against the SpeculativeTurnTracker and measures elapsed time against speculative_reopen_ms or unanswered_reopen_ms.
  3. Pending Registration: If reopenable, _begin_pending_reopen_if_needed creates a candidate revision via speculative_turns.begin_reopen_candidate().
  4. Confirmation: Upon receiving the next audio segment, _confirm_pending_reopen finalizes the revision update, or _cancel_pending_reopen aborts if the turn committed during the interval.
  5. Grace Closure: When the LLM commits a response, start_reopen_grace begins the final grace period before permanent turn closure.

Practical Configuration Examples

To inspect turn reopenability at runtime for debugging purposes:


# Assuming handler is an initialized VADHandler instance

if handler._should_reopen_current_turn(start_ms=handler._audio_ms):
    print("Current turn is soft-ended and can be reopened")
else:
    print("Turn is closed; a new turn will be started")

For manual reopen triggering during development:


# Register a pending reopen candidate

handler._begin_pending_reopen_if_needed(audio_start_ms=handler._audio_ms)

# Simulate confirmation of the next speech segment

handler._confirm_pending_reopen()

# The handler now uses the same turn_id with an incremented revision

Summary

  • Soft-ended turns prevent premature conversational fragmentation by maintaining reopenable states after VAD endpoints.
  • The VADHandlerArguments class configures grace windows through min_speech_continuation_ms, speculative_reopen_ms, and unanswered_reopen_ms.
  • _should_reopen_current_turn in vad_handler.py implements the decision logic, checking commit status and elapsed time against the appropriate thresholds.
  • SpeculativeTurnTracker manages revision history and enforces thread-safe reopen grace periods.
  • The system dynamically adjusts reopen windows based on whether the assistant is currently generating a response or has not yet responded.

Frequently Asked Questions

What is the difference between speculative_reopen_ms and unanswered_reopen_ms?

speculative_reopen_ms (default 1000 ms) applies while the assistant is actively generating a response, providing a short window for the user to interject. unanswered_reopen_ms (default 7000 ms) applies when no response has been generated yet, allowing longer natural pauses without turn fragmentation.

How does the pipeline prevent reopening a turn that the LLM has already processed?

The _should_reopen_current_turn method queries speculative_turns.is_committed() to verify that the LLM has not yet processed the turn. If is_committed returns True, the method returns False, forcing a new turn to start regardless of timing parameters.

Can users configure the hysteresis for speech continuation independently of reopen windows?

Yes. The min_speech_continuation_ms parameter (default 192 ms) specifically controls the hysteresis for detecting speech continuation within an active or reopenable turn, operating independently of the speculative_reopen_ms and unanswered_reopen_ms grace periods used for endpoint decisions.

What happens if speech resumes after the reopen grace period expires?

If speech resumes after speculative_reopen_ms or unanswered_reopen_ms has elapsed, _should_reopen_current_turn returns False, and the VADHandler treats the new audio as the start of a fresh turn with a new turn_id and initial revision. The previous turn is considered permanently closed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →