Architecture of the Speculative Turns System for Uncommitted Speech in Speech-to-Speech
The speculative turns system in Hugging Face's speech-to-speech uses a thread-safe SpeculativeTurnTracker to manage turn revisions, pending reopen candidates, and grace periods, allowing the system to retroactively assign late-arriving audio segments to prior turns without losing context.
Real-time speech-to-speech systems must handle users who start speaking before the assistant finishes responding. To prevent "cut-off" errors and allow the assistant to reopen turns previously thought complete, the huggingface/speech-to-speech repository implements a speculative-turns subsystem centered on the SpeculativeTurnTracker class in src/speech_to_speech/pipeline/speculative_turns.py.
Core Data Structures of the SpeculativeTurnTracker
The SpeculativeTurnTracker maintains four thread-safe mutable structures protected by a Condition lock:
self._latest_revision: OrderedDict[str, int]– Tracks the most recent revision index (0, 1, 2…) for each turn ID.self._committed_revision: dict[str, int]– Records the highest revision that has been finalized and can no longer be reopened.self._pending_reopen: dict[str, _PendingReopen]– Stores a single pending reopen candidate per turn, representing a speculative continuation that awaits confirmation.self._reopen_grace: dict[str, _ReopenGrace]– Manages active grace windows for turns that recently emitted final audio.
All public methods acquire this condition, manipulate the structures, and notify waiting threads to coordinate between the VAD layer and the downstream pipeline.
Turn Revision Lifecycle
Observing New Revisions
When the VAD detects speech activity, the system calls tracker.observe(turn_id, revision) to register the latest segment. This updates _latest_revision and establishes the turn's position in the revision history, making it eligible for speculative reopening if new audio arrives.
Generating Reopen Candidates
When the VAD detects speech that might belong to an earlier turn, VADHandler calls begin_reopen_candidate(turn_id, base_revision) in src/speech_to_speech/VAD/vad_handler.py. If the turn is uncommitted and the supplied revision matches the tracker's latest revision, the method stores a _PendingReopen entry and returns a new candidate revision (base_rev + 1). The candidate remains tentative until explicitly confirmed.
Confirming or Cancelling Candidates
The system resolves pending candidates through two distinct paths:
- Confirm:
confirm_reopen_candidate(turn_id, base_rev, cand_rev)promotes the candidate to the latest revision and removes the pending entry. - Cancel:
cancel_reopen_candidate(turn_id, cand_rev)deletes the pending entry, leaving the original revision untouched.
Both operations wake threads waiting for the pending reopen to resolve, as implemented in speculative_turns.py (lines 79-131).
The Reopen Grace Window
After emitting a final audio segment, VADHandler invokes start_reopen_grace(turn_id, turn_revision, timeout_seconds) with the configured speculative_reopen_ms value. This creates a _ReopenGrace entry that keeps the turn reopenable for a short duration even before the assistant commits a response, preventing premature finalization of the user's utterance (lines 83-101).
Committing Revisions
When the downstream LLM produces a response, the pipeline calls commit(turn_id, revision). The commit succeeds only if the revision remains the latest (no newer speculative revision exists). Once committed, the revision is recorded in _committed_revision and can no longer be reopened, ensuring downstream components receive stable turn boundaries.
Memory Management and Pruning
To prevent unbounded memory growth, the tracker enforces _MAX_TRACKED_TURNS (default 2048). The _prune_tracked_turns method removes oldest entries from _latest_revision that are neither pending nor in a grace window. This ensures active speculative turns persist even when the cap is reached, guaranteeing that a turn with a pending reopen is never dropped prematurely (lines 188-207).
VAD Integration and End-to-End Flow
The VADHandler orchestrates the speculative workflow defined in src/speech_to_speech/VAD/vad_handler.py:
- Speech detection triggers
_begin_pending_reopen_if_neededto evaluate whether to create a reopen candidate. - Candidate confirmation via
_confirm_pending_reopen(lines 126-136) promotes the revision when the VAD is confident the audio belongs to the prior turn. - New turn creation through
_start_new_turn(lines 176-185) generates fresh turn IDs when speech cannot be attributed to existing turns. - Turn finalization invokes
start_reopen_graceafter emitting final segments (lines 151-165), maintaining reopenability during the assistant's response generation.
Messages propagating through the pipeline carry (turn_id, turn_revision) pairs defined in src/speech_to_speech/pipeline/messages.py, while src/speech_to_speech/pipeline/events.py defines high-level SpeechStartedEvent and SpeechStoppedEvent notifications used by the VAD to inform other pipeline stages.
Practical Implementation Examples
Manual Tracker Usage
from speech_to_speech.pipeline.speculative_turns import SpeculativeTurnTracker
tracker = SpeculativeTurnTracker()
# Initialize a turn
tracker.observe("turn_1", 0)
# User continues speaking - create reopen candidate
candidate = tracker.begin_reopen_candidate("turn_1", 0) # Returns 1
# Confirm the reopen
tracker.confirm_reopen_candidate("turn_1", 0, candidate)
# Attempt to commit old revision (fails because revision 1 exists)
stale = tracker.commit_if_latest_after_pending_reopen("turn_1", 0) # False
# Commit current revision (succeeds)
current = tracker.commit_if_latest_after_pending_reopen("turn_1", 1) # True
VAD-Triggered Reopen Logic
# Inside VADHandler._ensure_turn_for_speech_start
if self._should_reopen_current_turn(audio_start_ms):
reopened_turn = self._reopen_current_turn()
if reopened_turn is not None:
return reopened_turn # (turn_id, new_revision, reopened=True)
The decision to reopen depends on the elapsed time since the last final audio (self._last_final_audio_ms) and the configured speculative_reopen_ms / unanswered_reopen_ms thresholds.
Pruning with Turn Limits
# Demonstrate pruning with reduced capacity
tracker = SpeculativeTurnTracker(max_tracked_turns=2)
tracker.observe("turn_1", 0)
tracker.observe("turn_2", 0)
tracker.observe("turn_3", 0) # Triggers pruning of turn_1
assert list(tracker._latest_revision) == ["turn_2", "turn_3"]
# Note: turn_1 would be preserved if it had a pending reopen or active grace window
Summary
- The SpeculativeTurnTracker maintains thread-safe revision history using four synchronized data structures protected by a
Conditionlock. - Pending reopen candidates allow the system to tentatively assign audio to prior turns before confirmation, preventing cut-offs.
- Reopen grace windows prevent premature turn finalization when the assistant hasn't yet committed a response.
- Commit operations finalize revisions atomically only when no newer speculative versions exist.
- Automatic pruning bounds memory to 2048 turns while preserving active speculative states that are pending or within grace periods.
Frequently Asked Questions
How does the speculative turns system prevent audio cut-offs?
The system prevents cut-offs by maintaining a reopen grace window via start_reopen_grace and supporting pending reopen candidates through begin_reopen_candidate and confirm_reopen_candidate. These mechanisms allow the VAD to retroactively assign late-arriving audio segments to previous turns rather than forcing them into new turns, ensuring continuous conversation flow even when users speak during assistant responses.
Is the SpeculativeTurnTracker thread-safe?
Yes, all public methods in speculative_turns.py acquire a Condition lock before manipulating the internal dictionaries. The tracker uses self._condition to protect _latest_revision, _committed_revision, _pending_reopen, and _reopen_grace, notifying waiting threads after state changes to coordinate between VAD processing and pipeline commitment operations.
What happens when the system reaches the maximum tracked turns limit?
When the number of tracked turns exceeds _MAX_TRACKED_TURNS (default 2048), the _prune_tracked_turns method removes the oldest turn entries from memory. However, turns with active pending reopens or ongoing grace windows are preserved, ensuring that speculative state is never dropped prematurely even under memory pressure or during long-running conversations.
Why are turn revisions monotonically increasing?
Revisions increment sequentially (0, 1, 2…) to establish a clear, ordered history of speech segments within a turn. This allows commit to verify that it is finalizing the latest revision and enables confirm_reopen_candidate to atomically promote speculative candidates while maintaining consistency across concurrent VAD detection and pipeline processing threads.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →