# Architecture of the Speculative Turns System for Uncommitted Speech in Speech-to-Speech

> Explore the architecture of Hugging Face's speculative turns system for uncommitted speech. Learn how it efficiently manages turn revisions and late-arriving audio for seamless context.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: architecture
- Published: 2026-07-09

---

**The speculative turns system in Hugging Face's speech-to-speech uses a thread-safe `SpeculativeTurnTracker` to manage turn revisions, pending reopen candidates, and grace periods, allowing the system to retroactively assign late-arriving audio segments to prior turns without losing context.**

Real-time speech-to-speech systems must handle users who start speaking before the assistant finishes responding. To prevent "cut-off" errors and allow the assistant to reopen turns previously thought complete, the `huggingface/speech-to-speech` repository implements a speculative-turns subsystem centered on the **`SpeculativeTurnTracker`** class in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py).

## Core Data Structures of the SpeculativeTurnTracker

The `SpeculativeTurnTracker` maintains four thread-safe mutable structures protected by a `Condition` lock:

- **`self._latest_revision: OrderedDict[str, int]`** – Tracks the most recent revision index (0, 1, 2…) for each turn ID.
- **`self._committed_revision: dict[str, int]`** – Records the highest revision that has been finalized and can no longer be reopened.
- **`self._pending_reopen: dict[str, _PendingReopen]`** – Stores a single pending reopen candidate per turn, representing a speculative continuation that awaits confirmation.
- **`self._reopen_grace: dict[str, _ReopenGrace]`** – Manages active grace windows for turns that recently emitted final audio.

All public methods acquire this condition, manipulate the structures, and notify waiting threads to coordinate between the VAD layer and the downstream pipeline.

## Turn Revision Lifecycle

### Observing New Revisions

When the VAD detects speech activity, the system calls `tracker.observe(turn_id, revision)` to register the latest segment. This updates `_latest_revision` and establishes the turn's position in the revision history, making it eligible for speculative reopening if new audio arrives.

### Generating Reopen Candidates

When the VAD detects speech that might belong to an earlier turn, `VADHandler` calls `begin_reopen_candidate(turn_id, base_revision)` in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py). If the turn is uncommitted and the supplied revision matches the tracker's latest revision, the method stores a `_PendingReopen` entry and returns a new candidate revision (`base_rev + 1`). The candidate remains tentative until explicitly confirmed.

### Confirming or Cancelling Candidates

The system resolves pending candidates through two distinct paths:

- **Confirm:** `confirm_reopen_candidate(turn_id, base_rev, cand_rev)` promotes the candidate to the latest revision and removes the pending entry.
- **Cancel:** `cancel_reopen_candidate(turn_id, cand_rev)` deletes the pending entry, leaving the original revision untouched.

Both operations wake threads waiting for the pending reopen to resolve, as implemented in [`speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/speculative_turns.py) (lines 79-131).

### The Reopen Grace Window

After emitting a final audio segment, `VADHandler` invokes `start_reopen_grace(turn_id, turn_revision, timeout_seconds)` with the configured `speculative_reopen_ms` value. This creates a `_ReopenGrace` entry that keeps the turn reopenable for a short duration even before the assistant commits a response, preventing premature finalization of the user's utterance (lines 83-101).

### Committing Revisions

When the downstream LLM produces a response, the pipeline calls `commit(turn_id, revision)`. The commit succeeds only if the revision remains the latest (no newer speculative revision exists). Once committed, the revision is recorded in `_committed_revision` and can no longer be reopened, ensuring downstream components receive stable turn boundaries.

## Memory Management and Pruning

To prevent unbounded memory growth, the tracker enforces `_MAX_TRACKED_TURNS` (default 2048). The `_prune_tracked_turns` method removes oldest entries from `_latest_revision` that are neither pending nor in a grace window. This ensures active speculative turns persist even when the cap is reached, guaranteeing that a turn with a pending reopen is never dropped prematurely (lines 188-207).

## VAD Integration and End-to-End Flow

The `VADHandler` orchestrates the speculative workflow defined in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py):

1. **Speech detection** triggers `_begin_pending_reopen_if_needed` to evaluate whether to create a reopen candidate.
2. **Candidate confirmation** via `_confirm_pending_reopen` (lines 126-136) promotes the revision when the VAD is confident the audio belongs to the prior turn.
3. **New turn creation** through `_start_new_turn` (lines 176-185) generates fresh turn IDs when speech cannot be attributed to existing turns.
4. **Turn finalization** invokes `start_reopen_grace` after emitting final segments (lines 151-165), maintaining reopenability during the assistant's response generation.

Messages propagating through the pipeline carry `(turn_id, turn_revision)` pairs defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py), while [`src/speech_to_speech/pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/events.py) defines high-level `SpeechStartedEvent` and `SpeechStoppedEvent` notifications used by the VAD to inform other pipeline stages.

## Practical Implementation Examples

### Manual Tracker Usage

```python
from speech_to_speech.pipeline.speculative_turns import SpeculativeTurnTracker

tracker = SpeculativeTurnTracker()

# Initialize a turn

tracker.observe("turn_1", 0)

# User continues speaking - create reopen candidate

candidate = tracker.begin_reopen_candidate("turn_1", 0)  # Returns 1

# Confirm the reopen

tracker.confirm_reopen_candidate("turn_1", 0, candidate)

# Attempt to commit old revision (fails because revision 1 exists)

stale = tracker.commit_if_latest_after_pending_reopen("turn_1", 0)  # False

# Commit current revision (succeeds)

current = tracker.commit_if_latest_after_pending_reopen("turn_1", 1)  # True

```

### VAD-Triggered Reopen Logic

```python

# Inside VADHandler._ensure_turn_for_speech_start

if self._should_reopen_current_turn(audio_start_ms):
    reopened_turn = self._reopen_current_turn()
    if reopened_turn is not None:
        return reopened_turn  # (turn_id, new_revision, reopened=True)

```

The decision to reopen depends on the elapsed time since the last final audio (`self._last_final_audio_ms`) and the configured `speculative_reopen_ms` / `unanswered_reopen_ms` thresholds.

### Pruning with Turn Limits

```python

# Demonstrate pruning with reduced capacity

tracker = SpeculativeTurnTracker(max_tracked_turns=2)

tracker.observe("turn_1", 0)
tracker.observe("turn_2", 0)
tracker.observe("turn_3", 0)  # Triggers pruning of turn_1

assert list(tracker._latest_revision) == ["turn_2", "turn_3"]

# Note: turn_1 would be preserved if it had a pending reopen or active grace window

```

## Summary

- The **SpeculativeTurnTracker** maintains thread-safe revision history using four synchronized data structures protected by a `Condition` lock.
- **Pending reopen candidates** allow the system to tentatively assign audio to prior turns before confirmation, preventing cut-offs.
- **Reopen grace windows** prevent premature turn finalization when the assistant hasn't yet committed a response.
- **Commit operations** finalize revisions atomically only when no newer speculative versions exist.
- **Automatic pruning** bounds memory to 2048 turns while preserving active speculative states that are pending or within grace periods.

## Frequently Asked Questions

### How does the speculative turns system prevent audio cut-offs?

The system prevents cut-offs by maintaining a **reopen grace window** via `start_reopen_grace` and supporting **pending reopen candidates** through `begin_reopen_candidate` and `confirm_reopen_candidate`. These mechanisms allow the VAD to retroactively assign late-arriving audio segments to previous turns rather than forcing them into new turns, ensuring continuous conversation flow even when users speak during assistant responses.

### Is the SpeculativeTurnTracker thread-safe?

Yes, all public methods in [`speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/speculative_turns.py) acquire a `Condition` lock before manipulating the internal dictionaries. The tracker uses `self._condition` to protect `_latest_revision`, `_committed_revision`, `_pending_reopen`, and `_reopen_grace`, notifying waiting threads after state changes to coordinate between VAD processing and pipeline commitment operations.

### What happens when the system reaches the maximum tracked turns limit?

When the number of tracked turns exceeds `_MAX_TRACKED_TURNS` (default 2048), the `_prune_tracked_turns` method removes the oldest turn entries from memory. However, turns with active pending reopens or ongoing grace windows are preserved, ensuring that speculative state is never dropped prematurely even under memory pressure or during long-running conversations.

### Why are turn revisions monotonically increasing?

Revisions increment sequentially (0, 1, 2…) to establish a clear, ordered history of speech segments within a turn. This allows `commit` to verify that it is finalizing the latest revision and enables `confirm_reopen_candidate` to atomically promote speculative candidates while maintaining consistency across concurrent VAD detection and pipeline processing threads.