# SpeculativeTurnTracker: How Hugging Face Reduces Latency in Speech-to-Speech Pipelines

> Discover how SpeculativeTurnTracker by Hugging Face slashes speech-to-speech latency. This revision manager enables speculative turns for faster audio responses, correcting errors instantly.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-10

---

**The `SpeculativeTurnTracker` is a thread-safe revision manager that enables real-time speech-to-speech systems to generate speculative audio responses from partial transcriptions while providing atomic mechanisms to correct or "reopen" turns when later audio invalidates the initial hypothesis.**

The `SpeculativeTurnTracker` is implemented in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py) within the Hugging Face speech-to-speech repository. It tracks the lifecycle of audio turns using UUID-based identifiers and monotonically increasing revision numbers, allowing the Text-to-Speech (TTS) and Large Language Model (LLM) components to begin processing immediately upon receiving the first audio chunks. By treating transcription as a streaming, revisable process rather than a batch operation, the tracker eliminates the perceptible delay typically caused by waiting for complete utterance finalization.

## What Is the SpeculativeTurnTracker?

At its core, the `SpeculativeTurnTracker` maintains an **OrderedDict** called `_latest_revision` that maps turn UUIDs to their current revision numbers. This structure, defined at lines 24-30 of [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py), provides O(1) access to the latest state of any active conversation turn. The tracker coordinates between Speech-to-Text (STT) handlers that report new audio revisions and downstream generation components that consume these revisions to produce audible responses.

The class uses a **threading.Condition** variable to synchronize access across concurrent pipeline stages, ensuring that checks for the latest revision and updates to turn state remain atomic even under heavy load.

## How Speculative Turns Reduce Perceived Latency

Traditional speech systems wait for an utterance to be fully transcribed before generating a response, creating a noticeable gap between the end of user speech and system feedback. The `SpeculativeTurnTracker` solves this by **decoupling generation from finalization**.

### Immediate Feedback Loop

When raw audio chunks arrive, the STT handler immediately calls `observe(turn_id, revision)` to register the newest revision (lines 38-48). Downstream components can then call `is_latest(turn_id, revision)` (lines 50-55) to verify if their working revision is still current. If confirmed, the TTS component begins generating audio output using only the partial transcription available at that moment, delivering audible feedback within milliseconds rather than seconds.

### Graceful Correction via Reopen

If subsequent audio chunks change the transcription meaning, the tracker initiates a **reopen** sequence. The `begin_reopen_candidate(turn_id, revision)` method (lines 50-78) creates a pending candidate revision. The system can then validate whether the new audio warrants replacing the speculative output. If validated, `confirm_reopen_candidate()` atomically promotes the new revision; if the change is minor or erroneous, `cancel_reopen_candidate()` (lines 118-131) rolls back to the previous stable state.

## Core Lifecycle Methods

The tracker manages a strict state machine for each turn, ensuring correctness while maximizing concurrency.

### Observing Audio Revisions

The `observe()` method updates the `_latest_revision` OrderedDict and triggers automatic pruning when the count exceeds `_MAX_TRACKED_TURNS` (default 2048). This prevents unbounded memory growth by removing the oldest turn entries (lines 90-104).

### Validating Speculative Work

Components check revision validity using:

- **`is_latest()`**: Blocks until the turn is not in a pending reopen state, then returns whether the queried revision matches current.
- **`try_is_latest_after_pending_reopen()`**: Non-blocking variant that returns immediately with a boolean status, allowing the pipeline to skip generation if a newer revision is already available.

### Managing Reopen Candidates

When audio continues after an initial endpoint:

1. `begin_reopen_candidate()` establishes a candidate revision and marks the turn as pending.
2. The system validates audio against this candidate.
3. `confirm_reopen_candidate(turn_id, base_revision, candidate_revision)` atomically swaps the latest revision if the base matches expectations (lines 79-116).

### Grace Periods and Final Commitment

The tracker supports a **reopen grace period** managed by `start_reopen_grace(turn_id, revision, grace_s)` (lines 83-95). During this window, the speculative output remains valid even as the candidate is evaluated. Once stable, `commit_if_latest_after_pending_reopen()` or `commit_if_latest_after_reopen_grace()` finalizes the revision, notifying any waiting threads and allowing cleanup of speculative state (lines 98-122).

## Thread Safety and Memory Management

The `SpeculativeTurnTracker` uses Python's `threading.Condition` to protect shared state across the pipeline's concurrent STT, LLM, and TTS threads. All public methods acquire the internal lock before modifying the `_latest_revision` dictionary or the `_pending_reopen` and `_reopen_grace` tracking structures.

Memory safety is enforced through automatic pruning. When tracked turns exceed 2048 entries, the `_prune_tracked_turns()` method removes the oldest items from the OrderedDict, ensuring the tracker can run indefinitely in long-running conversation sessions without leaking memory.

## Implementation Example

The following pattern demonstrates how the pipeline integrates the tracker:

```python
from speech_to_speech.pipeline.speculative_turns import SpeculativeTurnTracker

# Initialize shared tracker (typically passed to all pipeline components)

tracker = SpeculativeTurnTracker()

# STT component reports new audio for turn "abc"

tracker.observe("abc", revision=0)

# TTS component checks if revision 0 is still current before generating

if tracker.is_latest("abc", revision=0):
    audio = generate_tts(turn_text="Hello")  # Speculative generation

# Later audio arrives, triggering a reopen candidate

candidate = tracker.begin_reopen_candidate("abc", base_revision=0)
if candidate == 1:  # New revision 1 available

    # Validate and confirm the new revision

    tracker.confirm_reopen_candidate("abc", base_revision=0, candidate_revision=1)

# Commit final revision when stable

tracker.commit_if_latest_after_pending_reopen("abc", revision=1)

```

This same integration pattern appears in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 66-78), where the speculative turns instance is shared across pipeline stages.

## Summary

- The `SpeculativeTurnTracker` resides in [`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py) and manages UUID-identified conversation turns with monotonic revision numbers.
- It enables **speculative generation** by allowing TTS/LLM components to process partial transcriptions immediately while maintaining correctness through atomic reopen mechanisms.
- **Thread-safe operations** use Condition variables and OrderedDict structures to coordinate between concurrent STT handlers and generation components.
- **Memory bounds** are enforced via automatic pruning when tracked turns exceed 2048 entries.
- The **reopen lifecycle** (begin, confirm/cancel, commit) provides a graceful mechanism to replace speculative outputs when later audio changes transcription meaning.

## Frequently Asked Questions

### What happens when new audio arrives after speculative generation has already started?

When the STT handler detects additional audio for a turn, it calls `begin_reopen_candidate()` to create a pending revision. The system validates whether the new audio significantly alters the meaning. If so, `confirm_reopen_candidate()` atomically replaces the speculative response; otherwise, `cancel_reopen_candidate()` preserves the original output. This ensures users receive corrected information without perceiving the underlying latency of reprocessing.

### How does the SpeculativeTurnTracker prevent memory leaks in long conversations?

The tracker enforces a hard limit of `_MAX_TRACKED_TURNS` (default 2048). Each call to `observe()` invokes `_prune_tracked_turns()`, which removes the oldest entries from the internal OrderedDict once the threshold is exceeded. This automatic cleanup ensures memory usage remains constant regardless of conversation duration.

### What is the difference between `is_latest()` and `try_is_latest_after_pending_reopen()`?

`is_latest()` blocks the calling thread if the turn currently has a pending reopen candidate, waiting until the reopen is resolved or the grace period expires. In contrast, `try_is_latest_after_pending_reopen()` returns immediately with a boolean result, allowing non-blocking components to skip generation work when they detect that a newer revision is already available.

### Which pipeline components interact directly with the SpeculativeTurnTracker?

The STT handlers call `observe()` to report new transcriptions and `begin_reopen_candidate()` when audio continues past an initial endpoint. The LLM and TTS components call `is_latest()` or its non-blocking variant to validate revisions before generating responses, and the pipeline orchestrator calls `confirm_reopen_candidate()` or `commit_if_latest_after_pending_reopen()` to finalize turn states.