# What Is the ThreadManager and How It Orchestrates Handler Threads in Speech-to-Speech

> Discover how the ThreadManager orchestrates handler threads in Hugging Face speech-to-speech, managing worker thread lifecycles for efficient startup, monitoring, and shutdown.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-11

---

**The ThreadManager is a lightweight utility class in the Hugging Face speech-to-speech repository that coordinates the lifecycle of concurrent worker threads by standardizing how handler components are started, monitored, and gracefully shut down.**

The `speech-to-speech` library relies on multiple concurrent handlers to process real-time audio streams, from voice activity detection (VAD) to text-to-speech (TTS) generation. The ThreadManager, defined in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py), provides the essential infrastructure to manage these threads without implementing any audio processing logic itself.

## Core Lifecycle Methods in thread_manager.py

The ThreadManager class wraps a collection of handler objects that must expose a `run()` method and a `stop_event` attribute. It provides three public methods to control thread execution located in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py).

### Starting Threads with start()

The `start()` method iterates over the supplied handler objects and creates a `threading.Thread` for each using the handler's `run` method as the target. According to lines 18-24, it explicitly marks threads as non-daemon (ensuring the process waits for completion on shutdown), stores the thread objects internally, and immediately begins execution.

### Blocking Execution with wait()

When the main process must pause until all handlers complete their work, the `wait()` method calls `join()` on every managed thread. As implemented in lines 25-28, this blocks the caller until every handler thread has finished processing.

### Graceful Shutdown with stop()

The `stop()` method handles cleanup by first signaling each handler to terminate through its `stop_event` flag. Lines 30-39 implement a graceful shutdown sequence: each thread is joined with a **5-second timeout**, and a warning is logged if any thread fails to terminate within that window.

## Orchestration in the Speech-to-Speech Pipeline

The ThreadManager serves as the central coordination point when building speech-to-speech pipelines. In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the `build_pipeline` function aggregates all handlers—including the RealtimeServer, VAD, STT, language model, and TTS components—into a single list.

Lines 72-76 return a `ThreadManager` instance wrapping this collection:

```python
all_handlers: list[Any] = [realtime_server]
for unit in pool:
    all_handlers.extend(unit.handlers)
return ThreadManager(all_handlers)

```

In **Realtime** mode, the `RealtimeServer` (defined in [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py)) runs as a dedicated handler thread, while each `PipelineUnit` contributes its own set of handlers for concurrent processing. The ThreadManager ensures these components operate simultaneously while providing a unified interface for lifecycle management across local, websocket, and socket-based deployment modes.

## Practical Usage Example

To execute a complete speech-to-speech pipeline, instantiate the ThreadManager through `build_pipeline` and invoke its lifecycle methods:

```python

# Build the pipeline with all handlers

pipeline_tm = build_pipeline(
    module_kwargs,
    socket_receiver_kwargs,
    socket_sender_kwargs,
    websocket_streamer_kwargs,
    vad_handler_kwargs,
    whisper_stt_handler_kwargs,
    # ... additional handler configurations ...

    queues_and_events,
)

# Launch all threads concurrently

pipeline_tm.start()

# Block until all handlers complete

pipeline_tm.wait()

# Graceful shutdown with timeout handling

pipeline_tm.stop()

```

## Summary

- The ThreadManager in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py) provides a standardized interface for managing handler threads in the speech-to-speech pipeline.
- It creates non-daemon threads via `start()`, blocks until completion via `wait()`, and initiates graceful shutdown via `stop()`.
- The `build_pipeline` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) returns a ThreadManager instance that coordinates all pipeline components including the RealtimeServer and processing units.
- Each handler must implement a `run()` method and a `stop_event` flag to be compatible with the manager.
- The stop sequence includes a 5-second timeout per thread with warning logs for unresponsive handlers.

## Frequently Asked Questions

### What is the ThreadManager in speech-to-speech?

The ThreadManager is a utility class in the Hugging Face speech-to-speech repository that manages the lifecycle of worker threads. It handles the creation, execution, and termination of handler threads without containing any audio processing logic itself, acting purely as an orchestration layer that wraps the common "create → run → wait → stop" pattern.

### How does ThreadManager handle graceful shutdown?

When `stop()` is called, the ThreadManager sets the `stop_event` flag on each handler to request termination, then attempts to join each thread with a 5-second timeout. If a thread fails to terminate within this window, the manager logs a warning but continues the shutdown process for the remaining threads, preventing zombie processes while allowing the application to exit cleanly.

### Where is ThreadManager instantiated in the codebase?

The ThreadManager is instantiated in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) within the `build_pipeline` function. This function collects all handlers—including the RealtimeServer, VAD, STT, language model, and TTS components—and passes them to the ThreadManager constructor before returning the instance to the caller for lifecycle management.

### Why does ThreadManager use non-daemon threads?

The ThreadManager explicitly marks threads as non-daemon (as seen in lines 18-24 of [`thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/thread_manager.py)) to ensure the main process waits for all handler threads to complete before exiting. This prevents premature termination of audio processing tasks and guarantees that the cleanup operations in the `stop()` method can execute properly, avoiding data corruption or resource leaks during shutdown.