# What Is the ThreadManager in Hugging Face Speech-to-Speech?

> Learn about Hugging Face's ThreadManager, an orchestration utility that expertly manages handler threads in the Speech-to-Speech pipeline, ensuring smooth startup and graceful shutdown.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-10

---

**The ThreadManager is a lightweight orchestration utility that launches, monitors, and gracefully stops collections of handler threads in the Speech-to-Speech pipeline by managing their lifecycle through a standardized `stop_event` interface.**

The `ThreadManager` class, located in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py), serves as the central coordination mechanism for the Hugging Face `speech-to-speech` repository. It abstracts the complexity of thread lifecycle management, allowing the pipeline to run audio processing, transcription, and text-to-speech handlers concurrently while ensuring a clean shutdown when the application terminates.

## How ThreadManager Coordinates Handler Threads

The `ThreadManager` operates on a simple contract: every handler must expose a `run()` method for the thread target and a `stop_event` (`threading.Event`) that the handler checks to break out of its work loop.

### Construction and Handler Requirements

Upon instantiation, the `ThreadManager` stores references to the supplied handler sequence. Each handler must conform to the expected interface to participate in the managed lifecycle.

```python
def __init__(self, handlers: Sequence[Any]) -> None:
    self.handlers = handlers
    self.threads: List[threading.Thread] = []

```

In [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py), the `__init__` method initializes an empty list to track the `Thread` objects that will be spawned during startup.

### Starting Threads with start()

The `start()` method iterates over the handlers, creates a `threading.Thread` for each, and assigns the handler's `run` method as the thread target. Critically, threads are configured as **non-daemon** (`daemon = False`), ensuring the Python interpreter waits for them to complete before exiting.

```python
def start(self) -> None:
    for handler in self.handlers:
        thread = threading.Thread(target=handler.run, daemon=False)
        thread.start()
        self.threads.append(thread)

```

This logic appears in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py) (lines 18–24). Once called, each handler executes in its own thread, processing audio, transcriptions, or LLM responses independently.

### Graceful Shutdown with stop()

When the pipeline receives a shutdown signal, `stop()` coordinates a graceful exit. It first sets the `stop_event` on every handler to signal termination, then joins each thread with a **5-second timeout**. If a thread remains alive after the timeout, the manager logs a warning.

```python
def stop(self) -> None:
    for handler in self.handlers:
        handler.stop_event.set()
    for thread in self.threads:
        thread.join(timeout=5)
        if thread.is_alive():
            logger.warning("Thread did not stop in time")

```

This implementation in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py) (lines 29–40) prevents indefinite blocking while giving handlers a defined window to clean up resources.

### Blocking with wait()

For scenarios requiring the main thread to pause until all handlers finish—such as after `start()` but before a manual `stop()`—the `wait()` method provides indefinite blocking via `join()`.

```python
def wait(self) -> None:
    for thread in self.threads:
        thread.join()

```

This optional utility appears in [`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py) (lines 25–28).

## Integration Points in the Codebase

The `ThreadManager` acts as the central entry point for background work in two primary locations.

### Pipeline Construction

In [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), the pipeline assembles communication and processing handlers into a list, then returns a `ThreadManager` instance to control them. This design pattern decouples handler logic from thread management.

```python
from speech_to_speech.utils.thread_manager import ThreadManager

def build_manager(comms_handlers, pipeline_handlers):
    return ThreadManager([*comms_handlers, *pipeline_handlers])

```

### Real-time Server

The OpenAI-compatible real-time server in [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) utilizes the manager to run the FastAPI/uvicorn server within a dedicated thread, ensuring the server lifecycle respects the same graceful shutdown semantics as the audio handlers.

```python
from speech_to_speech.utils.thread_manager import ThreadManager

# Inside server setup

thread_manager = ThreadManager([uvicorn_handler])
thread_manager.start()

```

## Practical Implementation Example

Below is a complete, runnable example demonstrating the handler interface and `ThreadManager` coordination.

```python

# example_handler.py

import threading
import time

class ExampleHandler:
    def __init__(self):
        self.stop_event = threading.Event()

    def run(self):
        while not self.stop_event.is_set():
            print("Processing...")
            time.sleep(0.5)

# main.py

from speech_to_speech.utils.thread_manager import ThreadManager
from example_handler import ExampleHandler

handlers = [ExampleHandler() for _ in range(3)]
manager = ThreadManager(handlers)

# Launch all threads

manager.start()

# Simulate runtime

time.sleep(5)

# Signal shutdown and join

manager.stop()

```

This pattern mirrors the production implementation in `huggingface/speech-to-speech`, where handlers check `stop_event.is_set()` to exit their work loops cleanly.

## Summary

- **ThreadManager** ([`src/speech_to_speech/utils/thread_manager.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/thread_manager.py)) orchestrates the startup and shutdown of handler threads in the Speech-to-Speech pipeline.
- Handlers must implement a `run()` method and expose a `stop_event` (`threading.Event`) for coordination.
- The `start()` method spawns non-daemon threads, while `stop()` signals termination via the event and joins with a 5-second timeout.
- The class is used in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) for pipeline handlers and in [`server.py`](https://github.com/huggingface/speech-to-speech/blob/main/server.py) for the FastAPI real-time server.

## Frequently Asked Questions

### What interface must a handler implement to work with ThreadManager?

A handler must expose a `run()` method that serves as the thread entry point and a `stop_event` attribute of type `threading.Event`. The `ThreadManager` calls `run()` to start work and sets `stop_event` to request termination.

### How does ThreadManager ensure threads shut down gracefully?

The `stop()` method sets the `stop_event` for every handler, allowing each thread to break out of its processing loop. It then joins each thread with a 5-second timeout, logging a warning if any thread fails to terminate within that window.

### Where is ThreadManager instantiated in the speech-to-speech codebase?

The class is instantiated in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) to manage pipeline handlers and in [`src/speech_to_speech/api/openai_realtime/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/server.py) to run the FastAPI/uvicorn server thread.

### Why does ThreadManager use non-daemon threads?

Threads are created with `daemon = False` (the default) so that the Python interpreter waits for them to complete before shutting down. This prevents premature termination of audio processing or network communication threads while the application is still running.