# Live Transcription with Parakeet TDT Controls and Parameters: A Developer’s Guide to Real-Time Speech-to-Text

> Master live transcription with Parakeet TDT controls. This guide details Hugging Face pipeline parameters like enable_live_transcription for real-time speech-to-text.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-09

---

**Enable real-time partial transcription in the `huggingface/speech-to-speech` pipeline by activating `--enable_live_transcription` and tuning `--live_transcription_update_interval` when using the Parakeet TDT backend, which leverages device-specific compute locks and `SmartProgressiveStreamingHandler` to balance responsiveness with resource contention.**

Live transcription with Parakeet TDT controls and parameters allows the speech-to-speech pipeline to emit intermediate text while the user is still speaking, rather than waiting for a complete utterance. This feature is implemented in the `ParakeetTDTSTTHandler` class and configured through dedicated command-line arguments defined in the repository’s arguments classes. The following guide breaks down the architecture, configuration options, and practical implementation steps required to deploy this capability in production environments.

## Understanding the Parakeet TDT Live Transcription Architecture

### Voice Activity Detection and Turn Types

The pipeline differentiates between **progressive** and **final** audio turns via the VAD handler. When `enable_live_transcription` is set to `True`, the system processes progressive turns through a streaming handler to generate partial outputs, while final turns trigger a complete transcription pass. This logic resides in [[`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py)](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py), specifically within the `process()` method that inspects the `AudioInItem` type and routes it to the appropriate processing branch.

### The STT Handler Initialization Flow

During pipeline construction, `get_stt_handler()` in [[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) instantiates `ParakeetTDTSTTHandler` when `module_kwargs.stt == "parakeet-tdt"`. The handler’s `setup()` method initializes the underlying model—selecting either the MLX backend for Apple Silicon (`_setup_mlx()`) or the nano-parakeet fallback for CUDA/CPU devices—based on the `device` and `compute_type` parameters supplied via [`ParakeetTDTSTTHandlerArguments`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py).

### Progressive vs. Final Processing Modes

**Progressive mode** operates with a short-timeout compute lock (0.01 seconds) to minimize latency, invoking `SmartProgressiveStreamingHandler` to emit `PartialTranscription` objects at intervals defined by `live_transcription_update_interval`. **Final mode** acquires a longer-duration lock (5 seconds) to ensure exclusive access for the complete inference pass, returning a finalized `Transcription` object and resetting the streaming state. This dual-mode architecture prevents resource starvation on shared compute devices while maintaining real-time responsiveness.

## Key Parameters and Configuration Options

The following parameters control live transcription behavior and model performance:

- **`--enable_live_transcription`** – Boolean flag that activates the progressive transcription pipeline. When omitted, the handler only processes final turns.
- **`--live_transcription_update_interval`** – Floating-point value specifying the seconds between partial transcription emissions (default typically 0.5). Lower values increase update frequency but consume more compute.
- **`--parakeet_tdt_model_name`** – Model identifier string. Defaults to `"mlx-community/parakeet-tdt-0.6b-v3"` on Apple Silicon and `"nvidia/parakeet-tdt-0.6b-v3"` for CUDA/CPU.
- **`--parakeet_tdt_device`** – Compute device selection (`"auto"`, `"mps"`, `"cuda"`, or `"cpu"`). The `"auto"` setting defers to `torch` availability checks in the handler’s initialization logic.
- **`--parakeet_tdt_compute_type`** – Precision mode (`"float16"` or `"float32"`), affecting memory bandwidth and inference speed.
- **`--parakeet_tdt_language`** – Optional ISO-639-1 language code (e.g., `en`, `de`) for constrained decoding; if omitted, the model performs automatic language detection.

## Practical Implementation: Code Examples

### Command-Line Configuration

Execute the pipeline with live transcription enabled via the module’s CLI entry point:

```bash
python -m speech_to_speech.s2s_pipeline \
    --stt parakeet-tdt \
    --enable_live_transcription \
    --live_transcription_update_interval 0.3 \
    --parakeet_tdt_device auto \
    --parakeet_tdt_compute_type float16

```

For Apple Silicon deployments requiring explicit MLX device selection:

```bash
python -m speech_to_speech.s2s_pipeline \
    --stt parakeet-tdt \
    --parakeet_tdt_device mps \
    --enable_live_transcription \
    --parakeet_tdt_model_name mlx-community/parakeet-tdt-0.6b-v3

```

### Python API Integration

Integrate the pipeline programmatically to customize queue management or embed within larger applications:

```python
from speech_to_speech.s2s_pipeline import parse_arguments, prepare_all_args, build_pipeline
from speech_to_speech.utils.thread_manager import ThreadManager
from threading import Event
from queue import Queue

# Parse and prepare configuration

args = parse_arguments()
prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Initialize inter-thread communication queues

queues_and_events = {
    "stop_event": Event(),
    "should_listen": Event(),
    "recv_audio_chunks_queue": Queue(),
    "send_audio_chunks_queue": Queue(),
    "spoken_prompt_queue": Queue(),
    "stt_output_queue": Queue(),
    "text_prompt_queue": Queue(),
    "lm_response_queue": Queue(),
    "lm_processed_queue": Queue(),
    "text_output_queue": None,
}

# Construct and launch the pipeline

pipeline: ThreadManager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    queues_and_events,
)

pipeline.start()
pipeline.wait()  # Blocks until shutdown signal

```

This pattern mirrors the `main()` implementation in [[`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) while exposing the underlying `ThreadManager` for custom lifecycle control.

## Handling Compute Resources with MLX Lock Context

On Apple Silicon devices, the `ParakeetTDTSTTHandler` utilizes `MLXLockContext` from [[`src/speech_to_speech/utils/mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py)](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py) to serialize access to the MLX runtime. The handler acquires a **short timeout** (0.01s) for progressive updates and a **long timeout** (5.0s) for final transcription, preventing race conditions when multiple threads compete for the neural engine. This locking strategy is transparent to the user but critical for stable operation when `enable_live_transcription` is active.

## Summary

- **Live transcription with Parakeet TDT controls and parameters** requires setting `--enable_live_transcription` and optionally tuning `--live_transcription_update_interval` to control partial result frequency.
- The `ParakeetTDTSTTHandler` processes **progressive** turns via `SmartProgressiveStreamingHandler` with short compute locks, while **final** turns use full model passes with extended locks.
- Configuration is managed through `ParakeetTDTSTTHandlerArguments` in [[`parakeet_tdt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/parakeet_tdt_arguments.py)](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py), supporting automatic device detection and precision selection.
- Apple Silicon deployments rely on `MLXLockContext` to prevent resource contention during concurrent streaming and final inference operations.

## Frequently Asked Questions

### How do I enable live transcription with Parakeet TDT?

Pass `--enable_live_transcription` when launching the pipeline via the CLI or set `enable_live_transcription=True` in the `module_kwargs` dictionary when using the Python API. This flag activates the progressive processing branch in `ParakeetTDTSTTHandler.process()`, allowing partial transcriptions to emit while audio capture is ongoing.

### What is the optimal `live_transcription_update_interval` value?

The default value of `0.5` seconds provides a balance between UI responsiveness and compute overhead. Reducing this to `0.2` or `0.3` seconds yields faster updates suitable for real-time captioning, but increases CPU/GPU utilization; values below `0.1` may introduce audio latency on resource-constrained devices.

### Why does the pipeline use a compute lock on Apple Silicon?

MLX (Apple’s machine learning framework) requires exclusive access to the neural engine during model inference. The `MLXLockContext` in [[`mlx_lock.py`](https://github.com/huggingface/speech-to-speech/blob/main/mlx_lock.py)](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py) enforces this serialization to prevent runtime crashes or undefined behavior when progressive and final transcription threads attempt simultaneous execution.

### Can I use Parakeet TDT live transcription without MLX?

Yes. The handler automatically falls back to the `nano-parakeet` backend on CUDA or CPU devices when MLX is unavailable or when `--parakeet_tdt_device` is explicitly set to `cuda` or `cpu`. Live transcription functions identically across backends, though latency characteristics will vary based on the underlying hardware acceleration.