Live Transcription with Parakeet TDT Controls and Parameters: A Developer’s Guide to Real-Time Speech-to-Text
Enable real-time partial transcription in the huggingface/speech-to-speech pipeline by activating --enable_live_transcription and tuning --live_transcription_update_interval when using the Parakeet TDT backend, which leverages device-specific compute locks and SmartProgressiveStreamingHandler to balance responsiveness with resource contention.
Live transcription with Parakeet TDT controls and parameters allows the speech-to-speech pipeline to emit intermediate text while the user is still speaking, rather than waiting for a complete utterance. This feature is implemented in the ParakeetTDTSTTHandler class and configured through dedicated command-line arguments defined in the repository’s arguments classes. The following guide breaks down the architecture, configuration options, and practical implementation steps required to deploy this capability in production environments.
Understanding the Parakeet TDT Live Transcription Architecture
Voice Activity Detection and Turn Types
The pipeline differentiates between progressive and final audio turns via the VAD handler. When enable_live_transcription is set to True, the system processes progressive turns through a streaming handler to generate partial outputs, while final turns trigger a complete transcription pass. This logic resides in [src/speech_to_speech/STT/parakeet_tdt_handler.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py), specifically within the process() method that inspects the AudioInItem type and routes it to the appropriate processing branch.
The STT Handler Initialization Flow
During pipeline construction, get_stt_handler() in [src/speech_to_speech/s2s_pipeline.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) instantiates ParakeetTDTSTTHandler when module_kwargs.stt == "parakeet-tdt". The handler’s setup() method initializes the underlying model—selecting either the MLX backend for Apple Silicon (_setup_mlx()) or the nano-parakeet fallback for CUDA/CPU devices—based on the device and compute_type parameters supplied via ParakeetTDTSTTHandlerArguments.
Progressive vs. Final Processing Modes
Progressive mode operates with a short-timeout compute lock (0.01 seconds) to minimize latency, invoking SmartProgressiveStreamingHandler to emit PartialTranscription objects at intervals defined by live_transcription_update_interval. Final mode acquires a longer-duration lock (5 seconds) to ensure exclusive access for the complete inference pass, returning a finalized Transcription object and resetting the streaming state. This dual-mode architecture prevents resource starvation on shared compute devices while maintaining real-time responsiveness.
Key Parameters and Configuration Options
The following parameters control live transcription behavior and model performance:
--enable_live_transcription– Boolean flag that activates the progressive transcription pipeline. When omitted, the handler only processes final turns.--live_transcription_update_interval– Floating-point value specifying the seconds between partial transcription emissions (default typically 0.5). Lower values increase update frequency but consume more compute.--parakeet_tdt_model_name– Model identifier string. Defaults to"mlx-community/parakeet-tdt-0.6b-v3"on Apple Silicon and"nvidia/parakeet-tdt-0.6b-v3"for CUDA/CPU.--parakeet_tdt_device– Compute device selection ("auto","mps","cuda", or"cpu"). The"auto"setting defers totorchavailability checks in the handler’s initialization logic.--parakeet_tdt_compute_type– Precision mode ("float16"or"float32"), affecting memory bandwidth and inference speed.--parakeet_tdt_language– Optional ISO-639-1 language code (e.g.,en,de) for constrained decoding; if omitted, the model performs automatic language detection.
Practical Implementation: Code Examples
Command-Line Configuration
Execute the pipeline with live transcription enabled via the module’s CLI entry point:
python -m speech_to_speech.s2s_pipeline \
--stt parakeet-tdt \
--enable_live_transcription \
--live_transcription_update_interval 0.3 \
--parakeet_tdt_device auto \
--parakeet_tdt_compute_type float16
For Apple Silicon deployments requiring explicit MLX device selection:
python -m speech_to_speech.s2s_pipeline \
--stt parakeet-tdt \
--parakeet_tdt_device mps \
--enable_live_transcription \
--parakeet_tdt_model_name mlx-community/parakeet-tdt-0.6b-v3
Python API Integration
Integrate the pipeline programmatically to customize queue management or embed within larger applications:
from speech_to_speech.s2s_pipeline import parse_arguments, prepare_all_args, build_pipeline
from speech_to_speech.utils.thread_manager import ThreadManager
from threading import Event
from queue import Queue
# Parse and prepare configuration
args = parse_arguments()
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
)
# Initialize inter-thread communication queues
queues_and_events = {
"stop_event": Event(),
"should_listen": Event(),
"recv_audio_chunks_queue": Queue(),
"send_audio_chunks_queue": Queue(),
"spoken_prompt_queue": Queue(),
"stt_output_queue": Queue(),
"text_prompt_queue": Queue(),
"lm_response_queue": Queue(),
"lm_processed_queue": Queue(),
"text_output_queue": None,
}
# Construct and launch the pipeline
pipeline: ThreadManager = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.websocket_streamer_kwargs,
args.vad_handler_kwargs,
args.whisper_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
queues_and_events,
)
pipeline.start()
pipeline.wait() # Blocks until shutdown signal
This pattern mirrors the main() implementation in [s2s_pipeline.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) while exposing the underlying ThreadManager for custom lifecycle control.
Handling Compute Resources with MLX Lock Context
On Apple Silicon devices, the ParakeetTDTSTTHandler utilizes MLXLockContext from [src/speech_to_speech/utils/mlx_lock.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py) to serialize access to the MLX runtime. The handler acquires a short timeout (0.01s) for progressive updates and a long timeout (5.0s) for final transcription, preventing race conditions when multiple threads compete for the neural engine. This locking strategy is transparent to the user but critical for stable operation when enable_live_transcription is active.
Summary
- Live transcription with Parakeet TDT controls and parameters requires setting
--enable_live_transcriptionand optionally tuning--live_transcription_update_intervalto control partial result frequency. - The
ParakeetTDTSTTHandlerprocesses progressive turns viaSmartProgressiveStreamingHandlerwith short compute locks, while final turns use full model passes with extended locks. - Configuration is managed through
ParakeetTDTSTTHandlerArgumentsin [parakeet_tdt_arguments.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py), supporting automatic device detection and precision selection. - Apple Silicon deployments rely on
MLXLockContextto prevent resource contention during concurrent streaming and final inference operations.
Frequently Asked Questions
How do I enable live transcription with Parakeet TDT?
Pass --enable_live_transcription when launching the pipeline via the CLI or set enable_live_transcription=True in the module_kwargs dictionary when using the Python API. This flag activates the progressive processing branch in ParakeetTDTSTTHandler.process(), allowing partial transcriptions to emit while audio capture is ongoing.
What is the optimal live_transcription_update_interval value?
The default value of 0.5 seconds provides a balance between UI responsiveness and compute overhead. Reducing this to 0.2 or 0.3 seconds yields faster updates suitable for real-time captioning, but increases CPU/GPU utilization; values below 0.1 may introduce audio latency on resource-constrained devices.
Why does the pipeline use a compute lock on Apple Silicon?
MLX (Apple’s machine learning framework) requires exclusive access to the neural engine during model inference. The MLXLockContext in [mlx_lock.py](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/mlx_lock.py) enforces this serialization to prevent runtime crashes or undefined behavior when progressive and final transcription threads attempt simultaneous execution.
Can I use Parakeet TDT live transcription without MLX?
Yes. The handler automatically falls back to the nano-parakeet backend on CUDA or CPU devices when MLX is unavailable or when --parakeet_tdt_device is explicitly set to cuda or cpu. Live transcription functions identically across backends, though latency characteristics will vary based on the underlying hardware acceleration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →