# How to Configure Logging and Debugging for the Speech-to-Speech Pipeline

> Troubleshoot the speech-to-speech pipeline with Hugging Face. Learn about CLI verbosity, handler timing, pipeline prefixes, and torch-compile diagnostics using helpful logging and debugging tools.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**The Speech-to-Speech repository provides a flexible logging system with CLI verbosity controls, per-handler timing metrics, pipeline-index prefixes, and torch-compile diagnostics accessible through the `--log_level` flag and `setup_logger` function.**

The `huggingface/speech-to-speech` repository implements a comprehensive logging architecture designed for troubleshooting complex, real-time speech processing pipelines. Understanding these logging and debugging options enables you to isolate performance bottlenecks, track execution across multiple pipeline instances, and filter out noise from third-party dependencies.

## Command-Line Verbosity Control

The pipeline exposes a global verbosity setting through the **`ModuleArguments.log_level`** parameter defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py). This translates to the `--log_level` CLI flag, which accepts standard Python logging levels (`debug`, `info`, `warning`, `error`) and defaults to `info`.

When the pipeline starts, the `main` entry point in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) extracts `module_kwargs.log_level` and passes it to **`setup_logger`**. This function configures `logging.basicConfig` and sets the root logger level for the entire application.

```bash
python -m speech_to_speech.s2s_pipeline \
    --log_level debug \
    --mode realtime \
    --tts qwen3 \
    --stt parakeet-tdt

```

## Pipeline Context and Prefixing

For multi-pipeline deployments, the system injects pipeline identifiers into every log record using **`PipelineLogFilter`**. This custom filter, defined in [`src/speech_to_speech/pipeline/log_context.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/log_context.py), adds a `pipeline_prefix` attribute to each record based on the current context variable `pipeline_log_ctx`.

When handlers initialize, `setup_logger` attaches this filter to existing handlers. Each handler's `run` method sets the context (`pipeline_log_ctx.set(self.pipeline_index)`) before entering its processing loop, resulting in log entries automatically prefixed with `[pipeline-0]`, `[pipeline-1]`, etc.

## Per-Handler Timing Metrics

Every handler inherits from `BaseHandler`, which provides automatic performance tracking via the **`timing_log_level`** property. As implemented in [`src/speech_to_speech/baseHandler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/baseHandler.py) (lines 50-55), this defaults to `logging.DEBUG`.

The handler measures execution duration within its `process` method. If processing exceeds **`min_time_to_debug`** (0.001 seconds), the handler logs the elapsed time:

```text
2026-07-30 12:45:01,123 - [pipeline-0]speech_to_speech.VAD.vad_handler - DEBUG - VADHandler: 0.037 s
2026-07-30 12:45:01,200 - [pipeline-0]speech_to_speech.STT.whisper_stt_handler - DEBUG - WhisperSTTHandler: 0.214 s

```

This provides precise timing metrics for each pipeline stage without manual instrumentation.

## Torch Compile Diagnostics

When `log_level` is set to `debug`, the `setup_logger` function in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 34-52) automatically enables PyTorch-specific debugging by calling `torch._logging.set_logs(...)`. This emits detailed diagnostics about graph breaks, recompilations, and cuDAGraph status, which is critical for optimizing torch-compiled models in the pipeline.

## Granular Module-Level Control

The architecture uses Python's standard `logging.getLogger(__name__)` pattern, giving every component its own namespaced logger. This allows fine-grained control over individual modules while troubleshooting specific pipeline stages.

```python
import logging

# Silence the noisy pocket_tts library (as implemented in src/speech_to_speech/TTS/pocket_tts_handler.py lines 62-65)

logging.getLogger("pocket_tts").setLevel(logging.WARNING)

# Enable DEBUG only for the VAD handler

logging.getLogger("speech_to_speech.VAD.vad_handler").setLevel(logging.DEBUG)

```

Handlers like [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py), [`facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/facebookmms_handler.py), and [`chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/chatTTS_handler.py) explicitly set their third-party dependencies to `WARNING` or `DEBUG` to maintain clean output by default.

## Runtime Log Level Changes

You can programmatically adjust verbosity after pipeline initialization using the **`setup_logger`** function directly:

```python
import logging
from speech_to_speech.s2s_pipeline import setup_logger

# Switch to DEBUG for interactive troubleshooting

setup_logger("debug")

```

## Summary

- **`--log_level`** controls global verbosity via `ModuleArguments` and configures the root logger through `setup_logger` in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).
- **`PipelineLogFilter`** in [`log_context.py`](https://github.com/huggingface/speech-to-speech/blob/main/log_context.py) automatically prefixes logs with pipeline indices like `[pipeline-0]` for multi-instance tracking.
- **`BaseHandler.timing_log_level`** provides sub-millisecond timing metrics for every handler when the log level is set to DEBUG.
- **Debug mode** activates `torch._logging.set_logs` for PyTorch compilation diagnostics including graph breaks and recompilations.
- **Per-module loggers** allow suppressing noisy dependencies like `pocket_tts` or isolating specific components like VAD handlers.

## Frequently Asked Questions

### How do I enable PyTorch compilation debugging?

When you set `--log_level debug`, the `setup_logger` function automatically calls `torch._logging.set_logs(...)` to emit detailed graph-break, recompilation, and cuDAGraph diagnostics. This helps identify why `torch.compile` might be recompiling models or breaking graphs during the speech-to-speech pipeline execution.

### What does the [pipeline-0] prefix mean in my logs?

The prefix is injected by **`PipelineLogFilter`** in [`src/speech_to_speech/pipeline/log_context.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/log_context.py). It indicates which pipeline instance (0-N) generated the log record, making it possible to trace events when running multiple pipelines concurrently. The filter reads from the `pipeline_log_ctx` context variable that each handler sets before processing.

### How can I silence noisy third-party libraries while keeping debug logs for the pipeline?

Use Python's standard logging API to target specific namespaces. For example, `logging.getLogger("pocket_tts").setLevel(logging.WARNING)` suppresses INFO logs from the TTS library while keeping your pipeline handlers at DEBUG. This pattern is already applied in handlers like [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py) to keep output clean.

### Why don't I see timing information for every handler call?

Timing logs only appear when a handler's `process` method takes longer than **`min_time_to_debug`** (0.001 seconds) and your log level is set to DEBUG or the handler's specific `timing_log_level`. Fast operations below this threshold are silently skipped to reduce log volume. Check that your logger is set to DEBUG and that the handler operation exceeds the 1-millisecond threshold.