# LMOutputProcessor in Hugging Face Speech-to-Speech: Splitting Text and Tool Calls for TTS

> Learn how Hugging Face Speech-to-Speech's LMOutputProcessor splits LLM output, separating tool calls from text before sending clean speech data to TTS. Optimize your TTS pipeline.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-10

---

**The LMOutputProcessor is a pipeline handler that intercepts raw LLM output, routes tool calls and generated text to the client via WebSocket events, and forwards only sanitized text (stripped of tool metadata) to the Text-to-Speech subsystem.**

In the Hugging Face `speech-to-speech` repository, the `LMOutputProcessor` serves as the critical boundary between language model generation and voice synthesis. Implemented in [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py), this component ensures that tool invocations are visible to users in the UI while preventing them from contaminating the audio stream consumed by the TTS engine.

## What Is the LMOutputProcessor?

The `LMOutputProcessor` inherits from `BaseHandler[LLMOut, TTSIn]` and acts as the adapter between the LLM and TTS pipeline stages. It receives three distinct message types from upstream handlers: `LLMResponseChunk`, `TokenUsage`, and `EndOfResponse`.

The processor maintains two output paths:

- **Text side-channel**: A `text_output_queue` that feeds WebSocket events to the client (such as a web UI)
- **Audio pipeline**: A generator yielding `TTSInput` objects to downstream Text-to-Speech handlers

This dual-path architecture enables the separation of concerns: the client receives complete information including tool calls, while the TTS system receives only speakable text.

## How LMOutputProcessor Splits Text from Tool Calls

The splitting mechanism operates through three sequential responsibilities defined in the `process` method of [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py).

### Filtering Stale Speculative Turns

Before processing any output, the handler validates whether the current turn is still active. Through the `_turn_output_allowed` method (lines 49-53), it checks the `SpeculativeTurnTracker` to discard responses belonging to outdated turn revisions. This prevents stale or superseded speculative generations from reaching either the client or the TTS engine.

```python

# From lm_output_processor.py lines 49-53

if not self._turn_output_allowed(lm_output):
    # Drop this chunk – a newer revision exists

    return

```

### Emitting Text-Side-Channel Events with Tool Calls

For every valid `LLMResponseChunk`, the processor constructs an `AssistantTextEvent` containing both the generated text and any associated tool invocations. When `lm_output.tools` is non-empty, these `ToolCall` objects are attached directly to the event before placement on the `text_output_queue`.

```python

# From lm_output_processor.py lines 23-35

if self.text_output_queue is not None:
    event = AssistantTextEvent(
        text=lm_output.text,
        turn_id=lm_output.turn_id,
        turn_revision=lm_output.turn_revision,
        cancel_generation=lm_output.cancel_generation,
    )
    if lm_output.tools:
        event.tools = lm_output.tools  # Tool calls attached here

        logger.info(f"Sending to clients: text='{lm_output.text}', "
                   f"tools={[t.name for t in lm_output.tools]}")
    self.text_output_queue.put(event)

```

This ensures the client receives complete metadata about function calls while keeping the audio pathway separate.

### Sanitizing Input for Text-to-Speech

After dispatching the side-channel event, the processor determines whether audio synthesis is required. It imports `response_wants_audio` from `speech_to_speech.utils.utils` to evaluate the LLM response metadata. When audio is desired, it yields a `TTSInput` object containing **only the plain text**—deliberately excluding the tool payload.

```python

# From lm_output_processor.py lines 37-48

if lm_output.text and response_wants_audio(lm_output.response):
    logger.debug(f"Forwarding to TTS: '{lm_output.text}'")
    yield TTSInput(
        text=lm_output.text,              # Clean text only

        language_code=lm_output.language_code,
        runtime_config=lm_output.runtime_config,
        response=lm_output.response,
        turn_id=lm_output.turn_id,
        turn_revision=lm_output.turn_revision,
        speech_stopped_at_s=lm_output.speech_stopped_at_s,
        cancel_generation=lm_output.cancel_generation,
    )

```

This sanitization ensures that the TTS handler receives only natural language strings suitable for speech synthesis, avoiding the pronunciation of JSON tool schemas or function names.

## Pipeline Integration

The `LMOutputProcessor` is instantiated within the main pipeline construction in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) (lines 380-426). It bridges the LLM handler and the TTS handler, receiving speculative turn tracking state to manage concurrent generation scenarios.

```python

# From s2s_pipeline.py – pipeline construction

from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor

lm_processor = LMOutputProcessor(
    text_output_queue=text_output_queue,          # WebSocket output

    speculative_turns=speculative_turn_tracker,   # Turn state management

)

```

The processor sits in the handler chain: **VAD → STT → TranscriptionNotifier → LM → LMOutputProcessor → TTS**. This positioning allows it to act as the final filter before speech synthesis, ensuring that only approved, current-turn text enters the audio generation stage.

## Summary

- **Dual-path routing**: The `LMOutputProcessor` splits LLM output into a text side-channel (for UI display) and an audio pipeline (for TTS synthesis).
- **Tool call separation**: Tool invocations travel to the client via `AssistantTextEvent` but are intentionally omitted from `TTSInput` objects.
- **Stale turn filtering**: The `_turn_output_allowed` method prevents outdated speculative responses from contaminating either output stream.
- **Audio gating**: The `response_wants_audio` utility function determines whether a given response should trigger speech synthesis.
- **Source location**: Core logic resides in [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py) with dependencies on [`src/speech_to_speech/pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/events.py) for event types and [`src/speech_to_speech/utils/utils.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/utils.py) for audio decision logic.

## Frequently Asked Questions

### What happens when the LLM outputs both text and tool calls simultaneously?

The processor handles both components in parallel. It creates an `AssistantTextEvent` containing the text and attaches the `tools` list to the same event, then places this on the `text_output_queue` for the client. Separately, if `response_wants_audio` returns true, it yields a `TTSInput` containing only the text string. The tool metadata never reaches the TTS subsystem, preventing the voice engine from attempting to speak JSON function parameters.

### How does the LMOutputProcessor prevent stale responses from reaching the user?

Through the `_turn_output_allowed` method (lines 49-53), the processor checks the `SpeculativeTurnTracker` to verify that the current `turn_id` and `turn_revision` match the latest active generation. If the user has interrupted or superseded the current turn with new input, old chunks are silently discarded before any output queues are modified.

### Where is the decision made to generate audio for a response?

The audio generation decision is delegated to the `response_wants_audio` function imported from `speech_to_speech.utils.utils`. This utility examines the `LLMResponse` metadata to determine whether the assistant's output should be spoken. The `LMOutputProcessor` uses this boolean check to conditionally yield `TTSInput` objects; if false, the text is sent to the client without triggering the TTS pipeline.

### Can tool calls be sent to the Text-to-Speech engine in this architecture?

No. By design, the `TTSInput` objects yielded by `LMOutputProcessor` contain only the `text` string, `language_code`, and runtime configuration. The `tools` attribute from the original `LLMResponseChunk` is only attached to `AssistantTextEvent` objects destined for the text output queue. This architectural boundary ensures that functional tool metadata (like JSON arguments) never enters the audio synthesis path where it would create nonsensical speech output.