# How LMOutputProcessor Streams LLM Text to the TTS Pipeline

> Discover how LMOutputProcessor streams LLM text to TTS, forwarding chunks via WebSocket and routing audio-eligible text for synthesis. Learn about its dual-channel handling.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-30

---

**LMOutputProcessor acts as a dual-channel handler that immediately forwards streaming LLM chunks to the client via WebSocket while selectively routing audio-eligible text to the TTS synthesis pipeline.**

The `LMOutputProcessor` class in the `huggingface/speech-to-speech` repository manages the critical handoff between language model generation and voice synthesis. This handler processes streaming `LLMResponseChunk` objects to ensure low-latency text visibility for users while controlling which content reaches the text-to-speech engine. Understanding how LMOutputProcessor handles streaming text from the LLM before TTS synthesis is essential for building responsive speech-to-speech applications.

## Core Workflow of LMOutputProcessor

The processor operates through a five-step pipeline defined in [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py), coordinating between the LLM output stream and downstream consumers.

### Receiving and Validating Chunks

When the `process` method receives an `LLMOut` instance, it first validates the message type. For standard streaming tokens wrapped in `LLMResponseChunk` objects, the processor proceeds to staleness validation. The optional `SpeculativeTurnTracker` guards against stale outputs through the `_turn_output_allowed` check, dropping outdated chunks before further processing.

### Dual-Channel Output Architecture

Every valid chunk triggers two parallel distribution paths. First, an `AssistantTextEvent` containing the raw text and any extracted tool calls is constructed and pushed onto the `text_output_queue`, providing immediate visual feedback to the client via WebSocket. Second, the processor evaluates whether the current turn requires audio synthesis.

### Audio Routing Logic

The `response_wants_audio` helper—imported from [`src/speech_to_speech/utils/utils.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/utils.py)—examines the LLM response metadata to determine if the chunk should be spoken. When audio is required, the processor yields a `TTSInput` object that travels downstream to the TTS handler. If the check fails, the text reaches the client without triggering voice synthesis.

## Handling Special Message Types

Beyond standard response chunks, the processor manages conversation lifecycle events. `TokenUsage` messages update statistics and emit `TokenUsageEvent` objects to the text queue without audio generation. `EndOfResponse` markers forward any generation errors as `ResponseFailedEvent` instances and yield an `EndOfResponse` signal to close the voice stream, with both paths protected by the same staleness checks.

## Practical Implementation Example

The following example demonstrates initializing the processor and feeding streaming chunks:

```python
from queue import Queue
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from speech_to_speech.pipeline.messages import LLMResponseChunk, EndOfResponse

# Queue that would be consumed by the WebSocket layer

text_q = Queue()

# Initialise the processor (no speculative turn tracker needed for this demo)

processor = LMOutputProcessor()
processor.setup(text_output_queue=text_q)

# Simulate a streaming LLM response (3 chunks)

chunks = [
    LLMResponseChunk(text="Hello", tools=[], turn_id="t1", turn_revision=0, response={"audio": True}),
    LLMResponseChunk(text=", how ", tools=[], turn_id="t1", turn_revision=0, response={"audio": True}),
    LLMResponseChunk(text="are you?", tools=[], turn_id="t1", turn_revision=0, response={"audio": True}),
]

# Feed each chunk through the processor

for chunk in chunks:
    for tts_input in processor.process(chunk):
        # The processor yields a TTSInput only when audio is required

        print("→ TTS will synthesize:", tts_input.text)

# Finally send the end‑of‑response marker

for tts_input in processor.process(EndOfResponse(turn_id="t1", turn_revision=0)):
    print("→ TTS finished")

```

This implementation places each chunk as an `AssistantTextEvent` on `text_q` for client delivery while yielding `TTSInput` objects only when the response metadata contains `"audio": True`.

## Summary

- **LMOutputProcessor** serves as the intermediary between LLM generation and TTS synthesis in the `huggingface/speech-to-speech` architecture.
- The **dual-channel design** ensures immediate text visibility via WebSocket while selectively routing audio-eligible content.
- **Staleness checks** through `SpeculativeTurnTracker` prevent processing outdated speculative turns.
- **Special handlers** manage `TokenUsage` statistics and `EndOfResponse` lifecycle events separately from standard chunks.
- Source files include [`lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/lm_output_processor.py) for core logic, [`messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/messages.py) for data structures, [`utils.py`](https://github.com/huggingface/speech-to-speech/blob/main/utils.py) for audio routing decisions, and [`speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/speculative_turns.py) for turn tracking.

## Frequently Asked Questions

### What is the purpose of the text_output_queue in LMOutputProcessor?

The `text_output_queue` serves as the WebSocket-side channel that delivers `AssistantTextEvent` objects to the client immediately as LLM tokens arrive. This queue ensures users see text feedback with minimal latency, independent of whether the content is being synthesized into speech.

### How does LMOutputProcessor decide which text gets sent to TTS?

The processor calls `response_wants_audio(lm_output.response)` to inspect the LLM response metadata. Only chunks flagged for audio production yield a `TTSInput` object downstream, allowing the system to suppress voice output for turns marked as text-only or internal tool responses.

### What happens to LLM chunks from outdated speculative turns?

Before processing any chunk, the processor queries the optional `SpeculativeTurnTracker` via `_turn_output_allowed`. Stale chunks from superseded turns are silently dropped, ensuring the client and TTS engine only receive content from the current active conversation turn.

### Can LMOutputProcessor handle non-streaming LLM responses?

Yes. While optimized for `LLMResponseChunk` streaming tokens, the `process` method also handles `TokenUsage` and `EndOfResponse` message types. These special cases update usage statistics and signal conversation termination without triggering TTS synthesis.