# LMOutputProcessor: How It Transforms LLM Output for TTS in Speech-to-Speech

> Learn how LMOutputProcessor transforms LLM output for TTS in speech-to-speech. This essential component bridges LLMs and TTS, ensuring seamless conversion for real-time speech.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-11

---

**The LMOutputProcessor bridges the language model and text-to-speech stages by converting LLM response chunks into TTS-ready input objects while simultaneously routing UI events and filtering stale speculative turns.**

In the `huggingface/speech-to-speech` repository, the **LMOutputProcessor** serves as the critical middleware between the large language model (LLM) generation stage and the text-to-speech (TTS) synthesis stage. Implemented in [`speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/lm_output_processor.py), this handler ensures that model responses are simultaneously displayed in the user interface and prepared for audio synthesis, while managing complex conversational states like speculative turns and tool calls.

## What Is the LMOutputProcessor?

The **LMOutputProcessor** is a streaming message handler that sits between the LLM and TTS components in the processing chain. According to the pipeline architecture in [`speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/s2s_pipeline.py) (around line 425), it transforms incoming `LLMResponseChunk` objects into `TTSInput` instances that downstream TTS handlers can consume directly.

Unlike a simple pass-through filter, the processor maintains dual outputs: it yields audio-ready text to the TTS stage while pushing side-channel metadata (tool calls, token usage) to a separate `text_output_queue` consumed by the WebSocket router.

## Core Responsibilities of the LMOutputProcessor

The processor handles three primary tasks in the speech-to-speech pipeline.

### Side-Channel Event Routing

When the LLM emits non-audio metadata such as **tool-call metadata** or **token-usage statistics**, the processor packages these into typed events (`AssistantTextEvent`, `TokenUsageEvent`, or `ResponseFailedEvent`). These events are pushed onto the `text_output_queue`, allowing the client UI to display intermediate text or tool results in real time without blocking the audio synthesis pipeline.

### Speculative Turn Filtering

In conversational AI systems, speculative turns occur when the pipeline re-opens a turn after user interruptions. The processor validates each chunk against a `SpeculativeTurnTracker` instance. If a chunk's `turn_id` and `turn_revision` indicate it belongs to a stale speculative turn, the processor drops it to prevent outdated audio or text from reaching the user.

### TTS Input Generation

For each `LLMResponseChunk` containing user-visible text where `response` signals that audio is desired (`response_wants_audio`), the processor yields a **TTSInput** object. This object encapsulates the raw text, language code, runtime configuration, and bookkeeping fields (turn ID, revision) required by the downstream TTS handler.

## Implementation Details

The processor exposes two primary methods: `setup()` for initialization and `process()` for message handling.

During initialization, the processor receives references to the `text_output_queue` (for UI events) and optionally a `SpeculativeTurnTracker` (for stale turn detection). These connections enable the processor to route assistant text and token usage statistics to the WebSocket interface while filtering audio-bound content.

When processing messages, the handler discriminates between content types. Text chunks bound for speech synthesis are wrapped in `TTSInput` objects and yielded to the generator. Concurrently, `AssistantTextEvent` instances are queued for immediate UI display. For `TokenUsage` or `EndOfResponse` messages, the processor updates the UI side-channel and forwards termination signals downstream so the audio pipeline can release its slot and re-enable listening.

## Code Example: Wiring the LMOutputProcessor

The following example demonstrates how to initialize the processor and handle an LLM response chunk:

```python
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from queue import SimpleQueue
from speech_to_speech.pipeline.speculative_turns import SpeculativeTurnTracker
from speech_to_speech.pipeline.messages import LLMResponseChunk, TTSInput

# Queue that the WebSocket router reads from

text_queue = SimpleQueue()

# Optional speculative-turn tracker (used for re-opening turns)

speculative_tracker = SpeculativeTurnTracker()

# Initialize the processor

lm_processor = LMOutputProcessor()
lm_processor.setup(text_output_queue=text_queue,
                   speculative_turns=speculative_tracker)

# Simulate an LLM response chunk

chunk = LLMResponseChunk(
    text="Hello, how can I help you?",
    tools=None,
    turn_id="turn-1",
    turn_revision=0,
    cancel_generation=False,
    response="assistant",          # signals audio is desired

    language_code="en-US",
    runtime_config={},
    speech_stopped_at_s=None,
)

# Process the chunk – the processor yields a TTSInput for the TTS stage

for tts_input in lm_processor.process(chunk):
    assert isinstance(tts_input, TTSInput)
    # Pass `tts_input` to your TTS handler here

    print("TTS will synthesize:", tts_input.text)

# Meanwhile the UI side-channel receives an AssistantTextEvent

ui_event = text_queue.get()
print("UI will display:", ui_event.text)

```

## Handling Token Usage and End-of-Response Signals

The processor also handles non-content messages that control pipeline state. When receiving a `TokenUsage` object, it generates a `TokenUsageEvent` for the UI queue without yielding TTS output. Similarly, `EndOfResponse` messages trigger UI updates and forward termination signals downstream:

```python
from speech_to_speech.pipeline.messages import TokenUsage

usage = TokenUsage(
    input_tokens=15,
    output_tokens=30,
    turn_id="turn-1",
    turn_revision=0,
)

# The processor forwards a TokenUsageEvent to the UI queue

for _ in lm_processor.process(usage):
    pass  # no TTS output for token usage

ui_event = text_queue.get()
print("Tokens used – in:", ui_event.input_tokens, "out:", ui_event.output_tokens)

```

## Summary

- The **LMOutputProcessor** in [`speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/lm_output_processor.py) acts as the bridge between LLM generation and TTS synthesis in the Hugging Face Speech-to-Speech pipeline.
- It extracts side-channel information (tool calls, token usage) and routes it to the UI via `AssistantTextEvent` and `TokenUsageEvent` objects placed on the `text_output_queue`.
- The processor filters stale speculative turns using a `SpeculativeTurnTracker` to prevent outdated content from reaching users.
- Valid text chunks are transformed into `TTSInput` objects containing text, language codes, and turn metadata for downstream synthesis.
- The implementation in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py) (line 425) positions this handler between the LLM and TTS stages to enable simultaneous UI updates and audio generation.

## Frequently Asked Questions

### How does the LMOutputProcessor decide which text gets sent to TTS?

The processor checks the `response` field of each `LLMResponseChunk` to determine if audio is desired (`response_wants_audio`). When the chunk contains user-visible text and the response type indicates speech synthesis is required, the processor yields a `TTSInput` object. Tool metadata and token usage statistics are routed to the UI queue instead, ensuring only clean textual content reaches the TTS engine.

### What happens to LLM chunks from interrupted speculative turns?

The processor validates each chunk against a `SpeculativeTurnTracker` instance provided during setup. If the chunk's `turn_id` and `turn_revision` indicate it belongs to a stale speculative turn that has been superseded by user interruption, the processor drops the chunk entirely. This prevents the system from synthesizing outdated audio or displaying obsolete text after the conversation state has changed.

### Where does the LMOutputProcessor fit in the speech-to-speech pipeline?

According to the handler chain in [`speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/s2s_pipeline.py) around line 425, the processor sits immediately after the LLM generation stage and before the TTS synthesis stage. The flow is: `LLM → LMOutputProcessor → TTS`. This positioning allows the processor to intercept all LLM outputs, split them into UI-facing events and audio-facing inputs, and forward them to their respective destinations.

### Can the LMOutputProcessor handle tool-calling metadata from the LLM?

Yes. When the LLM emits tool-call metadata as part of an `LLMResponseChunk`, the processor packages this information into an `AssistantTextEvent` and pushes it onto the `text_output_queue`. This allows the client UI to display tool execution status or results in real time, while the audio pipeline continues processing the conversational text separately.