# LMOutputProcessor: How It Routes LLM Output to TTS and Text Events in Speech-to-Speech

> Discover how LMOutputProcessor routes LLM output to TTS and text events. This component ensures real-time sync between displayed text and generated speech in speech-to-speech.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-31

---

**The LMOutputProcessor splits LLM output into parallel text and audio streams by emitting `AssistantTextEvent` objects to a text queue while yielding `TTSInput` objects to the TTS subsystem, ensuring real-time synchronization between displayed text and generated speech.**

The **LMOutputProcessor** serves as the central routing component in the huggingface/speech-to-speech pipeline, bridging the gap between large language model inference and real-time audio generation. According to the source code in [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py), this processor ingests raw LLM chunks and intelligently distributes them between client-facing text events and the text-to-speech (TTS) subsystem. Understanding how the LMOutputProcessor splits output between TTS and text events is crucial for building responsive voice applications that maintain synchronization between visual text and spoken audio.

## What Is the LMOutputProcessor?

The **LMOutputProcessor** sits between the LLM stage and downstream consumers in the speech-to-speech architecture. Its primary responsibility is to consume objects of type `LLMResponseChunk`, `TokenUsage`, or `EndOfResponse` and route them appropriately. Unlike a simple passthrough component, it maintains turn synchronization through optional speculative turn tracking and makes routing decisions based on response metadata.

The processor operates as a dual-channel router. One channel feeds a `text_output_queue` with event objects for client consumption, while the other yields `TTSInput` instances directly into the TTS processing chain. This design decouples the audio generation pipeline from the text display logic while preserving ordering guarantees.

## How LMOutputProcessor Splits Output Between TTS and Text Events

The splitting logic follows a six-step pipeline implemented in the `process` method:

### Filtering Stale Turns with `_turn_output_allowed`

When a speculative turn tracker is present, the processor first validates that incoming chunks belong to the current turn. The `_turn_output_allowed` method discards any output that does not match the latest turn identifier. This prevents outdated chunks from reaching the client or TTS during turn revisions or interruptions.

### Emitting Text Side-Channel Events

For every valid chunk, the processor immediately places an `AssistantTextEvent` onto the `text_output_queue`. If the chunk contains tool calls, they are attached to this event. Additionally, `TokenUsageEvent` objects are emitted (lines 71–103 in [`lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/lm_output_processor.py)) to track token consumption, and `ResponseFailedEvent` objects are dispatched when errors occur (lines 121–136).

### Routing Audio via `response_wants_audio` and `TTSInput`

The decision to generate audio depends on the `response_wants_audio` helper at line 37. This function examines the LLM response metadata to determine if audio production is requested. When audio is needed, the processor extracts the textual content and yields a `TTSInput` instance (lines 139–148) containing the text, language, runtime configuration, and turn identifiers. These objects travel downstream to the TTS handlers for immediate synthesis.

### Handling End-of-Response and Errors

When an `EndOfResponse` signal arrives, the processor first emits a `ResponseFailedEvent` if the response indicates an error state (lines 81–108). It then yields the `EndOfResponse` object downstream to allow the audio pipeline to reset its state or release its processing slot. This ensures proper cleanup and resource management at the conclusion of each turn.

## LMOutputProcessor Implementation Example

The following example demonstrates setting up the processor and handling a typical LLM response chunk:

```python
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from queue import SimpleQueue

text_q = SimpleQueue()
processor = LMOutputProcessor()
processor.setup(text_output_queue=text_q, speculative_turns=None)

# Simulate a chunk coming from the LLM

chunk = LLMResponseChunk(
    text="Here is the weather forecast.",
    tools=None,
    turn_id="turn-1",
    turn_revision=0,
    response={"audio": True},
)

# Process the chunk

for tts_input in processor.process(chunk):
    # tts_input will be a TTSInput instance ready for the TTS handler

    print("Forward to TTS:", tts_input)

# Meanwhile, the text side-channel receives an AssistantTextEvent

event = text_q.get()
print("Send to client:", event)

```

For handling token usage and end-of-response signals:

```python

# Emit token usage statistics to the text queue

processor.process(TokenUsage(input_tokens=10, output_tokens=15, turn_id="turn-1"))

# Signal end of response to both text and audio channels

processor.process(EndOfResponse(turn_id="turn-1", turn_revision=0))

```

## Key Source Files and Dependencies

The implementation relies on several modules across the repository:

- **[`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py)**: Core implementation containing the `process` method and routing logic.
- **[`src/speech_to_speech/pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/events.py)**: Definitions for `AssistantTextEvent`, `TokenUsageEvent`, and `ResponseFailedEvent`.
- **[`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py)**: Message type definitions including `LLMResponseChunk`, `TTSInput`, and `EndOfResponse`.
- **[`src/speech_to_speech/pipeline/speculative_turns.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/speculative_turns.py)**: Turn-tracking logic used by `_turn_output_allowed` to filter stale output.
- **[`tests/test_lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/test_lm_output_processor.py)**: Unit tests verifying the split behavior between text and audio channels.

## Summary

- The **LMOutputProcessor** in [`lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/lm_output_processor.py) acts as a dual-channel router between the LLM and downstream consumers.
- It filters stale output using `_turn_output_allowed` when speculative turn tracking is enabled.
- Text events flow to `text_output_queue` as `AssistantTextEvent`, `TokenUsageEvent`, or `ResponseFailedEvent` objects for client delivery.
- Audio production is gated by `response_wants_audio` (line 37), which inspects response metadata.
- Clean text reaches the TTS subsystem via yielded `TTSInput` instances (lines 139–148).
- End-of-response handling (lines 81–108) ensures proper error propagation and pipeline reset.

## Frequently Asked Questions

### What is the primary purpose of the LMOutputProcessor in the speech-to-speech pipeline?

The **LMOutputProcessor** bridges the LLM inference stage with both the TTS subsystem and client-facing text channels. It receives raw LLM output chunks and splits them into two parallel streams: text events for display and `TTSInput` objects for speech synthesis. This separation allows the audio and text channels to operate independently while maintaining synchronization through shared turn identifiers.

### How does LMOutputProcessor determine whether to send text to the TTS pipeline?

The processor uses the `response_wants_audio` helper function at line 37 of [`lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/lm_output_processor.py) to inspect the LLM response metadata. If the metadata indicates that audio should be generated, the processor yields a `TTSInput` instance containing the text content. If audio is not requested, the text is only sent to the text output queue as an `AssistantTextEvent`, suppressing speech generation.

### What types of messages does the LMOutputProcessor handle?

The `process` method accepts three message types defined in [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py): `LLMResponseChunk` (containing generated text and tool calls), `TokenUsage` (tracking input and output token counts), and `EndOfResponse` (signaling completion). Each type triggers different side effects, with `TokenUsage` generating metrics events and `EndOfResponse` triggering cleanup logic.

### How does the processor prevent stale output from reaching the user?

When configured with a speculative turn tracker, the processor validates every chunk through the `_turn_output_allowed` method. This check compares the chunk's turn identifier against the current active turn. If the chunk belongs to a superseded turn, it is discarded before reaching either the text queue or TTS pipeline, ensuring users only receive content from the most recent request.