LMOutputProcessor: How It Transforms LLM Output for TTS in Speech-to-Speech

The LMOutputProcessor bridges the language model and text-to-speech stages by converting LLM response chunks into TTS-ready input objects while simultaneously routing UI events and filtering stale speculative turns.

In the huggingface/speech-to-speech repository, the LMOutputProcessor serves as the critical middleware between the large language model (LLM) generation stage and the text-to-speech (TTS) synthesis stage. Implemented in speech_to_speech/LLM/lm_output_processor.py, this handler ensures that model responses are simultaneously displayed in the user interface and prepared for audio synthesis, while managing complex conversational states like speculative turns and tool calls.

What Is the LMOutputProcessor?

The LMOutputProcessor is a streaming message handler that sits between the LLM and TTS components in the processing chain. According to the pipeline architecture in speech_to_speech/s2s_pipeline.py (around line 425), it transforms incoming LLMResponseChunk objects into TTSInput instances that downstream TTS handlers can consume directly.

Unlike a simple pass-through filter, the processor maintains dual outputs: it yields audio-ready text to the TTS stage while pushing side-channel metadata (tool calls, token usage) to a separate text_output_queue consumed by the WebSocket router.

Core Responsibilities of the LMOutputProcessor

The processor handles three primary tasks in the speech-to-speech pipeline.

Side-Channel Event Routing

When the LLM emits non-audio metadata such as tool-call metadata or token-usage statistics, the processor packages these into typed events (AssistantTextEvent, TokenUsageEvent, or ResponseFailedEvent). These events are pushed onto the text_output_queue, allowing the client UI to display intermediate text or tool results in real time without blocking the audio synthesis pipeline.

Speculative Turn Filtering

In conversational AI systems, speculative turns occur when the pipeline re-opens a turn after user interruptions. The processor validates each chunk against a SpeculativeTurnTracker instance. If a chunk's turn_id and turn_revision indicate it belongs to a stale speculative turn, the processor drops it to prevent outdated audio or text from reaching the user.

TTS Input Generation

For each LLMResponseChunk containing user-visible text where response signals that audio is desired (response_wants_audio), the processor yields a TTSInput object. This object encapsulates the raw text, language code, runtime configuration, and bookkeeping fields (turn ID, revision) required by the downstream TTS handler.

Implementation Details

The processor exposes two primary methods: setup() for initialization and process() for message handling.

During initialization, the processor receives references to the text_output_queue (for UI events) and optionally a SpeculativeTurnTracker (for stale turn detection). These connections enable the processor to route assistant text and token usage statistics to the WebSocket interface while filtering audio-bound content.

When processing messages, the handler discriminates between content types. Text chunks bound for speech synthesis are wrapped in TTSInput objects and yielded to the generator. Concurrently, AssistantTextEvent instances are queued for immediate UI display. For TokenUsage or EndOfResponse messages, the processor updates the UI side-channel and forwards termination signals downstream so the audio pipeline can release its slot and re-enable listening.

Code Example: Wiring the LMOutputProcessor

The following example demonstrates how to initialize the processor and handle an LLM response chunk:

from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from queue import SimpleQueue
from speech_to_speech.pipeline.speculative_turns import SpeculativeTurnTracker
from speech_to_speech.pipeline.messages import LLMResponseChunk, TTSInput

# Queue that the WebSocket router reads from

text_queue = SimpleQueue()

# Optional speculative-turn tracker (used for re-opening turns)

speculative_tracker = SpeculativeTurnTracker()

# Initialize the processor

lm_processor = LMOutputProcessor()
lm_processor.setup(text_output_queue=text_queue,
                   speculative_turns=speculative_tracker)

# Simulate an LLM response chunk

chunk = LLMResponseChunk(
    text="Hello, how can I help you?",
    tools=None,
    turn_id="turn-1",
    turn_revision=0,
    cancel_generation=False,
    response="assistant",          # signals audio is desired

    language_code="en-US",
    runtime_config={},
    speech_stopped_at_s=None,
)

# Process the chunk – the processor yields a TTSInput for the TTS stage

for tts_input in lm_processor.process(chunk):
    assert isinstance(tts_input, TTSInput)
    # Pass `tts_input` to your TTS handler here

    print("TTS will synthesize:", tts_input.text)

# Meanwhile the UI side-channel receives an AssistantTextEvent

ui_event = text_queue.get()
print("UI will display:", ui_event.text)

Handling Token Usage and End-of-Response Signals

The processor also handles non-content messages that control pipeline state. When receiving a TokenUsage object, it generates a TokenUsageEvent for the UI queue without yielding TTS output. Similarly, EndOfResponse messages trigger UI updates and forward termination signals downstream:

from speech_to_speech.pipeline.messages import TokenUsage

usage = TokenUsage(
    input_tokens=15,
    output_tokens=30,
    turn_id="turn-1",
    turn_revision=0,
)

# The processor forwards a TokenUsageEvent to the UI queue

for _ in lm_processor.process(usage):
    pass  # no TTS output for token usage

ui_event = text_queue.get()
print("Tokens used – in:", ui_event.input_tokens, "out:", ui_event.output_tokens)

Summary

  • The LMOutputProcessor in speech_to_speech/LLM/lm_output_processor.py acts as the bridge between LLM generation and TTS synthesis in the Hugging Face Speech-to-Speech pipeline.
  • It extracts side-channel information (tool calls, token usage) and routes it to the UI via AssistantTextEvent and TokenUsageEvent objects placed on the text_output_queue.
  • The processor filters stale speculative turns using a SpeculativeTurnTracker to prevent outdated content from reaching users.
  • Valid text chunks are transformed into TTSInput objects containing text, language codes, and turn metadata for downstream synthesis.
  • The implementation in s2s_pipeline.py (line 425) positions this handler between the LLM and TTS stages to enable simultaneous UI updates and audio generation.

Frequently Asked Questions

How does the LMOutputProcessor decide which text gets sent to TTS?

The processor checks the response field of each LLMResponseChunk to determine if audio is desired (response_wants_audio). When the chunk contains user-visible text and the response type indicates speech synthesis is required, the processor yields a TTSInput object. Tool metadata and token usage statistics are routed to the UI queue instead, ensuring only clean textual content reaches the TTS engine.

What happens to LLM chunks from interrupted speculative turns?

The processor validates each chunk against a SpeculativeTurnTracker instance provided during setup. If the chunk's turn_id and turn_revision indicate it belongs to a stale speculative turn that has been superseded by user interruption, the processor drops the chunk entirely. This prevents the system from synthesizing outdated audio or displaying obsolete text after the conversation state has changed.

Where does the LMOutputProcessor fit in the speech-to-speech pipeline?

According to the handler chain in speech_to_speech/s2s_pipeline.py around line 425, the processor sits immediately after the LLM generation stage and before the TTS synthesis stage. The flow is: LLM → LMOutputProcessor → TTS. This positioning allows the processor to intercept all LLM outputs, split them into UI-facing events and audio-facing inputs, and forward them to their respective destinations.

Can the LMOutputProcessor handle tool-calling metadata from the LLM?

Yes. When the LLM emits tool-call metadata as part of an LLMResponseChunk, the processor packages this information into an AssistantTextEvent and pushes it onto the text_output_queue. This allows the client UI to display tool execution status or results in real time, while the audio pipeline continues processing the conversational text separately.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →