How LMOutputProcessor Streams LLM Text to the TTS Pipeline

LMOutputProcessor acts as a dual-channel handler that immediately forwards streaming LLM chunks to the client via WebSocket while selectively routing audio-eligible text to the TTS synthesis pipeline.

The LMOutputProcessor class in the huggingface/speech-to-speech repository manages the critical handoff between language model generation and voice synthesis. This handler processes streaming LLMResponseChunk objects to ensure low-latency text visibility for users while controlling which content reaches the text-to-speech engine. Understanding how LMOutputProcessor handles streaming text from the LLM before TTS synthesis is essential for building responsive speech-to-speech applications.

Core Workflow of LMOutputProcessor

The processor operates through a five-step pipeline defined in src/speech_to_speech/LLM/lm_output_processor.py, coordinating between the LLM output stream and downstream consumers.

Receiving and Validating Chunks

When the process method receives an LLMOut instance, it first validates the message type. For standard streaming tokens wrapped in LLMResponseChunk objects, the processor proceeds to staleness validation. The optional SpeculativeTurnTracker guards against stale outputs through the _turn_output_allowed check, dropping outdated chunks before further processing.

Dual-Channel Output Architecture

Every valid chunk triggers two parallel distribution paths. First, an AssistantTextEvent containing the raw text and any extracted tool calls is constructed and pushed onto the text_output_queue, providing immediate visual feedback to the client via WebSocket. Second, the processor evaluates whether the current turn requires audio synthesis.

Audio Routing Logic

The response_wants_audio helper—imported from src/speech_to_speech/utils/utils.py—examines the LLM response metadata to determine if the chunk should be spoken. When audio is required, the processor yields a TTSInput object that travels downstream to the TTS handler. If the check fails, the text reaches the client without triggering voice synthesis.

Handling Special Message Types

Beyond standard response chunks, the processor manages conversation lifecycle events. TokenUsage messages update statistics and emit TokenUsageEvent objects to the text queue without audio generation. EndOfResponse markers forward any generation errors as ResponseFailedEvent instances and yield an EndOfResponse signal to close the voice stream, with both paths protected by the same staleness checks.

Practical Implementation Example

The following example demonstrates initializing the processor and feeding streaming chunks:

from queue import Queue
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from speech_to_speech.pipeline.messages import LLMResponseChunk, EndOfResponse

# Queue that would be consumed by the WebSocket layer

text_q = Queue()

# Initialise the processor (no speculative turn tracker needed for this demo)

processor = LMOutputProcessor()
processor.setup(text_output_queue=text_q)

# Simulate a streaming LLM response (3 chunks)

chunks = [
    LLMResponseChunk(text="Hello", tools=[], turn_id="t1", turn_revision=0, response={"audio": True}),
    LLMResponseChunk(text=", how ", tools=[], turn_id="t1", turn_revision=0, response={"audio": True}),
    LLMResponseChunk(text="are you?", tools=[], turn_id="t1", turn_revision=0, response={"audio": True}),
]

# Feed each chunk through the processor

for chunk in chunks:
    for tts_input in processor.process(chunk):
        # The processor yields a TTSInput only when audio is required

        print("→ TTS will synthesize:", tts_input.text)

# Finally send the end‑of‑response marker

for tts_input in processor.process(EndOfResponse(turn_id="t1", turn_revision=0)):
    print("→ TTS finished")

This implementation places each chunk as an AssistantTextEvent on text_q for client delivery while yielding TTSInput objects only when the response metadata contains "audio": True.

Summary

  • LMOutputProcessor serves as the intermediary between LLM generation and TTS synthesis in the huggingface/speech-to-speech architecture.
  • The dual-channel design ensures immediate text visibility via WebSocket while selectively routing audio-eligible content.
  • Staleness checks through SpeculativeTurnTracker prevent processing outdated speculative turns.
  • Special handlers manage TokenUsage statistics and EndOfResponse lifecycle events separately from standard chunks.
  • Source files include lm_output_processor.py for core logic, messages.py for data structures, utils.py for audio routing decisions, and speculative_turns.py for turn tracking.

Frequently Asked Questions

What is the purpose of the text_output_queue in LMOutputProcessor?

The text_output_queue serves as the WebSocket-side channel that delivers AssistantTextEvent objects to the client immediately as LLM tokens arrive. This queue ensures users see text feedback with minimal latency, independent of whether the content is being synthesized into speech.

How does LMOutputProcessor decide which text gets sent to TTS?

The processor calls response_wants_audio(lm_output.response) to inspect the LLM response metadata. Only chunks flagged for audio production yield a TTSInput object downstream, allowing the system to suppress voice output for turns marked as text-only or internal tool responses.

What happens to LLM chunks from outdated speculative turns?

Before processing any chunk, the processor queries the optional SpeculativeTurnTracker via _turn_output_allowed. Stale chunks from superseded turns are silently dropped, ensuring the client and TTS engine only receive content from the current active conversation turn.

Can LMOutputProcessor handle non-streaming LLM responses?

Yes. While optimized for LLMResponseChunk streaming tokens, the process method also handles TokenUsage and EndOfResponse message types. These special cases update usage statistics and signal conversation termination without triggering TTS synthesis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →