How LMOutputProcessor Streams LLM Text to the TTS Pipeline
LMOutputProcessor acts as a dual-channel handler that immediately forwards streaming LLM chunks to the client via WebSocket while selectively routing audio-eligible text to the TTS synthesis pipeline.
The LMOutputProcessor class in the huggingface/speech-to-speech repository manages the critical handoff between language model generation and voice synthesis. This handler processes streaming LLMResponseChunk objects to ensure low-latency text visibility for users while controlling which content reaches the text-to-speech engine. Understanding how LMOutputProcessor handles streaming text from the LLM before TTS synthesis is essential for building responsive speech-to-speech applications.
Core Workflow of LMOutputProcessor
The processor operates through a five-step pipeline defined in src/speech_to_speech/LLM/lm_output_processor.py, coordinating between the LLM output stream and downstream consumers.
Receiving and Validating Chunks
When the process method receives an LLMOut instance, it first validates the message type. For standard streaming tokens wrapped in LLMResponseChunk objects, the processor proceeds to staleness validation. The optional SpeculativeTurnTracker guards against stale outputs through the _turn_output_allowed check, dropping outdated chunks before further processing.
Dual-Channel Output Architecture
Every valid chunk triggers two parallel distribution paths. First, an AssistantTextEvent containing the raw text and any extracted tool calls is constructed and pushed onto the text_output_queue, providing immediate visual feedback to the client via WebSocket. Second, the processor evaluates whether the current turn requires audio synthesis.
Audio Routing Logic
The response_wants_audio helper—imported from src/speech_to_speech/utils/utils.py—examines the LLM response metadata to determine if the chunk should be spoken. When audio is required, the processor yields a TTSInput object that travels downstream to the TTS handler. If the check fails, the text reaches the client without triggering voice synthesis.
Handling Special Message Types
Beyond standard response chunks, the processor manages conversation lifecycle events. TokenUsage messages update statistics and emit TokenUsageEvent objects to the text queue without audio generation. EndOfResponse markers forward any generation errors as ResponseFailedEvent instances and yield an EndOfResponse signal to close the voice stream, with both paths protected by the same staleness checks.
Practical Implementation Example
The following example demonstrates initializing the processor and feeding streaming chunks:
from queue import Queue
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from speech_to_speech.pipeline.messages import LLMResponseChunk, EndOfResponse
# Queue that would be consumed by the WebSocket layer
text_q = Queue()
# Initialise the processor (no speculative turn tracker needed for this demo)
processor = LMOutputProcessor()
processor.setup(text_output_queue=text_q)
# Simulate a streaming LLM response (3 chunks)
chunks = [
LLMResponseChunk(text="Hello", tools=[], turn_id="t1", turn_revision=0, response={"audio": True}),
LLMResponseChunk(text=", how ", tools=[], turn_id="t1", turn_revision=0, response={"audio": True}),
LLMResponseChunk(text="are you?", tools=[], turn_id="t1", turn_revision=0, response={"audio": True}),
]
# Feed each chunk through the processor
for chunk in chunks:
for tts_input in processor.process(chunk):
# The processor yields a TTSInput only when audio is required
print("→ TTS will synthesize:", tts_input.text)
# Finally send the end‑of‑response marker
for tts_input in processor.process(EndOfResponse(turn_id="t1", turn_revision=0)):
print("→ TTS finished")
This implementation places each chunk as an AssistantTextEvent on text_q for client delivery while yielding TTSInput objects only when the response metadata contains "audio": True.
Summary
- LMOutputProcessor serves as the intermediary between LLM generation and TTS synthesis in the
huggingface/speech-to-speecharchitecture. - The dual-channel design ensures immediate text visibility via WebSocket while selectively routing audio-eligible content.
- Staleness checks through
SpeculativeTurnTrackerprevent processing outdated speculative turns. - Special handlers manage
TokenUsagestatistics andEndOfResponselifecycle events separately from standard chunks. - Source files include
lm_output_processor.pyfor core logic,messages.pyfor data structures,utils.pyfor audio routing decisions, andspeculative_turns.pyfor turn tracking.
Frequently Asked Questions
What is the purpose of the text_output_queue in LMOutputProcessor?
The text_output_queue serves as the WebSocket-side channel that delivers AssistantTextEvent objects to the client immediately as LLM tokens arrive. This queue ensures users see text feedback with minimal latency, independent of whether the content is being synthesized into speech.
How does LMOutputProcessor decide which text gets sent to TTS?
The processor calls response_wants_audio(lm_output.response) to inspect the LLM response metadata. Only chunks flagged for audio production yield a TTSInput object downstream, allowing the system to suppress voice output for turns marked as text-only or internal tool responses.
What happens to LLM chunks from outdated speculative turns?
Before processing any chunk, the processor queries the optional SpeculativeTurnTracker via _turn_output_allowed. Stale chunks from superseded turns are silently dropped, ensuring the client and TTS engine only receive content from the current active conversation turn.
Can LMOutputProcessor handle non-streaming LLM responses?
Yes. While optimized for LLMResponseChunk streaming tokens, the process method also handles TokenUsage and EndOfResponse message types. These special cases update usage statistics and signal conversation termination without triggering TTS synthesis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →