LMOutputProcessor: How It Routes LLM Output to TTS and Text Events in Speech-to-Speech
The LMOutputProcessor splits LLM output into parallel text and audio streams by emitting AssistantTextEvent objects to a text queue while yielding TTSInput objects to the TTS subsystem, ensuring real-time synchronization between displayed text and generated speech.
The LMOutputProcessor serves as the central routing component in the huggingface/speech-to-speech pipeline, bridging the gap between large language model inference and real-time audio generation. According to the source code in src/speech_to_speech/LLM/lm_output_processor.py, this processor ingests raw LLM chunks and intelligently distributes them between client-facing text events and the text-to-speech (TTS) subsystem. Understanding how the LMOutputProcessor splits output between TTS and text events is crucial for building responsive voice applications that maintain synchronization between visual text and spoken audio.
What Is the LMOutputProcessor?
The LMOutputProcessor sits between the LLM stage and downstream consumers in the speech-to-speech architecture. Its primary responsibility is to consume objects of type LLMResponseChunk, TokenUsage, or EndOfResponse and route them appropriately. Unlike a simple passthrough component, it maintains turn synchronization through optional speculative turn tracking and makes routing decisions based on response metadata.
The processor operates as a dual-channel router. One channel feeds a text_output_queue with event objects for client consumption, while the other yields TTSInput instances directly into the TTS processing chain. This design decouples the audio generation pipeline from the text display logic while preserving ordering guarantees.
How LMOutputProcessor Splits Output Between TTS and Text Events
The splitting logic follows a six-step pipeline implemented in the process method:
Filtering Stale Turns with _turn_output_allowed
When a speculative turn tracker is present, the processor first validates that incoming chunks belong to the current turn. The _turn_output_allowed method discards any output that does not match the latest turn identifier. This prevents outdated chunks from reaching the client or TTS during turn revisions or interruptions.
Emitting Text Side-Channel Events
For every valid chunk, the processor immediately places an AssistantTextEvent onto the text_output_queue. If the chunk contains tool calls, they are attached to this event. Additionally, TokenUsageEvent objects are emitted (lines 71–103 in lm_output_processor.py) to track token consumption, and ResponseFailedEvent objects are dispatched when errors occur (lines 121–136).
Routing Audio via response_wants_audio and TTSInput
The decision to generate audio depends on the response_wants_audio helper at line 37. This function examines the LLM response metadata to determine if audio production is requested. When audio is needed, the processor extracts the textual content and yields a TTSInput instance (lines 139–148) containing the text, language, runtime configuration, and turn identifiers. These objects travel downstream to the TTS handlers for immediate synthesis.
Handling End-of-Response and Errors
When an EndOfResponse signal arrives, the processor first emits a ResponseFailedEvent if the response indicates an error state (lines 81–108). It then yields the EndOfResponse object downstream to allow the audio pipeline to reset its state or release its processing slot. This ensures proper cleanup and resource management at the conclusion of each turn.
LMOutputProcessor Implementation Example
The following example demonstrates setting up the processor and handling a typical LLM response chunk:
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from queue import SimpleQueue
text_q = SimpleQueue()
processor = LMOutputProcessor()
processor.setup(text_output_queue=text_q, speculative_turns=None)
# Simulate a chunk coming from the LLM
chunk = LLMResponseChunk(
text="Here is the weather forecast.",
tools=None,
turn_id="turn-1",
turn_revision=0,
response={"audio": True},
)
# Process the chunk
for tts_input in processor.process(chunk):
# tts_input will be a TTSInput instance ready for the TTS handler
print("Forward to TTS:", tts_input)
# Meanwhile, the text side-channel receives an AssistantTextEvent
event = text_q.get()
print("Send to client:", event)
For handling token usage and end-of-response signals:
# Emit token usage statistics to the text queue
processor.process(TokenUsage(input_tokens=10, output_tokens=15, turn_id="turn-1"))
# Signal end of response to both text and audio channels
processor.process(EndOfResponse(turn_id="turn-1", turn_revision=0))
Key Source Files and Dependencies
The implementation relies on several modules across the repository:
src/speech_to_speech/LLM/lm_output_processor.py: Core implementation containing theprocessmethod and routing logic.src/speech_to_speech/pipeline/events.py: Definitions forAssistantTextEvent,TokenUsageEvent, andResponseFailedEvent.src/speech_to_speech/pipeline/messages.py: Message type definitions includingLLMResponseChunk,TTSInput, andEndOfResponse.src/speech_to_speech/pipeline/speculative_turns.py: Turn-tracking logic used by_turn_output_allowedto filter stale output.tests/test_lm_output_processor.py: Unit tests verifying the split behavior between text and audio channels.
Summary
- The LMOutputProcessor in
lm_output_processor.pyacts as a dual-channel router between the LLM and downstream consumers. - It filters stale output using
_turn_output_allowedwhen speculative turn tracking is enabled. - Text events flow to
text_output_queueasAssistantTextEvent,TokenUsageEvent, orResponseFailedEventobjects for client delivery. - Audio production is gated by
response_wants_audio(line 37), which inspects response metadata. - Clean text reaches the TTS subsystem via yielded
TTSInputinstances (lines 139–148). - End-of-response handling (lines 81–108) ensures proper error propagation and pipeline reset.
Frequently Asked Questions
What is the primary purpose of the LMOutputProcessor in the speech-to-speech pipeline?
The LMOutputProcessor bridges the LLM inference stage with both the TTS subsystem and client-facing text channels. It receives raw LLM output chunks and splits them into two parallel streams: text events for display and TTSInput objects for speech synthesis. This separation allows the audio and text channels to operate independently while maintaining synchronization through shared turn identifiers.
How does LMOutputProcessor determine whether to send text to the TTS pipeline?
The processor uses the response_wants_audio helper function at line 37 of lm_output_processor.py to inspect the LLM response metadata. If the metadata indicates that audio should be generated, the processor yields a TTSInput instance containing the text content. If audio is not requested, the text is only sent to the text output queue as an AssistantTextEvent, suppressing speech generation.
What types of messages does the LMOutputProcessor handle?
The process method accepts three message types defined in src/speech_to_speech/pipeline/messages.py: LLMResponseChunk (containing generated text and tool calls), TokenUsage (tracking input and output token counts), and EndOfResponse (signaling completion). Each type triggers different side effects, with TokenUsage generating metrics events and EndOfResponse triggering cleanup logic.
How does the processor prevent stale output from reaching the user?
When configured with a speculative turn tracker, the processor validates every chunk through the _turn_output_allowed method. This check compares the chunk's turn identifier against the current active turn. If the chunk belongs to a superseded turn, it is discarded before reaching either the text queue or TTS pipeline, ensuring users only receive content from the most recent request.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →