LMOutputProcessor in Hugging Face Speech-to-Speech: Splitting Text and Tool Calls for TTS

The LMOutputProcessor is a pipeline handler that intercepts raw LLM output, routes tool calls and generated text to the client via WebSocket events, and forwards only sanitized text (stripped of tool metadata) to the Text-to-Speech subsystem.

In the Hugging Face speech-to-speech repository, the LMOutputProcessor serves as the critical boundary between language model generation and voice synthesis. Implemented in src/speech_to_speech/LLM/lm_output_processor.py, this component ensures that tool invocations are visible to users in the UI while preventing them from contaminating the audio stream consumed by the TTS engine.

What Is the LMOutputProcessor?

The LMOutputProcessor inherits from BaseHandler[LLMOut, TTSIn] and acts as the adapter between the LLM and TTS pipeline stages. It receives three distinct message types from upstream handlers: LLMResponseChunk, TokenUsage, and EndOfResponse.

The processor maintains two output paths:

  • Text side-channel: A text_output_queue that feeds WebSocket events to the client (such as a web UI)
  • Audio pipeline: A generator yielding TTSInput objects to downstream Text-to-Speech handlers

This dual-path architecture enables the separation of concerns: the client receives complete information including tool calls, while the TTS system receives only speakable text.

How LMOutputProcessor Splits Text from Tool Calls

The splitting mechanism operates through three sequential responsibilities defined in the process method of src/speech_to_speech/LLM/lm_output_processor.py.

Filtering Stale Speculative Turns

Before processing any output, the handler validates whether the current turn is still active. Through the _turn_output_allowed method (lines 49-53), it checks the SpeculativeTurnTracker to discard responses belonging to outdated turn revisions. This prevents stale or superseded speculative generations from reaching either the client or the TTS engine.


# From lm_output_processor.py lines 49-53

if not self._turn_output_allowed(lm_output):
    # Drop this chunk – a newer revision exists

    return

Emitting Text-Side-Channel Events with Tool Calls

For every valid LLMResponseChunk, the processor constructs an AssistantTextEvent containing both the generated text and any associated tool invocations. When lm_output.tools is non-empty, these ToolCall objects are attached directly to the event before placement on the text_output_queue.


# From lm_output_processor.py lines 23-35

if self.text_output_queue is not None:
    event = AssistantTextEvent(
        text=lm_output.text,
        turn_id=lm_output.turn_id,
        turn_revision=lm_output.turn_revision,
        cancel_generation=lm_output.cancel_generation,
    )
    if lm_output.tools:
        event.tools = lm_output.tools  # Tool calls attached here

        logger.info(f"Sending to clients: text='{lm_output.text}', "
                   f"tools={[t.name for t in lm_output.tools]}")
    self.text_output_queue.put(event)

This ensures the client receives complete metadata about function calls while keeping the audio pathway separate.

Sanitizing Input for Text-to-Speech

After dispatching the side-channel event, the processor determines whether audio synthesis is required. It imports response_wants_audio from speech_to_speech.utils.utils to evaluate the LLM response metadata. When audio is desired, it yields a TTSInput object containing only the plain text—deliberately excluding the tool payload.


# From lm_output_processor.py lines 37-48

if lm_output.text and response_wants_audio(lm_output.response):
    logger.debug(f"Forwarding to TTS: '{lm_output.text}'")
    yield TTSInput(
        text=lm_output.text,              # Clean text only

        language_code=lm_output.language_code,
        runtime_config=lm_output.runtime_config,
        response=lm_output.response,
        turn_id=lm_output.turn_id,
        turn_revision=lm_output.turn_revision,
        speech_stopped_at_s=lm_output.speech_stopped_at_s,
        cancel_generation=lm_output.cancel_generation,
    )

This sanitization ensures that the TTS handler receives only natural language strings suitable for speech synthesis, avoiding the pronunciation of JSON tool schemas or function names.

Pipeline Integration

The LMOutputProcessor is instantiated within the main pipeline construction in src/speech_to_speech/s2s_pipeline.py (lines 380-426). It bridges the LLM handler and the TTS handler, receiving speculative turn tracking state to manage concurrent generation scenarios.


# From s2s_pipeline.py – pipeline construction

from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor

lm_processor = LMOutputProcessor(
    text_output_queue=text_output_queue,          # WebSocket output

    speculative_turns=speculative_turn_tracker,   # Turn state management

)

The processor sits in the handler chain: VAD → STT → TranscriptionNotifier → LM → LMOutputProcessor → TTS. This positioning allows it to act as the final filter before speech synthesis, ensuring that only approved, current-turn text enters the audio generation stage.

Summary

  • Dual-path routing: The LMOutputProcessor splits LLM output into a text side-channel (for UI display) and an audio pipeline (for TTS synthesis).
  • Tool call separation: Tool invocations travel to the client via AssistantTextEvent but are intentionally omitted from TTSInput objects.
  • Stale turn filtering: The _turn_output_allowed method prevents outdated speculative responses from contaminating either output stream.
  • Audio gating: The response_wants_audio utility function determines whether a given response should trigger speech synthesis.
  • Source location: Core logic resides in src/speech_to_speech/LLM/lm_output_processor.py with dependencies on src/speech_to_speech/pipeline/events.py for event types and src/speech_to_speech/utils/utils.py for audio decision logic.

Frequently Asked Questions

What happens when the LLM outputs both text and tool calls simultaneously?

The processor handles both components in parallel. It creates an AssistantTextEvent containing the text and attaches the tools list to the same event, then places this on the text_output_queue for the client. Separately, if response_wants_audio returns true, it yields a TTSInput containing only the text string. The tool metadata never reaches the TTS subsystem, preventing the voice engine from attempting to speak JSON function parameters.

How does the LMOutputProcessor prevent stale responses from reaching the user?

Through the _turn_output_allowed method (lines 49-53), the processor checks the SpeculativeTurnTracker to verify that the current turn_id and turn_revision match the latest active generation. If the user has interrupted or superseded the current turn with new input, old chunks are silently discarded before any output queues are modified.

Where is the decision made to generate audio for a response?

The audio generation decision is delegated to the response_wants_audio function imported from speech_to_speech.utils.utils. This utility examines the LLMResponse metadata to determine whether the assistant's output should be spoken. The LMOutputProcessor uses this boolean check to conditionally yield TTSInput objects; if false, the text is sent to the client without triggering the TTS pipeline.

Can tool calls be sent to the Text-to-Speech engine in this architecture?

No. By design, the TTSInput objects yielded by LMOutputProcessor contain only the text string, language_code, and runtime configuration. The tools attribute from the original LLMResponseChunk is only attached to AssistantTextEvent objects destined for the text output queue. This architectural boundary ensures that functional tool metadata (like JSON arguments) never enters the audio synthesis path where it would create nonsensical speech output.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →