How Tool Calls Work Through the LLM Integration in Hugging Face Speech-to-Speech

Tool calls flow from LLM text generation through parsing and validation, ultimately emitting OpenAI Realtime-compatible function_call events via WebSocket while ensuring TTS audio receives only clean text.

The huggingface/speech-to-speech repository implements a real-time four-stage pipeline (VAD → STT → LLM → TTS) where the LLM can invoke external functions. Understanding how tool calls work through the LLM integration requires tracing the path from raw text generation to validated pipeline events.

The Four-Stage Pipeline Architecture

Speech-to-Speech processes audio through sequential stages: Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS). Tool calls originate in the LLM stage when the model generates output containing code blocks like <code>tool_name(args)</code>. These blocks are extracted, parsed, and converted into structured events before reaching the TTS stage, ensuring spoken output remains natural while clients receive actionable tool instructions.

How Tool Calls Flow Through the LLM Integration

LLM Generation and Extraction

When the LLM generates a response containing tool invocations, the system uses extract_function_calls_from_text in src/speech_to_speech/LLM/language_model.py to locate code blocks. This function separates the conversational text from executable calls, returning a tuple of clean text and raw call strings.

The LLM is instructed via system prompts (constructed by build_tool_system_prompt in tool_prompt.py) to wrap function calls in <code> tags, enabling reliable extraction without complex regex on arbitrary text.

Parsing and Validation

Raw extracted calls are processed by parse_function_call in src/speech_to_speech/LLM/tool_call/function_call.py. This module implements:

  • _split_top_level_calls: Tokenizes the string safely, handling nested parentheses and string literals to separate multiple function calls
  • FunctionToolCall: A Pydantic model representing the parsed call with name and arguments
  • Validation: FunctionToolCall.to_realtime_function_tool_call() validates the call against available FunctionTool definitions from function_tool.py, producing an openai.types.ResponseFunctionToolCall object

This validation ensures only registered tools with correct schemas are forwarded to clients, preventing malformed or unauthorized function executions.

Pipeline Event Creation

The LMOutputProcessor in src/speech_to_speech/LLM/lm_output_processor.py bridges parsing and event emission. It receives LMOutput containing both the cleaned text and validated tool list, constructing an AssistantTextEvent (defined in src/speech_to_speech/pipeline/events.py).

This event carries:

  • type: "assistant_text"
  • text: The cleaned conversational content
  • tools: A list of ResponseFunctionToolCall objects
  • turn_id: The conversation turn identifier

The processor places this event on the internal text_output_queue for downstream consumption.

WebSocket Event Emission

The WebSocketRouter in src/speech_to_speech/connections/websocket_streamer.py consumes AssistantTextEvent objects from the queue. For each event, it:

  1. Sends the transcript text to the client for display
  2. Iterates through the tools list and emits separate tool-call events formatted as OpenAI Realtime protocol messages (type: "function_call")
  3. Increments usage counters (st.response_usage.tool_calls += len(event.tools)) in src/speech_to_speech/api/openai_realtime/handlers/response.py

Downstream TTS components receive only the cleaned text from the event, ensuring the voice output is not polluted by tool-call syntax or JSON arguments.

Implementation Examples

Configuring Tool Prompts

Configure the LLM to emit tool calls by building a system prompt with available function definitions:

from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool

search_tool = FunctionTool(
    name="web_search",
    description="Search the internet for current information",
    parameters={
        "type": "object",
        "properties": {"query": {"type": "string"}},
        "required": ["query"]
    }
)

system_prompt = build_tool_system_prompt([search_tool])

# LLM will now generate calls like: <code>web_search(query="weather today")</code>

Extracting and Validating Calls

Process raw LLM output to extract and validate function calls:

from speech_to_speech.LLM.language_model import extract_function_calls_from_text

llm_response = 'Let me search that for you.\n<code>web_search(query="Python documentation")</code>'
clean_text, tool_calls = extract_function_calls_from_text(llm_response)

# Convert to OpenAI-compatible format

if tool_calls:
    realtime_call = tool_calls[0].to_realtime_function_tool_call([search_tool])
    # realtime_call contains name, arguments, and call_id

Building Pipeline Events

Create the event object that flows through the pipeline:

from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from speech_to_speech.pipeline.events import AssistantTextEvent

processor = LMOutputProcessor(config={...})
event = AssistantTextEvent(
    text=clean_text,
    tools=[realtime_call],  # List[ResponseFunctionToolCall]

    turn_id="turn_42"
)

processor.text_output_queue.put(event)

Handling Events in the WebSocket Router

The router separates text from tool calls when transmitting to clients:


# Simplified logic from websocket_streamer.py

if isinstance(event, AssistantTextEvent):
    await websocket.send_text(event.text)
    
    for tool in event.tools:
        await websocket.send_json({
            "type": "function_call",
            "name": tool.name,
            "arguments": tool.arguments,
            "call_id": tool.call_id
        })

Summary

  • Tool calls are wrapped in <code> blocks by the LLM and extracted via extract_function_calls_from_text in language_model.py
  • Parsing occurs in function_call.py using _split_top_level_calls and validated against FunctionTool schemas to produce ResponseFunctionToolCall objects
  • Pipeline events are created by LMOutputProcessor as AssistantTextEvent objects containing both text and validated tools
  • WebSocket emission in websocket_streamer.py sends clean text to TTS while emitting separate function_call events to clients over the OpenAI Realtime protocol
  • Usage tracking increments tool_calls counters in response.py for monitoring and billing

Frequently Asked Questions

How are tool calls formatted in the LLM output?

The LLM wraps function invocations in XML-style <code> tags (e.g., <code>search(query="news")</code>). The extract_function_calls_from_text function in language_model.py specifically looks for these tags to separate executable code from conversational text, enabling reliable extraction without complex natural language parsing.

What prevents tool call syntax from being spoken by the TTS system?

The LMOutputProcessor creates AssistantTextEvent objects that separate the cleaned text field (sent to TTS) from the tools list (sent to the client). When the WebSocketRouter processes the event, it forwards only the text content to the TTS stage, ensuring the voice output contains natural language while the JSON tool arguments are transmitted silently to the client application.

How does the system handle multiple tool calls in a single response?

The _split_top_level_calls helper in function_call.py tokenizes the raw string to identify individual function calls separated by commas or newlines within the <code> block. Each call is parsed into a separate FunctionToolCall object, and the AssistantTextEvent carries a list of tools that the WebSocketRouter iterates through, emitting discrete function_call events for each tool invocation.

What protocol does the WebSocket router use for tool call events?

The WebSocketRouter in websocket_streamer.py emits tool calls using the OpenAI Realtime API protocol, sending JSON messages with "type": "function_call" containing the name, arguments, and call_id fields. This ensures compatibility with OpenAI's Realtime client SDKs while maintaining the flexibility to handle custom tool definitions validated against the FunctionTool schema.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →