How Tool Calls Work Through the LLM Integration in Hugging Face Speech-to-Speech
Tool calls flow from LLM text generation through parsing and validation, ultimately emitting OpenAI Realtime-compatible function_call events via WebSocket while ensuring TTS audio receives only clean text.
The huggingface/speech-to-speech repository implements a real-time four-stage pipeline (VAD → STT → LLM → TTS) where the LLM can invoke external functions. Understanding how tool calls work through the LLM integration requires tracing the path from raw text generation to validated pipeline events.
The Four-Stage Pipeline Architecture
Speech-to-Speech processes audio through sequential stages: Voice Activity Detection (VAD), Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS). Tool calls originate in the LLM stage when the model generates output containing code blocks like <code>tool_name(args)</code>. These blocks are extracted, parsed, and converted into structured events before reaching the TTS stage, ensuring spoken output remains natural while clients receive actionable tool instructions.
How Tool Calls Flow Through the LLM Integration
LLM Generation and Extraction
When the LLM generates a response containing tool invocations, the system uses extract_function_calls_from_text in src/speech_to_speech/LLM/language_model.py to locate code blocks. This function separates the conversational text from executable calls, returning a tuple of clean text and raw call strings.
The LLM is instructed via system prompts (constructed by build_tool_system_prompt in tool_prompt.py) to wrap function calls in <code> tags, enabling reliable extraction without complex regex on arbitrary text.
Parsing and Validation
Raw extracted calls are processed by parse_function_call in src/speech_to_speech/LLM/tool_call/function_call.py. This module implements:
_split_top_level_calls: Tokenizes the string safely, handling nested parentheses and string literals to separate multiple function callsFunctionToolCall: A Pydantic model representing the parsed call with name and arguments- Validation:
FunctionToolCall.to_realtime_function_tool_call()validates the call against availableFunctionTooldefinitions fromfunction_tool.py, producing anopenai.types.ResponseFunctionToolCallobject
This validation ensures only registered tools with correct schemas are forwarded to clients, preventing malformed or unauthorized function executions.
Pipeline Event Creation
The LMOutputProcessor in src/speech_to_speech/LLM/lm_output_processor.py bridges parsing and event emission. It receives LMOutput containing both the cleaned text and validated tool list, constructing an AssistantTextEvent (defined in src/speech_to_speech/pipeline/events.py).
This event carries:
type:"assistant_text"text: The cleaned conversational contenttools: A list ofResponseFunctionToolCallobjectsturn_id: The conversation turn identifier
The processor places this event on the internal text_output_queue for downstream consumption.
WebSocket Event Emission
The WebSocketRouter in src/speech_to_speech/connections/websocket_streamer.py consumes AssistantTextEvent objects from the queue. For each event, it:
- Sends the transcript text to the client for display
- Iterates through the
toolslist and emits separate tool-call events formatted as OpenAI Realtime protocol messages (type: "function_call") - Increments usage counters (
st.response_usage.tool_calls += len(event.tools)) insrc/speech_to_speech/api/openai_realtime/handlers/response.py
Downstream TTS components receive only the cleaned text from the event, ensuring the voice output is not polluted by tool-call syntax or JSON arguments.
Implementation Examples
Configuring Tool Prompts
Configure the LLM to emit tool calls by building a system prompt with available function definitions:
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool
search_tool = FunctionTool(
name="web_search",
description="Search the internet for current information",
parameters={
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"]
}
)
system_prompt = build_tool_system_prompt([search_tool])
# LLM will now generate calls like: <code>web_search(query="weather today")</code>
Extracting and Validating Calls
Process raw LLM output to extract and validate function calls:
from speech_to_speech.LLM.language_model import extract_function_calls_from_text
llm_response = 'Let me search that for you.\n<code>web_search(query="Python documentation")</code>'
clean_text, tool_calls = extract_function_calls_from_text(llm_response)
# Convert to OpenAI-compatible format
if tool_calls:
realtime_call = tool_calls[0].to_realtime_function_tool_call([search_tool])
# realtime_call contains name, arguments, and call_id
Building Pipeline Events
Create the event object that flows through the pipeline:
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from speech_to_speech.pipeline.events import AssistantTextEvent
processor = LMOutputProcessor(config={...})
event = AssistantTextEvent(
text=clean_text,
tools=[realtime_call], # List[ResponseFunctionToolCall]
turn_id="turn_42"
)
processor.text_output_queue.put(event)
Handling Events in the WebSocket Router
The router separates text from tool calls when transmitting to clients:
# Simplified logic from websocket_streamer.py
if isinstance(event, AssistantTextEvent):
await websocket.send_text(event.text)
for tool in event.tools:
await websocket.send_json({
"type": "function_call",
"name": tool.name,
"arguments": tool.arguments,
"call_id": tool.call_id
})
Summary
- Tool calls are wrapped in
<code>blocks by the LLM and extracted viaextract_function_calls_from_textinlanguage_model.py - Parsing occurs in
function_call.pyusing_split_top_level_callsand validated againstFunctionToolschemas to produceResponseFunctionToolCallobjects - Pipeline events are created by
LMOutputProcessorasAssistantTextEventobjects containing both text and validated tools - WebSocket emission in
websocket_streamer.pysends clean text to TTS while emitting separatefunction_callevents to clients over the OpenAI Realtime protocol - Usage tracking increments
tool_callscounters inresponse.pyfor monitoring and billing
Frequently Asked Questions
How are tool calls formatted in the LLM output?
The LLM wraps function invocations in XML-style <code> tags (e.g., <code>search(query="news")</code>). The extract_function_calls_from_text function in language_model.py specifically looks for these tags to separate executable code from conversational text, enabling reliable extraction without complex natural language parsing.
What prevents tool call syntax from being spoken by the TTS system?
The LMOutputProcessor creates AssistantTextEvent objects that separate the cleaned text field (sent to TTS) from the tools list (sent to the client). When the WebSocketRouter processes the event, it forwards only the text content to the TTS stage, ensuring the voice output contains natural language while the JSON tool arguments are transmitted silently to the client application.
How does the system handle multiple tool calls in a single response?
The _split_top_level_calls helper in function_call.py tokenizes the raw string to identify individual function calls separated by commas or newlines within the <code> block. Each call is parsed into a separate FunctionToolCall object, and the AssistantTextEvent carries a list of tools that the WebSocketRouter iterates through, emitting discrete function_call events for each tool invocation.
What protocol does the WebSocket router use for tool call events?
The WebSocketRouter in websocket_streamer.py emits tool calls using the OpenAI Realtime API protocol, sending JSON messages with "type": "function_call" containing the name, arguments, and call_id fields. This ensures compatibility with OpenAI's Realtime client SDKs while maintaining the flexibility to handle custom tool definitions validated against the FunctionTool schema.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →