# How Tool Calls Work Through the LLM Integration in Hugging Face Speech-to-Speech

> Understand LLM tool calls in Hugging Face Speech-to-Speech. Learn how tool calls are processed, validated, and emit OpenAI Realtime compatible function_call events via WebSocket.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-30

---

**Tool calls flow from LLM text generation through parsing and validation, ultimately emitting OpenAI Realtime-compatible `function_call` events via WebSocket while ensuring TTS audio receives only clean text.**

The huggingface/speech-to-speech repository implements a real-time **four-stage pipeline** (VAD → STT → LLM → TTS) where the LLM can invoke external functions. Understanding how tool calls work through the LLM integration requires tracing the path from raw text generation to validated pipeline events.

## The Four-Stage Pipeline Architecture

Speech-to-Speech processes audio through sequential stages: **Voice Activity Detection (VAD)**, **Speech-to-Text (STT)**, **Large Language Model (LLM)**, and **Text-to-Speech (TTS)**. Tool calls originate in the LLM stage when the model generates output containing code blocks like `<code>tool_name(args)</code>`. These blocks are extracted, parsed, and converted into structured events before reaching the TTS stage, ensuring spoken output remains natural while clients receive actionable tool instructions.

## How Tool Calls Flow Through the LLM Integration

### LLM Generation and Extraction

When the LLM generates a response containing tool invocations, the system uses **`extract_function_calls_from_text`** in [`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py) to locate code blocks. This function separates the conversational text from executable calls, returning a tuple of clean text and raw call strings.

The LLM is instructed via system prompts (constructed by `build_tool_system_prompt` in [`tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/tool_prompt.py)) to wrap function calls in `<code>` tags, enabling reliable extraction without complex regex on arbitrary text.

### Parsing and Validation

Raw extracted calls are processed by **`parse_function_call`** in [`src/speech_to_speech/LLM/tool_call/function_call.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/function_call.py). This module implements:

- **`_split_top_level_calls`**: Tokenizes the string safely, handling nested parentheses and string literals to separate multiple function calls
- **`FunctionToolCall`**: A Pydantic model representing the parsed call with name and arguments
- **Validation**: `FunctionToolCall.to_realtime_function_tool_call()` validates the call against available `FunctionTool` definitions from [`function_tool.py`](https://github.com/huggingface/speech-to-speech/blob/main/function_tool.py), producing an `openai.types.ResponseFunctionToolCall` object

This validation ensures only registered tools with correct schemas are forwarded to clients, preventing malformed or unauthorized function executions.

### Pipeline Event Creation

The **`LMOutputProcessor`** in [`src/speech_to_speech/LLM/lm_output_processor.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/lm_output_processor.py) bridges parsing and event emission. It receives `LMOutput` containing both the cleaned text and validated tool list, constructing an **`AssistantTextEvent`** (defined in [`src/speech_to_speech/pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/events.py)).

This event carries:
- `type`: `"assistant_text"`
- `text`: The cleaned conversational content
- `tools`: A list of `ResponseFunctionToolCall` objects
- `turn_id`: The conversation turn identifier

The processor places this event on the internal `text_output_queue` for downstream consumption.

### WebSocket Event Emission

The **WebSocketRouter** in [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) consumes `AssistantTextEvent` objects from the queue. For each event, it:

1. Sends the transcript text to the client for display
2. Iterates through the `tools` list and emits separate **tool-call events** formatted as OpenAI Realtime protocol messages (`type: "function_call"`)
3. Increments usage counters (`st.response_usage.tool_calls += len(event.tools)`) in [`src/speech_to_speech/api/openai_realtime/handlers/response.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/handlers/response.py)

Downstream TTS components receive only the cleaned text from the event, ensuring the voice output is not polluted by tool-call syntax or JSON arguments.

## Implementation Examples

### Configuring Tool Prompts

Configure the LLM to emit tool calls by building a system prompt with available function definitions:

```python
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool

search_tool = FunctionTool(
    name="web_search",
    description="Search the internet for current information",
    parameters={
        "type": "object",
        "properties": {"query": {"type": "string"}},
        "required": ["query"]
    }
)

system_prompt = build_tool_system_prompt([search_tool])

# LLM will now generate calls like: <code>web_search(query="weather today")</code>

```

### Extracting and Validating Calls

Process raw LLM output to extract and validate function calls:

```python
from speech_to_speech.LLM.language_model import extract_function_calls_from_text

llm_response = 'Let me search that for you.\n<code>web_search(query="Python documentation")</code>'
clean_text, tool_calls = extract_function_calls_from_text(llm_response)

# Convert to OpenAI-compatible format

if tool_calls:
    realtime_call = tool_calls[0].to_realtime_function_tool_call([search_tool])
    # realtime_call contains name, arguments, and call_id

```

### Building Pipeline Events

Create the event object that flows through the pipeline:

```python
from speech_to_speech.LLM.lm_output_processor import LMOutputProcessor
from speech_to_speech.pipeline.events import AssistantTextEvent

processor = LMOutputProcessor(config={...})
event = AssistantTextEvent(
    text=clean_text,
    tools=[realtime_call],  # List[ResponseFunctionToolCall]

    turn_id="turn_42"
)

processor.text_output_queue.put(event)

```

### Handling Events in the WebSocket Router

The router separates text from tool calls when transmitting to clients:

```python

# Simplified logic from websocket_streamer.py

if isinstance(event, AssistantTextEvent):
    await websocket.send_text(event.text)
    
    for tool in event.tools:
        await websocket.send_json({
            "type": "function_call",
            "name": tool.name,
            "arguments": tool.arguments,
            "call_id": tool.call_id
        })

```

## Summary

- **Tool calls** are wrapped in `<code>` blocks by the LLM and extracted via `extract_function_calls_from_text` in [`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py)
- **Parsing** occurs in [`function_call.py`](https://github.com/huggingface/speech-to-speech/blob/main/function_call.py) using `_split_top_level_calls` and validated against `FunctionTool` schemas to produce `ResponseFunctionToolCall` objects
- **Pipeline events** are created by `LMOutputProcessor` as `AssistantTextEvent` objects containing both text and validated tools
- **WebSocket emission** in [`websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_streamer.py) sends clean text to TTS while emitting separate `function_call` events to clients over the OpenAI Realtime protocol
- **Usage tracking** increments `tool_calls` counters in [`response.py`](https://github.com/huggingface/speech-to-speech/blob/main/response.py) for monitoring and billing

## Frequently Asked Questions

### How are tool calls formatted in the LLM output?

The LLM wraps function invocations in XML-style `<code>` tags (e.g., `<code>search(query="news")</code>`). The `extract_function_calls_from_text` function in [`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py) specifically looks for these tags to separate executable code from conversational text, enabling reliable extraction without complex natural language parsing.

### What prevents tool call syntax from being spoken by the TTS system?

The `LMOutputProcessor` creates `AssistantTextEvent` objects that separate the cleaned `text` field (sent to TTS) from the `tools` list (sent to the client). When the WebSocketRouter processes the event, it forwards only the text content to the TTS stage, ensuring the voice output contains natural language while the JSON tool arguments are transmitted silently to the client application.

### How does the system handle multiple tool calls in a single response?

The `_split_top_level_calls` helper in [`function_call.py`](https://github.com/huggingface/speech-to-speech/blob/main/function_call.py) tokenizes the raw string to identify individual function calls separated by commas or newlines within the `<code>` block. Each call is parsed into a separate `FunctionToolCall` object, and the `AssistantTextEvent` carries a list of `tools` that the WebSocketRouter iterates through, emitting discrete `function_call` events for each tool invocation.

### What protocol does the WebSocket router use for tool call events?

The `WebSocketRouter` in [`websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_streamer.py) emits tool calls using the **OpenAI Realtime API protocol**, sending JSON messages with `"type": "function_call"` containing the `name`, `arguments`, and `call_id` fields. This ensures compatibility with OpenAI's Realtime client SDKs while maintaining the flexibility to handle custom tool definitions validated against the `FunctionTool` schema.