# How to Implement Tool Calling and Function Calling with the Speech-to-Speech Voice Pipeline

> Learn to implement tool calling and function calling in the Hugging Face Speech-to-Speech pipeline. Discover how function calls are parsed and validated for seamless integration.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-07

---

**The Hugging Face Speech-to-Speech pipeline implements tool calling through OpenAI Realtime-compatible function calls wrapped in `<code>...</code>` delimiters, parsed during streaming and validated against JSON Schema definitions before emission as `ResponseFunctionToolCall` objects.**

The `huggingface/speech-to-speech` repository provides a complete voice-to-voice stack that mirrors the OpenAI Realtime protocol for **tool calling and function calling with the voice pipeline**. This guide walks through the implementation layers, from declaring tools to parsing LLM-generated calls during real-time streaming.

## Understanding the Tool-Calling Architecture

The implementation spans three architectural layers, each with dedicated source files:

| Layer | Purpose | Source File |
|-------|---------|-------------|
| **Tool definition** | `FunctionTool` subclasses `RealtimeFunctionTool` and renders Python-style signatures for system prompts | [`src/speech_to_speech/LLM/tool_call/function_tool.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/function_tool.py) |
| **Prompt construction** | `build_tool_system_prompt` injects tool definitions and `<code>`/`</code>` delimiters into the system prompt | [`src/speech_to_speech/LLM/tool_call/tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/tool_prompt.py) |
| **Runtime parsing** | `extract_function_calls_from_text` tokenizes output, parses calls, validates arguments, and produces `ResponseFunctionToolCall` objects | [`src/speech_to_speech/LLM/tool_call/function_call.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/function_call.py) and [`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py) |

## Declaring a Tool with FunctionTool

Tools inherit from `RealtimeFunctionTool` through the `FunctionTool` class. Define the tool's JSON Schema parameters and metadata:

```python
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool

light_tool = FunctionTool()
light_tool.name = "set_lights"
light_tool.description = "Set the smart-home lights to a colour or brightness."
light_tool.type = "function"
light_tool.parameters = {
    "type": "object",
    "properties": {
        "room": {"type": "string", "description": "Room identifier, e.g. 'kitchen'"},
        "color": {"type": "string", "enum": ["red", "green", "blue"], "description": "Desired colour"},
        "brightness": {"type": "integer", "minimum": 0, "maximum": 100}
    },
    "required": ["room"]
}

```

The `parameters` field follows standard JSON Schema. The `FunctionTool` class automatically renders this into a Python-style function signature for inclusion in the LLM's system prompt.

## Attaching Tools to a Realtime Session

Register tools in the session configuration before processing begins. The `BaseLanguageModelHandler._apply_instructions` method extracts these definitions and builds the tool-aware system prompt:

```python
from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig

runtime_cfg = RuntimeConfig()
runtime_cfg.session.tools = [light_tool]
runtime_cfg.session.tool_choice = "auto"  # or "none", "required", or a specific tool name

```

The `tool_choice` parameter controls invocation behavior:
- `"auto"` — Model decides whether to call a tool
- `"none"` — Disables tool calling
- `"required"` — Forces at least one tool call
- Specific tool name — Restricts to that tool only

## Building the Tool-Aware System Prompt

Generate the prompt section using `build_tool_system_prompt`:

```python
from speech_to_speech.LLM.tool_call.tool_prompt import (
    build_tool_system_prompt,
    ENTER_CODE,
    END_CODE,
)

tool_prompt = build_tool_system_prompt(
    tools=[light_tool],
    text_only=False,  # False = voice mode with "speak first" instruction

)

```

Setting `text_only=False` adds voice-specific instructions requiring the model to emit natural speech before any `<code>` block. The resulting prompt resembles:

```

Available tools:

def set_lights(room: str, color: str = None, brightness: int = None):
    """Set the smart-home lights to a colour or brightness.

    Args:
        room: Room identifier, e.g. 'kitchen'
        color: Desired colour
        brightness: 0-100
    """
...
To call a tool, put exactly one named-argument function call inside <code>...</code>:
<code>function_name(arg='value')</code>

```

## Parsing Tool Calls During Streaming

When the LLM emits output containing a tool call, the pipeline handles it in `BaseLanguageModelHandler._process_printable_text`. The method detects the `ENTER_CODE` delimiter (`<code>`), splits pre-code text into sent-tokenized speech chunks, and delegates parsing to `extract_function_calls_from_text`:

```python
from speech_to_speech.LLM.tool_call.function_call import extract_function_calls_from_text

# Inside the streaming handler

outside_text, tool_calls = extract_function_calls_from_text(
    text_with_code,           # Raw LLM output containing <code>...</code>

    block_regex=ctx.block_regex  # Compiled pattern matching delimiters

)

```

The `parse_function_call` function uses a **tokenizer-aware parser** to handle nested parentheses, quoted strings, tuples, and dictionaries—critical for robust extraction from streaming text.

## Validating and Converting Tool Calls

Each extracted `FunctionToolCall` undergoes validation via `to_realtime_function_tool_call`:

- **Positional arguments rejected** — Only named arguments `(key=value)` permitted
- **Schema validation** — Required parameters verified, undeclared parameters discarded
- **Tool existence check** — Called name must exist in `ctx.function_tools`
- **JSON serialization** — Arguments encoded and unique `call_id` assigned

Failed validation logs a warning and drops the call, allowing the LLM to continue speaking rather than crashing the pipeline.

The validated call becomes a `ResponseFunctionToolCall` stored in `ctx.tools` and yielded in an `LLMResponseChunk`:

```python

# From messages.py - LLMResponseChunk carries tool calls downstream

@dataclass
class LLMResponseChunk:
    text: Optional[str] = None
    tools: Optional[List[ResponseFunctionToolCall]] = None
    usage: Optional[TokenUsage] = None
    finish_reason: Optional[str] = None

```

## Emitting Tool Calls to the Client

The Realtime server serializes tool calls into `RealtimeConversationItemFunctionCall` events:

```python
original_chat.add_item(
    RealtimeConversationItemFunctionCall(
        type="function_call",
        id=tool.id,
        call_id=tool.call_id,      # Unique identifier for this invocation

        name=tool.name,            # Tool name

        arguments=tool.arguments,  # JSON-encoded parameters

        status=tool.status,        # "in_progress", "completed", etc.

    )
)

```

The client receives this event, executes the requested action, and may return a tool result message to continue the conversation.

## Complete Implementation Example

```python
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.language_model import LanguageModelHandler
from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig
from speech_to_speech.pipeline.handler_types import GenerateResponseRequest

# 1. Define the tool

weather_tool = FunctionTool()
weather_tool.name = "get_weather"
weather_tool.description = "Retrieve current weather for a location."
weather_tool.parameters = {
    "type": "object",
    "properties": {
        "location": {"type": "string", "description": "City name"},
        "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
    },
    "required": ["location"]
}

# 2. Configure session with tool

runtime_cfg = RuntimeConfig()
runtime_cfg.session.tools = [weather_tool]
runtime_cfg.session.tool_choice = "auto"

# 3. Build system prompt for voice mode

system_prompt = build_tool_system_prompt([weather_tool], text_only=False)

# 4. Initialize handler

handler = LanguageModelHandler()
handler.setup(
    model_name="Qwen/Qwen2.5-7B-Instruct",
    device="cuda",
    torch_dtype="bfloat16",
)

# 5. Process request and handle tool calls

request = GenerateResponseRequest(
    turn_id="turn-001",
    runtime_config=runtime_cfg,
)

for chunk in handler.process(request):
    if chunk.tools:
        for tool in chunk.tools:
            print(f"CALL {tool.call_id}: {tool.name}({tool.arguments})")
            # Execute tool, then send result back to conversation

    elif chunk.text:
        print(f"ASSISTANT: {chunk.text}")

```

## Enforcing Turn-Taking and Safety Rules

The pipeline implements three critical constraints:

- **Single call per response** — `BaseLanguageModelHandler._process_printable_text` filters extra calls with a warning
- **Voice preamble required** — The voice prompt template mandates natural speech before `<code>` blocks
- **Speculative cancellation** — New user turns can abort in-progress generation, maintaining responsiveness

These rules align with OpenAI Realtime protocol behavior and ensure **tool calling and function calling with the voice pipeline** remains robust in real-time scenarios.

## Key Source Files Reference

| File | Function |
|------|----------|
| [`src/speech_to_speech/LLM/tool_call/function_tool.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/function_tool.py) | `FunctionTool` class definition and signature rendering |
| [`src/speech_to_speech/LLM/tool_call/tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/tool_prompt.py) | `build_tool_system_prompt()` and delimiter constants (`ENTER_CODE`, `END_CODE`) |
| [`src/speech_to_speech/LLM/tool_call/function_call.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/function_call.py) | `extract_function_calls_from_text()`, `parse_function_call()`, `FunctionToolCall` validation |
| [`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py) | `BaseLanguageModelHandler` with `_process_printable_text()` and `_apply_instructions()` |
| [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py) | `LLMResponseChunk`, `ResponseFunctionToolCall` data models |

## Summary

- **Tool calling and function calling with the voice pipeline** uses `<code>...</code>` delimiters compatible with OpenAI Realtime
- `FunctionTool` declares schemas, `build_tool_system_prompt` injects them, and `extract_function_calls_from_text` parses streaming output
- The tokenizer-aware parser handles complex nested structures robustly
- Validation enforces named-only arguments, schema compliance, and declared tool existence
- Tool calls propagate through `LLMResponseChunk` to `RealtimeConversationItemFunctionCall` events

## Frequently Asked Questions

### What delimiter format does the Speech-to-Speech pipeline use for tool calls?

The pipeline uses `<code>...</code>` as the default delimiter pair, defined by `ENTER_CODE` and `END_CODE` constants in [`tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/tool_prompt.py). The LLM must wrap function calls inside these tags: `<code>tool_name(arg='value')</code>`.

### Can I force the model to use a specific tool?

Yes. Set `runtime_cfg.session.tool_choice = "your_tool_name"` instead of `"auto"`. This restricts the model to that exact tool, matching OpenAI's `tool_choice` behavior.

### How does the pipeline handle invalid tool calls?

Invalid calls are logged and discarded. The `to_realtime_function_tool_call` method validates against the tool schema; failures return `None` and emit a warning, allowing the conversation to continue rather than erroring.

### Is text-only mode available without voice instructions?

Yes. Pass `text_only=True` to `build_tool_system_prompt`. This removes the "speak first" requirement, suitable for chatbot interfaces without TTS or when integrating with non-voice clients.