# OpenAI Realtime Event System: Tool Calling Differences Between Local LLM and API Backends

> Explore OpenAI tool calling differences: API backends use discrete events, local LLMs stream deltas. Learn how the realtime event system manages state.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-10

---

**The OpenAI Realtime event system handles tool calling through discrete protocol events in the OpenAI API backend, while local LLM backends process tool calls as streaming deltas within the Chat Completions flow, using an internal pending-call map to track state.**

The huggingface/speech-to-speech library implements dual execution paths for tool calling, diverging significantly depending on whether you use a local LLM backend or the OpenAI Realtime API. Understanding these architectural differences is crucial for debugging function calling behavior and optimizing latency in speech-to-speech pipelines.

## Local LLM Backend: Streaming Delta Processing

Local LLM integrations rely on the **Chat** class to drive generation and manage tool state through continuous streaming.

### Trigger and Generation Flow

The generation is driven by the library’s own `Chat` class in [`src/speech_to_speech/LLM/chat.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat.py), which interfaces with models through the Chat Completions protocol. Unlike the discrete event system of the Realtime API, local LLMs emit tool calls as part of the standard streaming response.

### Tool-Call Format and State Tracking

The LLM emits *delta* objects inside the streamed `choices[].delta.tool_calls` field, exactly as defined by the Chat Completions spec. The `Chat` class stores each pending call in `Chat._pending_tool_calls`:

```python

# Inside Chat (src/speech_to_speech/LLM/chat.py)

# 1️⃣ Generation produces a delta with a tool call

if delta.tool_calls:
    for tc in delta.tool_calls:
        # Store pending call

        self._pending_tool_calls[tc.id] = tc

# 2️⃣ Later the model returns a function‑call output delta

if delta.function_call_output:
    fc = self._pending_tool_calls.pop(delta.call_id)
    # Convert to internal ToolCall object for the pipeline

    tool_call = ToolCall(item=fc)
    self.buffer.append(tool_call)

```

When the model later produces a `function_call_output` delta, the pending entry is removed and the result is turned into a `ToolCall` object that the pipeline consumes.

Tools are announced to the local LLM via the system prompt generated by `tool_prompt.build_tool_system_prompt` in [`src/speech_to_speech/LLM/tool_call/tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/tool_prompt.py). The `FunctionTool` Pydantic model in [`src/speech_to_speech/LLM/tool_call/function_tool.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/function_tool.py) defines the tool schema used for both backends.

## OpenAI Realtime API Backend: Discrete Event Processing

The OpenAI Realtime API backend isolates tool calls into distinct protocol events that the service layer explicitly creates and forwards.

### Event-Driven Architecture

The generation is driven by the OpenAI Realtime service ([`src/speech_to_speech/api/openai_realtime/service.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/service.py)). The Realtime protocol emits discrete events rather than streaming deltas:

- `response.function_call_arguments.done` (arguments)
- `conversation.item.create` (function‑call output)

### Event Conversion and Handling

`ResponseHandler` receives an `AssistantTextEvent` that already contains a list of `ResponseFunctionToolCall` objects (populated by the LLM). It converts each one into a `ResponseFunctionCallArgumentsDoneEvent` in the loop starting at line 24 of [`src/speech_to_speech/api/openai_realtime/handlers/response.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/handlers/response.py):

```python

# In ResponseHandler.handle_assistant_text (src/speech_to_speech/api/openai_realtime/handlers/response.py)

if event.tools:
    st.response_usage.tool_calls += len(event.tools)
    for tool in event.tools:
        events.append(
            ResponseFunctionCallArgumentsDoneEvent(
                type="response.function_call_arguments.done",
                event_id=self._next_event_id(),
                call_id=tool.call_id,
                name=tool.name,
                arguments=tool.arguments,
                item_id=item_id,
                output_index=output_idx,
                response_id=resp_id,
            )
        )
        output_idx += 1

```

The `ResponseHandler` increments the usage counter (`st.response_usage.tool_calls`) for each tool call and relies on the Realtime protocol to deliver the output as a separate conversation item.

Tools are sent to the Realtime API as part of the `response.create` payload (`RealtimeResponseCreateParams.tools`), using the same `FunctionTool` model for serialization.

## Key Architectural Differences

| Aspect | Local LLM Backend | OpenAI Realtime API Backend |
|--------|------------------|------------------------------|
| **Driver** | `Chat` class ([`src/speech_to_speech/LLM/chat.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat.py)) | `RealtimeService` ([`src/speech_to_speech/api/openai_realtime/service.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/service.py)) |
| **Tool Format** | Streaming deltas (`choices[].delta.tool_calls`) | Discrete events (`response.function_call_arguments.done`) |
| **State Management** | Internal `_pending_tool_calls` map | Protocol-level event stream |
| **Conversion Layer** | `ChatCompletionsLanguageModel` translates to `ResponseFunctionToolCall` | `ResponseHandler` creates `ResponseFunctionCallArgumentsDoneEvent` |
| **Tool Definition** | System prompt via `tool_prompt.build_tool_system_prompt` | `RealtimeResponseCreateParams.tools` payload |

## Summary

- **Local LLM backends** process tool calls as part of the same streaming generation, using an internal `Chat._pending_tool_calls` map to track state until function outputs arrive.
- **OpenAI Realtime API backends** isolate tool calls into distinct protocol events, with `ResponseHandler` explicitly creating `ResponseFunctionCallArgumentsDoneEvent` objects for the event stream.
- The **Chat** class in [`chat.py`](https://github.com/huggingface/speech-to-speech/blob/main/chat.py) manages local tool state, while the **RealtimeService** in [`service.py`](https://github.com/huggingface/speech-to-speech/blob/main/service.py) orchestrates API event translation.
- Tool definitions flow through [`tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/tool_prompt.py) for local models and via `RealtimeResponseCreateParams` for the API backend, both utilizing the shared `FunctionTool` model.

## Frequently Asked Questions

### How does the OpenAI Realtime event system represent tool calls differently than local LLM streaming?

The OpenAI Realtime event system represents tool calls as discrete events (`response.function_call_arguments.done`), whereas local LLM backends emit them as streaming deltas within the Chat Completions `choices[].delta.tool_calls` field. This distinction requires different handling logic in the `ResponseHandler` versus the `Chat` class.

### Where does the speech-to-speech library track pending tool calls for local LLMs?

The library tracks pending tool calls in `Chat._pending_tool_calls`, a dictionary defined in [`src/speech_to_speech/LLM/chat.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat.py). This map stores tool call objects by ID until the model returns a corresponding `function_call_output` delta, at which point the entry is popped and converted to a `ToolCall` object.

### What converts internal tool representations to OpenAI Realtime wire format?

The `ResponseHandler` in [`src/speech_to_speech/api/openai_realtime/handlers/response.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/handlers/response.py) converts `AssistantTextEvent.tools` (containing `ResponseFunctionToolCall` objects) into `ResponseFunctionCallArgumentsDoneEvent` instances. These events map directly to the OpenAI Realtime protocol specification.

### How are tool definitions transmitted to each backend?

For local LLMs, tools are announced via a system prompt generated by `tool_prompt.build_tool_system_prompt` in [`src/speech_to_speech/LLM/tool_call/tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/tool_prompt.py). For the OpenAI Realtime API, tools are sent as part of the `response.create` payload using `RealtimeResponseCreateParams.tools`.