OpenAI Realtime Event System: Tool Calling Differences Between Local LLM and API Backends

The OpenAI Realtime event system handles tool calling through discrete protocol events in the OpenAI API backend, while local LLM backends process tool calls as streaming deltas within the Chat Completions flow, using an internal pending-call map to track state.

The huggingface/speech-to-speech library implements dual execution paths for tool calling, diverging significantly depending on whether you use a local LLM backend or the OpenAI Realtime API. Understanding these architectural differences is crucial for debugging function calling behavior and optimizing latency in speech-to-speech pipelines.

Local LLM Backend: Streaming Delta Processing

Local LLM integrations rely on the Chat class to drive generation and manage tool state through continuous streaming.

Trigger and Generation Flow

The generation is driven by the library’s own Chat class in src/speech_to_speech/LLM/chat.py, which interfaces with models through the Chat Completions protocol. Unlike the discrete event system of the Realtime API, local LLMs emit tool calls as part of the standard streaming response.

Tool-Call Format and State Tracking

The LLM emits delta objects inside the streamed choices[].delta.tool_calls field, exactly as defined by the Chat Completions spec. The Chat class stores each pending call in Chat._pending_tool_calls:


# Inside Chat (src/speech_to_speech/LLM/chat.py)

# 1️⃣ Generation produces a delta with a tool call

if delta.tool_calls:
    for tc in delta.tool_calls:
        # Store pending call

        self._pending_tool_calls[tc.id] = tc

# 2️⃣ Later the model returns a function‑call output delta

if delta.function_call_output:
    fc = self._pending_tool_calls.pop(delta.call_id)
    # Convert to internal ToolCall object for the pipeline

    tool_call = ToolCall(item=fc)
    self.buffer.append(tool_call)

When the model later produces a function_call_output delta, the pending entry is removed and the result is turned into a ToolCall object that the pipeline consumes.

Tools are announced to the local LLM via the system prompt generated by tool_prompt.build_tool_system_prompt in src/speech_to_speech/LLM/tool_call/tool_prompt.py. The FunctionTool Pydantic model in src/speech_to_speech/LLM/tool_call/function_tool.py defines the tool schema used for both backends.

OpenAI Realtime API Backend: Discrete Event Processing

The OpenAI Realtime API backend isolates tool calls into distinct protocol events that the service layer explicitly creates and forwards.

Event-Driven Architecture

The generation is driven by the OpenAI Realtime service (src/speech_to_speech/api/openai_realtime/service.py). The Realtime protocol emits discrete events rather than streaming deltas:

  • response.function_call_arguments.done (arguments)
  • conversation.item.create (function‑call output)

Event Conversion and Handling

ResponseHandler receives an AssistantTextEvent that already contains a list of ResponseFunctionToolCall objects (populated by the LLM). It converts each one into a ResponseFunctionCallArgumentsDoneEvent in the loop starting at line 24 of src/speech_to_speech/api/openai_realtime/handlers/response.py:


# In ResponseHandler.handle_assistant_text (src/speech_to_speech/api/openai_realtime/handlers/response.py)

if event.tools:
    st.response_usage.tool_calls += len(event.tools)
    for tool in event.tools:
        events.append(
            ResponseFunctionCallArgumentsDoneEvent(
                type="response.function_call_arguments.done",
                event_id=self._next_event_id(),
                call_id=tool.call_id,
                name=tool.name,
                arguments=tool.arguments,
                item_id=item_id,
                output_index=output_idx,
                response_id=resp_id,
            )
        )
        output_idx += 1

The ResponseHandler increments the usage counter (st.response_usage.tool_calls) for each tool call and relies on the Realtime protocol to deliver the output as a separate conversation item.

Tools are sent to the Realtime API as part of the response.create payload (RealtimeResponseCreateParams.tools), using the same FunctionTool model for serialization.

Key Architectural Differences

Aspect Local LLM Backend OpenAI Realtime API Backend
Driver Chat class (src/speech_to_speech/LLM/chat.py) RealtimeService (src/speech_to_speech/api/openai_realtime/service.py)
Tool Format Streaming deltas (choices[].delta.tool_calls) Discrete events (response.function_call_arguments.done)
State Management Internal _pending_tool_calls map Protocol-level event stream
Conversion Layer ChatCompletionsLanguageModel translates to ResponseFunctionToolCall ResponseHandler creates ResponseFunctionCallArgumentsDoneEvent
Tool Definition System prompt via tool_prompt.build_tool_system_prompt RealtimeResponseCreateParams.tools payload

Summary

  • Local LLM backends process tool calls as part of the same streaming generation, using an internal Chat._pending_tool_calls map to track state until function outputs arrive.
  • OpenAI Realtime API backends isolate tool calls into distinct protocol events, with ResponseHandler explicitly creating ResponseFunctionCallArgumentsDoneEvent objects for the event stream.
  • The Chat class in chat.py manages local tool state, while the RealtimeService in service.py orchestrates API event translation.
  • Tool definitions flow through tool_prompt.py for local models and via RealtimeResponseCreateParams for the API backend, both utilizing the shared FunctionTool model.

Frequently Asked Questions

How does the OpenAI Realtime event system represent tool calls differently than local LLM streaming?

The OpenAI Realtime event system represents tool calls as discrete events (response.function_call_arguments.done), whereas local LLM backends emit them as streaming deltas within the Chat Completions choices[].delta.tool_calls field. This distinction requires different handling logic in the ResponseHandler versus the Chat class.

Where does the speech-to-speech library track pending tool calls for local LLMs?

The library tracks pending tool calls in Chat._pending_tool_calls, a dictionary defined in src/speech_to_speech/LLM/chat.py. This map stores tool call objects by ID until the model returns a corresponding function_call_output delta, at which point the entry is popped and converted to a ToolCall object.

What converts internal tool representations to OpenAI Realtime wire format?

The ResponseHandler in src/speech_to_speech/api/openai_realtime/handlers/response.py converts AssistantTextEvent.tools (containing ResponseFunctionToolCall objects) into ResponseFunctionCallArgumentsDoneEvent instances. These events map directly to the OpenAI Realtime protocol specification.

How are tool definitions transmitted to each backend?

For local LLMs, tools are announced via a system prompt generated by tool_prompt.build_tool_system_prompt in src/speech_to_speech/LLM/tool_call/tool_prompt.py. For the OpenAI Realtime API, tools are sent as part of the response.create payload using RealtimeResponseCreateParams.tools.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →