How to Handle and Stream Tool Calls from LLM in a Speech-to-Speech Pipeline
The Speech-to-Speech pipeline handles real-time tool calls by injecting tool definitions into the LLM system prompt, detecting code block delimiters in streamed tokens using pre-compiled regexes, and safely parsing arguments via AST evaluation before dispatching to concrete function implementations.
Handling and streaming tool calls from an LLM in a speech-to-speech pipeline requires tight coordination between prompt engineering, streaming text parsing, and secure code execution. The huggingface/speech-to-speech repository implements this through a dedicated tool-calling layer that intercepts LLM outputs in real-time, identifies function invocations wrapped in markdown-style code blocks, and converts them into executable commands while maintaining low latency for voice interactions.
Tool Call Architecture Overview
The implementation spans four core components that work together during streaming inference:
FunctionTool— Defines tool schemas using the OpenAI function-calling format and serializes them into code blocks that the LLM can emit.tool_prompt— Constructs system prompts that instruct the LLM when and how to invoke tools, and generates regex patterns for detecting tool calls in the stream.BaseOpenAICompatibleLanguageModel— Manages the streaming chat loop, applying pre-compiled regexes to each chunk to detect tool call delimiters.function_call— Safely parses raw tool arguments usingast.parseand converts tool definitions to the OpenAI Realtime API format.
Defining Tools with FunctionTool
Tools are represented by the FunctionTool dataclass in src/speech_to_speech/LLM/tool_call/function_tool.py. It mirrors the OpenAI function schema and provides serialization methods for both the system prompt and LLM output.
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool
camera_tool = FunctionTool(
name="camera",
description="Capture a picture from the webcam with a specified resolution.",
parameters={
"type": "object",
"properties": {
"resolution": {
"type": "string",
"enum": ["720p", "1080p"]
}
},
"required": ["resolution"]
},
type="function"
)
The to_code_prompt method renders the tool as a markdown code block using the ENTER_CODE and END_CODE delimiters (both defined as "```"). This creates a predictable pattern that the streaming parser can detect.
Building the System Prompt
The build_tool_system_prompt function in src/speech_to_speech/LLM/tool_call/tool_prompt.py assembles the instructions that tell the LLM how to use available tools.
Key capabilities:
- Voice mode: Instructs the LLM to prepend a brief natural sentence before calling a tool (important for user experience in voice assistants).
- Text-only mode: Sets
text_only=Trueto remove spoken lead-ins for pure text interfaces. - Constraint enforcement: Reminds the model that only one tool call may appear per response.
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
tools = [camera_tool]
system_prompt = build_tool_system_prompt(tools, text_only=False)
# System prompt includes available tools and instructions on delimiters
The module also exports build_block_regex(name), which compiles a regex pattern that captures everything between ```<tool_name>( and )``` , tolerant of whitespace and newlines.
Detecting Tool Calls in Streaming Responses
Streaming detection happens in src/speech_to_speech/LLM/base_openai_compatible_language_model.py. The BaseOpenAICompatibleLanguageModel initializes pre-compiled regexes for each tool during construction, then inspects every chunk in the streaming loop.
# Simplified extraction from BaseOpenAICompatibleLanguageModel.__init__ and chat()
self._tool_regexes = {
tool.name: build_block_regex(tool.name)
for tool in self.tools
}
async def chat(self, messages, *, stream=False):
async for chunk in self._stream_chat(messages):
text = chunk.get("content", "")
for name, regex in self._tool_regexes.items():
match = regex.search(text)
if match:
args_text = match.group("args")
# Parse and dispatch...
chunk["detected_tool"] = {"name": name, "args_raw": args_text}
yield chunk
The regex uses a non-greedy capture (.*?) to stop at the first closing delimiter, ensuring that nested structures or subsequent text do not interfere with detection.
Parsing and Converting Tool Arguments
Once a tool block is detected, parse_function_call in src/speech_to_speech/LLM/tool_call/function_call.py safely evaluates the argument string without executing arbitrary code.
Parsing strategy:
- Primary: Uses
ast.parsewithmode="eval"followed byast.literal_evalto handle Python literals (dicts, lists, strings, numbers). - Fallback: Attempts JSON parsing for LLMs that emit double-quoted JSON instead of Python syntax.
from speech_to_speech.LLM.tool_call.function_call import parse_function_call
raw_call = "camera(resolution='1080p')"
parsed = parse_function_call(raw_call)
# Returns [ParsedFunctionCall(function_name='camera', args={'resolution': '1080p'})]
For integration with the OpenAI Realtime API, to_realtime_function_tool_call converts FunctionTool objects into the required payload format:
from speech_to_speech.LLM.tool_call.function_call import to_realtime_function_tool_call
realtime_payload = to_realtime_function_tool_call([camera_tool])
# Output: [{"type": "function", "function": {"name": "camera", ...}}]
End-to-End Implementation Example
This example demonstrates initializing the pipeline with tools, streaming a conversation, and handling detected tool calls:
import asyncio
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool
from speech_to_speech.LLM.base_openai_compatible_language_model import BaseOpenAICompatibleLanguageModel
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.tool_call.function_call import parse_function_call
# 1. Define tools
camera = FunctionTool(
name="camera",
description="Take a photo",
parameters={"type": "object", "properties": {"res": {"type": "string"}}}
)
# 2. Build system prompt
system_prompt = build_tool_system_prompt([camera], text_only=False)
# 3. Initialize LLM with tools (implementation specific to your provider)
llm = BaseOpenAICompatibleLanguageModel(tools=[camera])
async def handle_conversation():
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": "Take a picture in 1080p"}
]
async for chunk in llm.chat(messages, stream=True):
content = chunk.get("content", "")
print(content, end="", flush=True)
# 4. Handle detected tool
if "detected_tool" in chunk:
tool_info = chunk["detected_tool"]
parsed = parse_function_call(
f"{tool_info['name']}({tool_info['args_raw']})"
)[0]
print(f"\n[Executing tool: {parsed.function_name}]")
print(f"Arguments: {parsed.args}")
# Dispatch to actual implementation here...
# Run
asyncio.run(handle_conversation())
Summary
- Tool definitions use the
FunctionTooldataclass to maintain OpenAI-compatible schemas and serialize to markdown code blocks. - System prompts are constructed via
build_tool_system_prompt, which instructs the LLM on tool usage and delimiter format while supporting voice-specific lead-in requirements. - Streaming detection relies on pre-compiled regexes from
build_block_regexthat scan each chunk inBaseOpenAICompatibleLanguageModel.chatfor```tool_name(...)```patterns. - Safe parsing leverages
ast.parseandast.literal_evalinparse_function_callto evaluate arguments without code execution vulnerabilities. - Realtime API compatibility is achieved through
to_realtime_function_tool_call, which transforms tools into the JSON format expected by OpenAI's Realtime endpoints.
Frequently Asked Questions
How does the pipeline detect tool calls while streaming text is still being generated?
The BaseOpenAICompatibleLanguageModel maintains a dictionary of pre-compiled regex patterns (one per tool) generated by build_block_regex. As each chunk arrives from the provider, it searches the accumulated text for the pattern ```tool_name(...)```. Because the regex uses non-greedy matching, it identifies complete tool blocks as soon as the closing delimiter appears, even if the LLM continues generating additional conversational text afterward.
What specific format must the LLM use to emit tool calls?
The LLM must emit tool calls inside markdown code fences using triple backticks. The expected format is ```tool_name(argument="value")```, where the tool name matches the name attribute of a registered FunctionTool. The constants ENTER_CODE and END_CODE in tool_prompt.py both define the delimiter as "```", creating a robust boundary that the streaming parser can detect without ambiguity.
How are tool arguments safely parsed from raw LLM output?
The parse_function_call function in function_call.py uses Python's ast module to safely evaluate the argument string. It first attempts ast.parse with mode="eval", then extracts the value node for literal evaluation. This approach supports complex nested structures (dictionaries, lists) while preventing code injection. If AST parsing fails, the function falls back to JSON parsing to accommodate LLMs that emit JSON-formatted arguments.
Can this architecture integrate with the OpenAI Realtime API?
Yes. The to_realtime_function_tool_call utility converts FunctionTool instances into the exact JSON structure required by the OpenAI Realtime API's tools field. Each tool is transformed into a dictionary containing type, function.name, function.description, and function.parameters, allowing seamless substitution between local tool definitions and external Realtime API configurations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →