How to Implement Tool Calling and Function Calling with the Language Model
The huggingface/speech-to-speech library enables language models to execute Python functions by injecting JSON-schema tool definitions into system prompts, parsing delimited code blocks with AST tokenization, and validating arguments against declared parameters before returning OpenAI-compatible ResponseFunctionToolCall objects.
Tool calling extends a language model's capabilities by allowing it to request execution of external functions during inference. The speech-to-speech repository provides a production-ready implementation that bridges schema-based tool definitions with real-time LLM interactions. This guide examines the complete pipeline from prompt construction to validated function execution.
Building the Tool System Prompt
The first stage converts FunctionTool definitions into natural language instructions that guide the model to emit properly formatted function calls.
FunctionTool and JSON Schema Definition
In src/speech_to_speech/LLM/tool_call/function_tool.py, the FunctionTool class extends openai.types.realtime.RealtimeFunctionTool to store a tool's name, description, and JSON schema (parameters). Each tool wraps a Python function's interface using standard JSON Schema syntax:
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool
schema = {
"type": "object",
"properties": {
"app_name": {"type": "string", "description": "Name of the app to open"},
"timeout": {"type": "number", "description": "Seconds before giving up"},
},
"required": ["app_name"],
}
mobile_open = FunctionTool(
type="function",
name="mobile.open_app",
description="Open an app on the phone",
parameters=schema,
)
Template Rendering with Jinja2
The build_tool_system_prompt function in src/speech_to_speech/LLM/tool_call/tool_prompt.py selects a Jinja2 template (voice-aware or text-only) and injects the tool list along with delimiter instructions. The FunctionTool.to_code_prompt() method renders a Python-style signature using signature_from_schema.py to convert JSON types into Python type hints:
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
prompt = build_tool_system_prompt([mobile_open])
The resulting system prompt instructs the model to wrap all function calls inside <code> and </code> tags, producing outputs like:
<code>mobile.open_app(app_name='drupe', timeout=5)</code>
Parsing LLM Output
When the model returns text containing function calls, the system must extract and parse the embedded code blocks without executing arbitrary code.
Block Detection and Extraction
The extract_function_calls_from_text function in src/speech_to_speech/LLM/tool_call/function_call.py uses a configurable regex (default r"<code>.*?</code>") to locate delimited blocks. It separates the surrounding conversational text from the code content, returning both the outside text and a list of extracted call strings.
Tokenization and AST Parsing
To safely parse potentially nested function arguments, the system employs Python's tokenize module via _split_top_level_calls. This tokenizer walks the character stream while tracking parentheses depth, yielding clean func(arg...) substrings even when arguments contain nested tuples, dictionaries, or quotes.
If tokenization fails due to malformed input, the _split_simple_calls_with_regex fallback uses _LENIENT_CALL_RE to recover sibling calls.
Each extracted string undergoes ast.parse(..., mode='eval') processing:
_extract_function_nameresolves dotted module paths (e.g.,mobile.open_app)_literal_from_astrecursively converts AST nodes into native Python literals, handling lists, tuples, dictionaries, and unary numbers
The parsed data populates a FunctionToolCall Pydantic model containing the function name, a dictionary of named and positional arguments, and the original source string.
from speech_to_speech.LLM.tool_call.function_call import (
extract_function_calls_from_text,
parse_function_call,
)
response = """
Sure, let me do that.
<code>mobile.open_app(app_name='drupe', timeout=5)</code>
Done!
"""
outside, calls = extract_function_calls_from_text(
response,
block_regex=r"<code>.*?</code>",
)
# calls[0] is a FunctionToolCall object
Validating and Converting Arguments
Before execution, arguments must conform to the original JSON schema to prevent type errors and injection attacks.
Safety Steps and Schema Enforcement
The FunctionToolCall.to_realtime_function_tool_call method in src/speech_to_speech/LLM/tool_call/function_call.py performs three validation stages:
- Positional argument stripping: Removes synthetic keys matching
^__arg_\d+__$that represent positional arguments, logging warnings for each removal - Undeclared parameter removal: When a tool schema is provided, strips any argument not listed in the schema's
propertiesobject - Required field enforcement: Compares the set of remaining keys against the schema's
requiredlist, raisingValueErrorif mandatory fields are missing
OpenAI Realtime Compatibility
After validation, the arguments are JSON-encoded and wrapped in a ResponseFunctionToolCall instance from openai.types.responses. This object includes a unique ID generated via _generate_id from src/speech_to_speech/utils/utils.py, making it compatible with the OpenAI Realtime API:
validated = calls[0].to_realtime_function_tool_call(function_tools=[mobile_open])
# Returns: ResponseFunctionToolCall(name='mobile.open_app', arguments='{"app_name":"drupe","timeout":5}', ...)
End-to-End Implementation Example
Integrating the full pipeline within a real-time server requires connecting the parsing logic to your execution context:
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.tool_call.function_call import extract_function_calls_from_text
# Define available tools
available_tools = [
FunctionTool(
type="function",
name="mobile.open_app",
description="Open an app",
parameters={
"type": "object",
"properties": {"app_name": {"type": "string"}},
"required": ["app_name"]
}
)
]
# Build system prompt once
system_prompt = build_tool_system_prompt(available_tools)
def handle_llm_output(text: str):
# Extract tool calls from model response
conversational_text, tool_calls = extract_function_calls_from_text(
text,
block_regex=r"<code>.*?</code>",
)
# Validate and convert to OpenAI format
realtime_calls = [
tc.to_realtime_function_tool_call(function_tools=available_tools)
for tc in tool_calls
]
# Execute or dispatch
for call in realtime_calls:
execute_tool(call) # Your custom dispatcher
send_to_client(call) # OpenAI Realtime websocket
Summary
- Tool definition: Use
FunctionToolwith JSON Schema parameters insrc/speech_to_speech/LLM/tool_call/function_tool.pyto declare callable functions - Prompt construction: Call
build_tool_system_promptfromsrc/speech_to_speech/LLM/tool_call/tool_prompt.pyto generate instructions with<code>delimiters - Safe parsing: Use
extract_function_calls_from_textinsrc/speech_to_speech/LLM/tool_call/function_call.pywhich tokenizes and AST-parses arguments to prevent code injection - Validation: Invoke
to_realtime_function_tool_callto strip positional args, filter undeclared parameters, and enforce required fields against the JSON schema - Integration: The resulting
ResponseFunctionToolCallobjects are compatible with OpenAI's Realtime API and ready for execution
Frequently Asked Questions
How does the system prevent the LLM from injecting malicious code?
The parser in src/speech_to_speech/LLM/tool_call/function_call.py uses Python's ast module with mode='eval' to parse function calls as expression syntax only, preventing statement execution. Additionally, the _literal_from_ast converter restricts allowed AST node types to literals (strings, numbers, lists, dicts) and rejects arbitrary code blocks, function definitions, or imports.
What happens if the model provides arguments not defined in the JSON schema?
When to_realtime_function_tool_call receives a tool schema, it automatically filters out any arguments not listed in the schema's properties dictionary, logging a warning for each removed key. This ensures that only declared parameters reach your execution layer, preventing interface breakage from hallucinated arguments.
Can the system handle complex nested arguments like lists of dictionaries?
Yes. The _literal_from_ast function in src/speech_to_speech/LLM/tool_call/function_call.py recursively traverses AST nodes, converting ast.List to Python lists and ast.Dict to dictionaries. The tokenization-based _split_top_level_calls correctly tracks parentheses depth to handle nested structures like func(items=[{"key": "value"}]) without breaking parsing.
How do I integrate this with my own LLM backend instead of OpenAI Realtime?
The architecture is backend-agnostic. Replace the OpenAI-specific ResponseFunctionToolCall export with your own response format, or simply use the FunctionToolCall objects directly after parsing. The core parsing logic (extract_function_calls_from_text and parse_function_call) works with any text-generating model that can follow the <code> delimiter instructions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →