How to Implement Custom Tool Definitions with the Local Transformers/MLX-LM LLM Backend
Wrap your Python functions in FunctionTool objects, pass them to the SpeechToSpeechPipeline via the tools parameter, and provide a tool_implementations mapping to enable local LLMs running on transformers or mlx-lm backends to call external capabilities.
The speech-to-speech library enables voice-driven interactions with local large language models. When using the transformers or mlx-lm backend, you can extend the LLM's capabilities by implementing custom tool definitions that allow the model to execute Python functions during conversations.
Understanding the Tool Architecture
The system treats external capabilities as function tools using a backend-agnostic architecture. At its core, the FunctionTool class (defined in src/speech_to_speech/LLM/tool_call/function_tool.py) stores your function's metadata—name, description, and JSON schema parameters—and renders it into a code-style prompt that the LLM can understand.
Key components orchestrate this process:
FunctionTool: Stores tool metadata and converts it to a Python function signaturesignature_from_schema: Translates JSON schemas into Python-style type hints (located insrc/speech_to_speech/LLM/tool_call/signature_from_schema.py)tool_prompt.py: Assembles the system prompt containing all tool definitions as code blockslanguage_model.py: Routes prompts to the local backend and parses tool calls from model outputss2s_pipeline.py: Orchestrates the end-to-end flow and injects tools into the LLM handler
Creating a Custom Function Tool
To define a custom tool, instantiate the FunctionTool class and populate its schema attributes. The following example creates a tool for generating matplotlib plots:
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool
def make_plot_tool() -> FunctionTool:
tool = FunctionTool()
tool.name = "draw_plot"
tool.description = "Draw a simple line plot with matplotlib."
tool.type = "function"
tool.parameters = {
"type": "object",
"properties": {
"x": {"type": "array", "items": {"type": "number"},
"description": "X coordinates"},
"y": {"type": "array", "items": {"type": "number"},
"description": "Y coordinates"},
"title": {"type": "string", "description": "Plot title"},
},
"required": ["x", "y"],
}
return tool
The FunctionTool.to_code_prompt() method automatically converts this schema into a Python function signature using signature_from_schema.py, generating a code block that appears in the system prompt.
Integrating Tools into the Pipeline
Register your tools with the SpeechToSpeechPipeline by passing them to the tools parameter. You must also provide a tool_implementations dictionary that maps tool names to their executable Python functions.
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from my_tool import make_plot_tool
import matplotlib.pyplot as plt
import io, base64
# Define the handler function
def draw_plot(x: list, y: list, title: str = "") -> str:
plt.figure()
plt.plot(x, y)
if title:
plt.title(title)
buf = io.BytesIO()
plt.savefig(buf, format="png")
plt.close()
return base64.b64encode(buf.getvalue()).decode()
# Initialize pipeline with local backend
pipeline = SpeechToSpeechPipeline(
llm_backend="mlx-lm", # or "transformers"
tools=[make_plot_tool()],
tool_implementations={"draw_plot": draw_plot},
# ... other configuration
)
The pipeline supports both transformers and mlx-lm backends interchangeably. The language_model.py handler selects the appropriate backend while using the same tool prompt format, ensuring your custom tool definitions work regardless of which local LLM engine you choose.
How Tool Calling Works
The tool execution flow follows four distinct stages managed by the pipeline:
-
Prompt Generation: The
build_tool_system_prompt()function insrc/speech_to_speech/LLM/tool_call/tool_prompt.pyconverts eachFunctionToolinto a Python code block. The system prompt instructs the model to emit at most one tool call per turn. -
Model Emission: When the local LLM decides to use a tool, it outputs a code block resembling:
draw_plot(x=[1, 2, 3], y=[4, 5, 6], title="Demo") -
Parsing: The
language_model.pywrapper extracts the function name and arguments from the model's response, looking up the corresponding implementation intool_implementations. -
Execution and Feedback: The pipeline executes the Python handler and returns the result (e.g., a base64-encoded image) as a tool result message, allowing the LLM to continue the conversation with the function output.
Summary
- Define tools using the
FunctionToolclass with JSON schema parameters insrc/speech_to_speech/LLM/tool_call/function_tool.py - Pass tool instances to
SpeechToSpeechPipelinevia thetoolsparameter - Provide executable handlers through the
tool_implementationsdictionary to map tool names to Python functions - The architecture works identically across both
transformersandmlx-lmbackends - Tool prompts are generated as Python code blocks that local models can parse consistently using the shared
LLMpackage
Frequently Asked Questions
Can I use the same tool definition for both transformers and mlx-lm backends?
Yes. The speech-to-speech library uses backend-agnostic prompt generation in tool_prompt.py. Whether you set llm_backend="transformers" or llm_backend="mlx-lm", the FunctionTool definitions and system prompts remain identical. The only difference is the underlying inference engine; the tool-calling interface is unified.
How does the model know when to call a tool versus respond normally?
The system prompt generated by build_tool_system_prompt() includes behavioral instructions that guide the model to emit at most one tool call per response. When the model determines a tool is needed, it outputs a Python code block that the language_model.py parser detects and extracts for execution.
What format should the tool implementation handler return?
Handlers should return strings that can be fed back into the conversation context. For text-based tools, return plain text results. For binary data like images, return base64-encoded strings. The language_model.py wrapper inserts this return value as a tool result message that the LLM receives in the next turn.
Where is the JSON schema converted to a Python signature?
The conversion happens in src/speech_to_speech/LLM/tool_call/signature_from_schema.py. This module translates your parameters schema into a Python-style function signature like (x: list, y: list, title: str = ""), which appears in the generated code prompt that the model sees.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →