# How to Implement Custom Tool Definitions with the Local Transformers/MLX-LM LLM Backend

> Learn to implement custom tool definitions with the local transformers/mlx-lm LLM backend. Wrap Python functions in FunctionTool and enable local LLMs to call external capabilities.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: tutorial
- Published: 2026-07-10

---

**Wrap your Python functions in `FunctionTool` objects, pass them to the `SpeechToSpeechPipeline` via the `tools` parameter, and provide a `tool_implementations` mapping to enable local LLMs running on `transformers` or `mlx-lm` backends to call external capabilities.**

The `speech-to-speech` library enables voice-driven interactions with local large language models. When using the **`transformers`** or **`mlx-lm`** backend, you can extend the LLM's capabilities by implementing custom tool definitions that allow the model to execute Python functions during conversations.

## Understanding the Tool Architecture

The system treats external capabilities as **function tools** using a backend-agnostic architecture. At its core, the `FunctionTool` class (defined in [`src/speech_to_speech/LLM/tool_call/function_tool.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/function_tool.py)) stores your function's metadata—name, description, and JSON schema parameters—and renders it into a code-style prompt that the LLM can understand.

Key components orchestrate this process:

- **`FunctionTool`**: Stores tool metadata and converts it to a Python function signature
- **`signature_from_schema`**: Translates JSON schemas into Python-style type hints (located in [`src/speech_to_speech/LLM/tool_call/signature_from_schema.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/signature_from_schema.py))
- **[`tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/tool_prompt.py)**: Assembles the system prompt containing all tool definitions as code blocks
- **[`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py)**: Routes prompts to the local backend and parses tool calls from model outputs
- **[`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)**: Orchestrates the end-to-end flow and injects tools into the LLM handler

## Creating a Custom Function Tool

To define a custom tool, instantiate the `FunctionTool` class and populate its schema attributes. The following example creates a tool for generating matplotlib plots:

```python
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool

def make_plot_tool() -> FunctionTool:
    tool = FunctionTool()
    tool.name = "draw_plot"
    tool.description = "Draw a simple line plot with matplotlib."
    tool.type = "function"
    tool.parameters = {
        "type": "object",
        "properties": {
            "x": {"type": "array", "items": {"type": "number"},
                  "description": "X coordinates"},
            "y": {"type": "array", "items": {"type": "number"},
                  "description": "Y coordinates"},
            "title": {"type": "string", "description": "Plot title"},
        },
        "required": ["x", "y"],
    }
    return tool

```

The `FunctionTool.to_code_prompt()` method automatically converts this schema into a Python function signature using [`signature_from_schema.py`](https://github.com/huggingface/speech-to-speech/blob/main/signature_from_schema.py), generating a code block that appears in the system prompt.

## Integrating Tools into the Pipeline

Register your tools with the `SpeechToSpeechPipeline` by passing them to the `tools` parameter. You must also provide a `tool_implementations` dictionary that maps tool names to their executable Python functions.

```python
from speech_to_speech.s2s_pipeline import SpeechToSpeechPipeline
from my_tool import make_plot_tool
import matplotlib.pyplot as plt
import io, base64

# Define the handler function

def draw_plot(x: list, y: list, title: str = "") -> str:
    plt.figure()
    plt.plot(x, y)
    if title:
        plt.title(title)
    buf = io.BytesIO()
    plt.savefig(buf, format="png")
    plt.close()
    return base64.b64encode(buf.getvalue()).decode()

# Initialize pipeline with local backend

pipeline = SpeechToSpeechPipeline(
    llm_backend="mlx-lm",  # or "transformers"

    tools=[make_plot_tool()],
    tool_implementations={"draw_plot": draw_plot},
    # ... other configuration

)

```

The pipeline supports both `transformers` and `mlx-lm` backends interchangeably. The [`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py) handler selects the appropriate backend while using the same tool prompt format, ensuring your custom tool definitions work regardless of which local LLM engine you choose.

## How Tool Calling Works

The tool execution flow follows four distinct stages managed by the pipeline:

1. **Prompt Generation**: The `build_tool_system_prompt()` function in [`src/speech_to_speech/LLM/tool_call/tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/tool_prompt.py) converts each `FunctionTool` into a Python code block. The system prompt instructs the model to emit at most one tool call per turn.

2. **Model Emission**: When the local LLM decides to use a tool, it outputs a code block resembling:
   ```python
   draw_plot(x=[1, 2, 3], y=[4, 5, 6], title="Demo")
   ```

3. **Parsing**: The [`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py) wrapper extracts the function name and arguments from the model's response, looking up the corresponding implementation in `tool_implementations`.

4. **Execution and Feedback**: The pipeline executes the Python handler and returns the result (e.g., a base64-encoded image) as a tool result message, allowing the LLM to continue the conversation with the function output.

## Summary

- Define tools using the `FunctionTool` class with JSON schema parameters in [`src/speech_to_speech/LLM/tool_call/function_tool.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/function_tool.py)
- Pass tool instances to `SpeechToSpeechPipeline` via the `tools` parameter
- Provide executable handlers through the `tool_implementations` dictionary to map tool names to Python functions
- The architecture works identically across both `transformers` and `mlx-lm` backends
- Tool prompts are generated as Python code blocks that local models can parse consistently using the shared `LLM` package

## Frequently Asked Questions

### Can I use the same tool definition for both transformers and mlx-lm backends?

Yes. The `speech-to-speech` library uses backend-agnostic prompt generation in [`tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/tool_prompt.py). Whether you set `llm_backend="transformers"` or `llm_backend="mlx-lm"`, the `FunctionTool` definitions and system prompts remain identical. The only difference is the underlying inference engine; the tool-calling interface is unified.

### How does the model know when to call a tool versus respond normally?

The system prompt generated by `build_tool_system_prompt()` includes behavioral instructions that guide the model to emit at most one tool call per response. When the model determines a tool is needed, it outputs a Python code block that the [`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py) parser detects and extracts for execution.

### What format should the tool implementation handler return?

Handlers should return strings that can be fed back into the conversation context. For text-based tools, return plain text results. For binary data like images, return base64-encoded strings. The [`language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/language_model.py) wrapper inserts this return value as a tool result message that the LLM receives in the next turn.

### Where is the JSON schema converted to a Python signature?

The conversion happens in [`src/speech_to_speech/LLM/tool_call/signature_from_schema.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/signature_from_schema.py). This module translates your `parameters` schema into a Python-style function signature like `(x: list, y: list, title: str = "")`, which appears in the generated code prompt that the model sees.