# How to Integrate Unsloth's Tool Calling and Code Execution Features into AI Models

> Learn how to integrate Unsloth's tool calling and code execution features into AI models. Execute Python code, web search, and terminal commands in a secure sandbox.

- Repository: [Unsloth AI/unsloth](https://github.com/unslothai/unsloth)
- Tags: how-to-guide
- Published: 2026-03-20

---

**TLDR:** Unsloth provides a built-in framework that allows LLMs to execute external tools—such as web search, Python code, and terminal commands—through a sandboxed environment with automatic result integration back into the conversation flow.

Unsloth equips open-source LLMs with native tool-calling capabilities through a self-contained backend architecture. By integrating features from the `unslothai/unsloth` repository, developers can enable AI models to perform autonomous research, run safe code execution, and interact with system shells without requiring external orchestration services.

## Understanding the Three-Layer Architecture

Unsloth implements tool calling across three distinct layers that handle definition, configuration, and execution.

### Tool Definition and Sandboxed Execution

The core tool logic resides in [`studio/backend/core/inference/tools.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/core/inference/tools.py), which declares three supported tools: `web_search`, `python`, and `terminal`. This module provides the `execute_tool` dispatcher function that routes calls to appropriate executors while enforcing safety constraints.

The sandbox creates a **per-session working directory** (`_workdirs`) for each chat session, isolating file system operations. For Python execution, the system runs `_check_code_safety` to perform static analysis that blocks signal-timer abuse and dangerous exception handling before code runs. Terminal commands pass through a hard-coded filter that blocks destructive operations like `rm` and `sudo`.

### Model-Side Configuration Flags

The OpenAI-compatible inference schema in [`studio/backend/models/inference.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/models/inference.py) exposes Unsloth-specific extension fields that control tool availability:

- `supports_tools`: Boolean indicating model capability
- `enable_tools`: Activation flag (set `true` to enable, `null` for auto-detection)
- `enabled_tools`: Optional whitelist (e.g., `["web_search", "python"]`)
- `auto_heal_tool_calls`: Automatic repair for malformed tool calls
- `max_tool_calls_per_message`: Loop iteration limit (default prevents infinite loops)
- `tool_call_timeout`: Maximum execution seconds per tool

When `enable_tools` is active, the backend signals availability to the model through the chat template system.

### The Agentic Inference Loop

The generation routine `generate_chat_completion_with_tools` in [`studio/backend/core/inference/llama_cpp.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/core/inference/llama_cpp.py) implements the agentic loop that coordinates between the LLM and tool executors.

The loop first requests a model response. If the output contains a `tool_calls` JSON array (OpenAI format) or XML-style `<tool_call>` blocks, the parser extracts the tool name and arguments. The backend then invokes `execute_tool`, streams a status message to the client, and appends the truncated result (limited to 8,000 characters) to the conversation history. The loop repeats up to `max_tool_calls_per_message` iterations, allowing the model to reason over fresh data iteratively.

## Enabling Tool Calling in Your Application

Integration requires configuring the request payload and ensuring your model supports the capability.

### Step 1: Configure the Request Payload

Send a chat completion request with Unsloth extension flags to activate the tool-calling pipeline:

```python
import httpx

payload = {
    "model": "unsloth-llama-3.1-8B",
    "messages": [
        {"role": "user", "content": "What is the latest news on quantum computing?"},
    ],
    "stream": False,
    "enable_tools": True,
    "enabled_tools": ["web_search"],
    "max_tool_calls_per_message": 5,
    "tool_call_timeout": 120,
}

resp = httpx.post("http://localhost:8000/v1/chat/completions", json=payload)
print(resp.json())

```

The `ChatCompletionRequest` schema processes these fields in [`studio/backend/models/inference.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/models/inference.py), automatically detecting tool support when `enable_tools` is `null`.

### Step 2: Model Prompt Integration

Unsloth's Ollama-style chat templates in [`unsloth/ollama_template_mappers.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/ollama_template_mappers.py) inject system prompts describing available tools when `supports_tools` evaluates to true. Model-specific implementations in [`unsloth/models/vision.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/vision.py) and [`unsloth/models/llama.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/llama.py) include the `[x-unsloth] Enable tool calling` flag in their UI representations, ensuring the model understands its capability to request external actions.

### Step 3: Handling Tool Call Responses

When the model decides to use a tool, it returns a structured response:

```json
{
  "role": "assistant",
  "content": "",
  "tool_calls": [
    {
      "id": "call_1",
      "type": "function",
      "function": {
        "name": "web_search",
        "arguments": {"query": "quantum computing breakthrough 2024"}
      }
    }
  ]
}

```

The backend detects this format, executes the corresponding tool, and feeds the DuckDuckGo search results or execution output back into the next generation context as a `tool_end` event.

## Safe Code Execution and Security Controls

Unsloth implements multiple security layers for code and terminal execution to prevent system abuse.

### Python Sandbox Safety

Before executing Python code, the `execute_tool` function runs `_check_code_safety` to analyze the abstract syntax tree for dangerous patterns, specifically detecting signal-timer manipulation and unsafe exception handling tricks that could break sandbox isolation.

```python
code = """
import math
print("sqrt(2) =", math.sqrt(2))
"""

result = execute_tool(
    name="python",
    arguments={"code": code},
    timeout=30,
    session_id="chat-42"
)

```

Each execution occurs within a temporary directory tied to the specific `session_id`, preventing cross-session file system contamination.

### Terminal Command Restrictions

The terminal tool maintains a blocklist of destructive commands including `rm`, `sudo`, and other system-modifying operations. This restriction ensures that shell access remains read-only or bounded to safe operations only.

## Summary

- **Three-layer architecture**: Tool definitions in [`tools.py`](https://github.com/unslothai/unsloth/blob/main/tools.py), configuration flags in [`inference.py`](https://github.com/unslothai/unsloth/blob/main/inference.py), and the agentic loop in [`llama_cpp.py`](https://github.com/unslothai/unsloth/blob/main/llama_cpp.py) work together to enable autonomous tool use.
- **Sandboxed execution**: Python and terminal tools run in isolated per-session directories with static analysis safety checks and command blocklists.
- **Flexible configuration**: Control tool availability through `enable_tools`, `enabled_tools`, and timeout parameters in the chat completion request.
- **Self-contained pipeline**: No external services required beyond optional DuckDuckGo integration for web search functionality.
- **Automatic healing**: The `auto_heal_tool_calls` flag attempts to repair malformed tool call formats automatically during the inference loop.

## Frequently Asked Questions

### How does Unsloth handle malformed tool calls from the model?

The backend includes an `auto_heal_tool_calls` configuration option that attempts to parse and repair malformed tool call structures during the `generate_chat_completion_with_tools` execution loop. If enabled, the system tries to extract tool names and arguments from improperly formatted JSON or XML blocks before rejecting the request.

### What safety measures prevent malicious code execution in the Python tool?

Unsloth employs a multi-layered approach: static analysis via `_check_code_safety` detects signal-timer abuse and dangerous exception patterns, execution occurs within a temporary per-session working directory (`_workdirs`), and a configurable timeout kills long-running processes. These measures isolate potentially harmful code without requiring external sandbox services.

### Can I restrict which tools are available to specific models or conversations?

Yes. The `enabled_tools` parameter accepts an array of tool names (e.g., `["web_search"]` or `["python", "terminal"]`) that acts as a whitelist for the current request. Combined with the `supports_tools` model flag and `enable_tools` boolean, you can granularly control tool availability per conversation or deployment.

### Which file contains the main logic for executing web search queries?

The `web_search` tool implementation resides in [`studio/backend/core/inference/tools.py`](https://github.com/unslothai/unsloth/blob/main/studio/backend/core/inference/tools.py), which uses the `ddgs` library to execute DuckDuckGo queries. This file also contains the `execute_tool` dispatcher function that routes all tool invocations to their appropriate handlers.