How to Integrate Unsloth's Tool Calling and Code Execution Features into AI Models

TLDR: Unsloth provides a built-in framework that allows LLMs to execute external tools—such as web search, Python code, and terminal commands—through a sandboxed environment with automatic result integration back into the conversation flow.

Unsloth equips open-source LLMs with native tool-calling capabilities through a self-contained backend architecture. By integrating features from the unslothai/unsloth repository, developers can enable AI models to perform autonomous research, run safe code execution, and interact with system shells without requiring external orchestration services.

Understanding the Three-Layer Architecture

Unsloth implements tool calling across three distinct layers that handle definition, configuration, and execution.

Tool Definition and Sandboxed Execution

The core tool logic resides in studio/backend/core/inference/tools.py, which declares three supported tools: web_search, python, and terminal. This module provides the execute_tool dispatcher function that routes calls to appropriate executors while enforcing safety constraints.

The sandbox creates a per-session working directory (_workdirs) for each chat session, isolating file system operations. For Python execution, the system runs _check_code_safety to perform static analysis that blocks signal-timer abuse and dangerous exception handling before code runs. Terminal commands pass through a hard-coded filter that blocks destructive operations like rm and sudo.

Model-Side Configuration Flags

The OpenAI-compatible inference schema in studio/backend/models/inference.py exposes Unsloth-specific extension fields that control tool availability:

  • supports_tools: Boolean indicating model capability
  • enable_tools: Activation flag (set true to enable, null for auto-detection)
  • enabled_tools: Optional whitelist (e.g., ["web_search", "python"])
  • auto_heal_tool_calls: Automatic repair for malformed tool calls
  • max_tool_calls_per_message: Loop iteration limit (default prevents infinite loops)
  • tool_call_timeout: Maximum execution seconds per tool

When enable_tools is active, the backend signals availability to the model through the chat template system.

The Agentic Inference Loop

The generation routine generate_chat_completion_with_tools in studio/backend/core/inference/llama_cpp.py implements the agentic loop that coordinates between the LLM and tool executors.

The loop first requests a model response. If the output contains a tool_calls JSON array (OpenAI format) or XML-style <tool_call> blocks, the parser extracts the tool name and arguments. The backend then invokes execute_tool, streams a status message to the client, and appends the truncated result (limited to 8,000 characters) to the conversation history. The loop repeats up to max_tool_calls_per_message iterations, allowing the model to reason over fresh data iteratively.

Enabling Tool Calling in Your Application

Integration requires configuring the request payload and ensuring your model supports the capability.

Step 1: Configure the Request Payload

Send a chat completion request with Unsloth extension flags to activate the tool-calling pipeline:

import httpx

payload = {
    "model": "unsloth-llama-3.1-8B",
    "messages": [
        {"role": "user", "content": "What is the latest news on quantum computing?"},
    ],
    "stream": False,
    "enable_tools": True,
    "enabled_tools": ["web_search"],
    "max_tool_calls_per_message": 5,
    "tool_call_timeout": 120,
}

resp = httpx.post("http://localhost:8000/v1/chat/completions", json=payload)
print(resp.json())

The ChatCompletionRequest schema processes these fields in studio/backend/models/inference.py, automatically detecting tool support when enable_tools is null.

Step 2: Model Prompt Integration

Unsloth's Ollama-style chat templates in unsloth/ollama_template_mappers.py inject system prompts describing available tools when supports_tools evaluates to true. Model-specific implementations in unsloth/models/vision.py and unsloth/models/llama.py include the [x-unsloth] Enable tool calling flag in their UI representations, ensuring the model understands its capability to request external actions.

Step 3: Handling Tool Call Responses

When the model decides to use a tool, it returns a structured response:

{
  "role": "assistant",
  "content": "",
  "tool_calls": [
    {
      "id": "call_1",
      "type": "function",
      "function": {
        "name": "web_search",
        "arguments": {"query": "quantum computing breakthrough 2024"}
      }
    }
  ]
}

The backend detects this format, executes the corresponding tool, and feeds the DuckDuckGo search results or execution output back into the next generation context as a tool_end event.

Safe Code Execution and Security Controls

Unsloth implements multiple security layers for code and terminal execution to prevent system abuse.

Python Sandbox Safety

Before executing Python code, the execute_tool function runs _check_code_safety to analyze the abstract syntax tree for dangerous patterns, specifically detecting signal-timer manipulation and unsafe exception handling tricks that could break sandbox isolation.

code = """
import math
print("sqrt(2) =", math.sqrt(2))
"""

result = execute_tool(
    name="python",
    arguments={"code": code},
    timeout=30,
    session_id="chat-42"
)

Each execution occurs within a temporary directory tied to the specific session_id, preventing cross-session file system contamination.

Terminal Command Restrictions

The terminal tool maintains a blocklist of destructive commands including rm, sudo, and other system-modifying operations. This restriction ensures that shell access remains read-only or bounded to safe operations only.

Summary

  • Three-layer architecture: Tool definitions in tools.py, configuration flags in inference.py, and the agentic loop in llama_cpp.py work together to enable autonomous tool use.
  • Sandboxed execution: Python and terminal tools run in isolated per-session directories with static analysis safety checks and command blocklists.
  • Flexible configuration: Control tool availability through enable_tools, enabled_tools, and timeout parameters in the chat completion request.
  • Self-contained pipeline: No external services required beyond optional DuckDuckGo integration for web search functionality.
  • Automatic healing: The auto_heal_tool_calls flag attempts to repair malformed tool call formats automatically during the inference loop.

Frequently Asked Questions

How does Unsloth handle malformed tool calls from the model?

The backend includes an auto_heal_tool_calls configuration option that attempts to parse and repair malformed tool call structures during the generate_chat_completion_with_tools execution loop. If enabled, the system tries to extract tool names and arguments from improperly formatted JSON or XML blocks before rejecting the request.

What safety measures prevent malicious code execution in the Python tool?

Unsloth employs a multi-layered approach: static analysis via _check_code_safety detects signal-timer abuse and dangerous exception patterns, execution occurs within a temporary per-session working directory (_workdirs), and a configurable timeout kills long-running processes. These measures isolate potentially harmful code without requiring external sandbox services.

Can I restrict which tools are available to specific models or conversations?

Yes. The enabled_tools parameter accepts an array of tool names (e.g., ["web_search"] or ["python", "terminal"]) that acts as a whitelist for the current request. Combined with the supports_tools model flag and enable_tools boolean, you can granularly control tool availability per conversation or deployment.

Which file contains the main logic for executing web search queries?

The web_search tool implementation resides in studio/backend/core/inference/tools.py, which uses the ddgs library to execute DuckDuckGo queries. This file also contains the execute_tool dispatcher function that routes all tool invocations to their appropriate handlers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →