Implementing Function Calling with LLMs for Tool Use: A Production-Ready Registry

Function calling with LLMs requires a robust tool registry that validates arguments against JSON schemas, executes functions in parallel, and returns structured error observations—enabling safe, multi-turn agent workflows.

Function calling transforms large language models from passive text generators into active agents that can interact with external systems. This guide examines a production-grade implementation from the rohitg00/ai-engineering-from-scratch repository, specifically phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py, which provides a stdlib-only ToolRegistry that bridges LLM reasoning with reliable tool execution.

Core Architecture: The Tool Interface

The foundation of reliable LLM tool use lies in explicit schema definitions that allow models to select appropriate functions. Each tool is encapsulated in a ToolDef object containing a name, human-readable description, and JSON Schema-based input_schema.

JSON Schema Definitions

In main.py, schema definitions capture required fields, primitive types, enums, and numeric limits. This uniform description lets the model understand tool capabilities before generation. The ToolDef dataclass stores these definitions alongside the executor function, creating a declarative interface that supports complex validation rules including type constraints and allowed values.

The ToolRegistry Class

The registry is implemented using only Python standard library components, exposing three public methods:

  • register() – Accepts a ToolDef and stores the executor, description, and schema
  • catalog() – Returns a JSON-serializable list of tool definitions ready for LLM consumption
  • dispatch_many() – Executes multiple tool calls in parallel while preserving correlation IDs

This lightweight design eliminates external dependencies while supporting the "Tool Interface → Function Calling → Execution → Observation" loop standard across OpenAI, Anthropic, and Gemini APIs.

Validation and Type Coercion

Before execution, incoming arguments undergo strict validation against their declared schemas. The private _coerce() helper in phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py performs safe type conversion—such as transforming the string "4" into integer 4—while enforcing enum constraints and numeric bounds.

When coercion fails, the system returns precise error messages (e.g., "validation error: missing required: a") rather than raising exceptions. This approach treats validation failures as observations that the LLM can read and correct in subsequent turns, maintaining conversation flow without breaking the agent loop.

Parallel Execution and Correlation

Modern LLMs support multiple tool calls within a single generation. The dispatch_many() method handles these parallel invocations by accepting a list of ToolCall objects, each carrying a unique tool_use_id that ties the request to its result.

The method returns a list of ToolResult objects containing the same correlation IDs, an ok boolean indicator, and string content. This design matches the parallel-call semantics described in docs/en.md, ensuring that results can be matched to their originating requests even when functions complete out of order.

Security and Sandboxing Best Practices

The implementation emphasizes per-tool sandbox policies rather than generic execution privileges. The lesson documentation specifically warns against implementing generic "run_shell" tools, recommending instead explicit verbs like git_status with narrowly constrained read/write surfaces, network access limits, and execution timeouts.

This granular approach minimizes attack surfaces while providing the LLM with specific, verifiable capabilities. Each tool should declare its resource requirements explicitly, allowing the agent framework to enforce least-privilege execution environments.

Alignment with BFCL V4 Standards

The registry design targets the Berkeley Function Calling Leaderboard V4 (BFCL V4) evaluation criteria, which tests agents across five axes: agentic reasoning, multi-turn dialogue, live function calling, non-live execution, and hallucination detection. The strict input_schema validation, parallel dispatch capabilities, and structured error reporting directly support these benchmarks, providing a solid foundation for tackling open problems in memory, dynamic decision-making, and long-horizon chaining.

Complete Implementation Example

The following runnable example demonstrates registration, parallel dispatch with intentional errors, and structured observation handling:

from dataclasses import dataclass
from typing import Any, Callable, Dict, List

# Simplified class definitions for clarity

@dataclass
class ToolDef:
    name: str
    description: str
    input_schema: Dict[str, Any]
    executor: Callable

@dataclass
class ToolCall:
    tool_use_id: str
    name: str
    arguments: Dict[str, Any]

@dataclass
class ToolResult:
    tool_use_id: str
    ok: bool
    content: str

class ToolRegistry:
    # Simplified implementation showing the interface

    def __init__(self):
        self._tools: Dict[str, ToolDef] = {}
    
    def register(self, tool: ToolDef) -> None:
        self._tools[tool.name] = tool
    
    def catalog(self) -> List[Dict]:
        return [{"name": t.name, "description": t.description, "input_schema": t.input_schema} 
                for t in self._tools.values()]
    
    def dispatch_many(self, calls: List[ToolCall]) -> List[ToolResult]:
        results = []
        for call in calls:
            if call.name not in self._tools:
                results.append(ToolResult(call.tool_use_id, False, f"unknown tool: {call.name}"))
                continue
            tool = self._tools[call.name]
            try:
                result = tool.executor(**call.arguments)
                results.append(ToolResult(call.tool_use_id, True, str(result)))
            except Exception as e:
                results.append(ToolResult(call.tool_use_id, False, f"execution error: {e}"))
        return results

# Register three example tools

reg = ToolRegistry()
reg.register(
    ToolDef(
        name="add",
        description="Add two integers a and b. Use for any integer addition.",
        input_schema={"type": "object",
                      "properties": {"a": {"type": "integer"},
                                     "b": {"type": "integer"}},
                      "required": ["a", "b"]},
        executor=lambda a, b: str(a + b),
    )
)
reg.register(
    ToolDef(
        name="multiply",
        description="Multiply two integers a and b. Prefer multiplication over looped addition.",
        input_schema={"type": "object",
                      "properties": {"a": {"type": "integer"},
                                     "b": {"type": "integer"}},
                      "required": ["a", "b"]},
        executor=lambda a, b: str(a * b),
    )
)
reg.register(
    ToolDef(
        name="classify",
        description="Classify a status as one of the allowed labels.",
        input_schema={"type": "object",
                      "properties": {"status": {"type": "string",
                                                "enum": ["open", "closed", "pending"]}},
                      "required": ["status"]},
        executor=lambda status: f"classified as {status}",
    )
)

# Build a parallel batch of tool calls (including intentional errors)

calls = [
    ToolCall("u01", "add", {"a": 2, "b": 3}),
    ToolCall("u02", "multiply", {"a": "4", "b": 5}),      # string will be coerced

    ToolCall("u03", "classify", {"status": "in_progress"}),  # enum violation

    ToolCall("u04", "classify", {"status": "open"}),
    ToolCall("u05", "subtract", {"a": 1, "b": 2}),       # unknown tool

]

# Dispatch all calls in parallel

results = reg.dispatch_many(calls)
for r in results:
    tag = "OK" if r.ok else "ERR"
    print(f"{r.tool_use_id} {tag}: {r.content}")

Running this snippet produces a catalog of available tools, then processes each call—returning validation errors as readable observations that the LLM can consume and retry.

Summary

  • ToolRegistry provides a stdlib-only foundation for LLM function calling with JSON Schema validation and parallel execution capabilities
  • The dispatch_many() method handles parallel tool calls while preserving correlation through tool_use_id fields
  • Input validation and type coercion occur via _coerce() logic before function execution, supporting enums and numeric bounds
  • Errors return as structured observation strings rather than exceptions, enabling LLM self-correction in multi-turn conversations
  • Sandboxing policies should restrict tool capabilities to specific verbs and explicit resource limits, avoiding generic shell execution
  • The implementation aligns with BFCL V4 evaluation standards for production agent systems

Frequently Asked Questions

What is the ToolRegistry in the ai-engineering-from-scratch repository?

The ToolRegistry is a lightweight Python class implemented in phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py that manages function registration, schema validation, and execution for LLM tool use. It provides the register(), catalog(), and dispatch_many() methods to bridge LLM outputs with actual function calls without requiring external dependencies beyond the Python standard library.

How does the registry handle invalid function arguments?

Instead of raising exceptions that would break the agent loop, the registry validates arguments against JSON schemas and returns structured error strings (e.g., "validation error: missing required: a") as observation content. This allows the LLM to receive feedback, reason about the failure, and retry with corrected parameters in the next turn, following the "error observations as structured strings" design principle documented in docs/en.md.

Can the registry execute multiple tools simultaneously?

Yes, the dispatch_many() method processes ToolCall objects in parallel, where each call includes a unique tool_use_id. The method returns a list of ToolResult objects containing the same correlation IDs, an ok status boolean, and string content. This enables the LLM to match specific results with their originating requests even when functions complete out of order, supporting the parallel tool call semantics found in modern LLM APIs.

The lesson documentation in phases/14-agent-engineering/06-tool-use-and-function-calling/docs/en.md emphasizes implementing per-tool sandbox policies that explicitly define read/write surfaces, network access, and execution timeouts. Avoid generic command execution tools; instead, expose specific verbs like git_status or read_file with narrowly defined parameters to minimize attack surfaces and maintain strict control over resource access.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →