# Implementing Function Calling with LLMs for Tool Use: A Production-Ready Registry

> Learn to implement function calling with LLMs for efficient tool use. Build a production-ready registry for safe, multi-turn agent workflows by validating arguments, executing functions, and handling errors.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**Function calling with LLMs requires a robust tool registry that validates arguments against JSON schemas, executes functions in parallel, and returns structured error observations—enabling safe, multi-turn agent workflows.**

Function calling transforms large language models from passive text generators into active agents that can interact with external systems. This guide examines a production-grade implementation from the `rohitg00/ai-engineering-from-scratch` repository, specifically [`phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py), which provides a **stdlib-only** `ToolRegistry` that bridges LLM reasoning with reliable tool execution.

## Core Architecture: The Tool Interface

The foundation of reliable LLM tool use lies in explicit schema definitions that allow models to select appropriate functions. Each tool is encapsulated in a **ToolDef** object containing a name, human-readable description, and JSON Schema-based `input_schema`.

### JSON Schema Definitions

In [`main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/main.py), schema definitions capture required fields, primitive types, enums, and numeric limits. This uniform description lets the model understand tool capabilities before generation. The `ToolDef` dataclass stores these definitions alongside the executor function, creating a declarative interface that supports complex validation rules including type constraints and allowed values.

### The ToolRegistry Class

The registry is implemented using only Python standard library components, exposing three public methods:

- **`register()`** – Accepts a `ToolDef` and stores the executor, description, and schema
- **`catalog()`** – Returns a JSON-serializable list of tool definitions ready for LLM consumption
- **`dispatch_many()`** – Executes multiple tool calls in parallel while preserving correlation IDs

This lightweight design eliminates external dependencies while supporting the "Tool Interface → Function Calling → Execution → Observation" loop standard across OpenAI, Anthropic, and Gemini APIs.

## Validation and Type Coercion

Before execution, incoming arguments undergo strict validation against their declared schemas. The private `_coerce()` helper in [`phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py) performs safe type conversion—such as transforming the string `"4"` into integer `4`—while enforcing enum constraints and numeric bounds.

When coercion fails, the system returns precise error messages (e.g., `"validation error: missing required: a"`) rather than raising exceptions. This approach treats validation failures as **observations** that the LLM can read and correct in subsequent turns, maintaining conversation flow without breaking the agent loop.

## Parallel Execution and Correlation

Modern LLMs support multiple tool calls within a single generation. The `dispatch_many()` method handles these parallel invocations by accepting a list of **ToolCall** objects, each carrying a unique `tool_use_id` that ties the request to its result.

The method returns a list of **ToolResult** objects containing the same correlation IDs, an `ok` boolean indicator, and string `content`. This design matches the parallel-call semantics described in [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md), ensuring that results can be matched to their originating requests even when functions complete out of order.

## Security and Sandboxing Best Practices

The implementation emphasizes per-tool sandbox policies rather than generic execution privileges. The lesson documentation specifically warns against implementing generic "run_shell" tools, recommending instead explicit verbs like `git_status` with narrowly constrained read/write surfaces, network access limits, and execution timeouts.

This granular approach minimizes attack surfaces while providing the LLM with specific, verifiable capabilities. Each tool should declare its resource requirements explicitly, allowing the agent framework to enforce least-privilege execution environments.

## Alignment with BFCL V4 Standards

The registry design targets the **Berkeley Function Calling Leaderboard V4** (BFCL V4) evaluation criteria, which tests agents across five axes: agentic reasoning, multi-turn dialogue, live function calling, non-live execution, and hallucination detection. The strict `input_schema` validation, parallel dispatch capabilities, and structured error reporting directly support these benchmarks, providing a solid foundation for tackling open problems in memory, dynamic decision-making, and long-horizon chaining.

## Complete Implementation Example

The following runnable example demonstrates registration, parallel dispatch with intentional errors, and structured observation handling:

```python
from dataclasses import dataclass
from typing import Any, Callable, Dict, List

# Simplified class definitions for clarity

@dataclass
class ToolDef:
    name: str
    description: str
    input_schema: Dict[str, Any]
    executor: Callable

@dataclass
class ToolCall:
    tool_use_id: str
    name: str
    arguments: Dict[str, Any]

@dataclass
class ToolResult:
    tool_use_id: str
    ok: bool
    content: str

class ToolRegistry:
    # Simplified implementation showing the interface

    def __init__(self):
        self._tools: Dict[str, ToolDef] = {}
    
    def register(self, tool: ToolDef) -> None:
        self._tools[tool.name] = tool
    
    def catalog(self) -> List[Dict]:
        return [{"name": t.name, "description": t.description, "input_schema": t.input_schema} 
                for t in self._tools.values()]
    
    def dispatch_many(self, calls: List[ToolCall]) -> List[ToolResult]:
        results = []
        for call in calls:
            if call.name not in self._tools:
                results.append(ToolResult(call.tool_use_id, False, f"unknown tool: {call.name}"))
                continue
            tool = self._tools[call.name]
            try:
                result = tool.executor(**call.arguments)
                results.append(ToolResult(call.tool_use_id, True, str(result)))
            except Exception as e:
                results.append(ToolResult(call.tool_use_id, False, f"execution error: {e}"))
        return results

# Register three example tools

reg = ToolRegistry()
reg.register(
    ToolDef(
        name="add",
        description="Add two integers a and b. Use for any integer addition.",
        input_schema={"type": "object",
                      "properties": {"a": {"type": "integer"},
                                     "b": {"type": "integer"}},
                      "required": ["a", "b"]},
        executor=lambda a, b: str(a + b),
    )
)
reg.register(
    ToolDef(
        name="multiply",
        description="Multiply two integers a and b. Prefer multiplication over looped addition.",
        input_schema={"type": "object",
                      "properties": {"a": {"type": "integer"},
                                     "b": {"type": "integer"}},
                      "required": ["a", "b"]},
        executor=lambda a, b: str(a * b),
    )
)
reg.register(
    ToolDef(
        name="classify",
        description="Classify a status as one of the allowed labels.",
        input_schema={"type": "object",
                      "properties": {"status": {"type": "string",
                                                "enum": ["open", "closed", "pending"]}},
                      "required": ["status"]},
        executor=lambda status: f"classified as {status}",
    )
)

# Build a parallel batch of tool calls (including intentional errors)

calls = [
    ToolCall("u01", "add", {"a": 2, "b": 3}),
    ToolCall("u02", "multiply", {"a": "4", "b": 5}),      # string will be coerced

    ToolCall("u03", "classify", {"status": "in_progress"}),  # enum violation

    ToolCall("u04", "classify", {"status": "open"}),
    ToolCall("u05", "subtract", {"a": 1, "b": 2}),       # unknown tool

]

# Dispatch all calls in parallel

results = reg.dispatch_many(calls)
for r in results:
    tag = "OK" if r.ok else "ERR"
    print(f"{r.tool_use_id} {tag}: {r.content}")

```

Running this snippet produces a catalog of available tools, then processes each call—returning validation errors as readable observations that the LLM can consume and retry.

## Summary

- **ToolRegistry** provides a stdlib-only foundation for LLM function calling with JSON Schema validation and parallel execution capabilities
- The `dispatch_many()` method handles parallel tool calls while preserving correlation through `tool_use_id` fields
- Input validation and type coercion occur via `_coerce()` logic before function execution, supporting enums and numeric bounds
- Errors return as structured observation strings rather than exceptions, enabling LLM self-correction in multi-turn conversations
- Sandboxing policies should restrict tool capabilities to specific verbs and explicit resource limits, avoiding generic shell execution
- The implementation aligns with **BFCL V4** evaluation standards for production agent systems

## Frequently Asked Questions

### What is the ToolRegistry in the ai-engineering-from-scratch repository?

The **ToolRegistry** is a lightweight Python class implemented in [`phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py) that manages function registration, schema validation, and execution for LLM tool use. It provides the `register()`, `catalog()`, and `dispatch_many()` methods to bridge LLM outputs with actual function calls without requiring external dependencies beyond the Python standard library.

### How does the registry handle invalid function arguments?

Instead of raising exceptions that would break the agent loop, the registry validates arguments against JSON schemas and returns structured error strings (e.g., `"validation error: missing required: a"`) as observation content. This allows the LLM to receive feedback, reason about the failure, and retry with corrected parameters in the next turn, following the "error observations as structured strings" design principle documented in [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md).

### Can the registry execute multiple tools simultaneously?

Yes, the `dispatch_many()` method processes **ToolCall** objects in parallel, where each call includes a unique `tool_use_id`. The method returns a list of **ToolResult** objects containing the same correlation IDs, an `ok` status boolean, and string content. This enables the LLM to match specific results with their originating requests even when functions complete out of order, supporting the parallel tool call semantics found in modern LLM APIs.

### What security measures are recommended when implementing LLM tool use?

The lesson documentation in [`phases/14-agent-engineering/06-tool-use-and-function-calling/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/06-tool-use-and-function-calling/docs/en.md) emphasizes implementing per-tool sandbox policies that explicitly define read/write surfaces, network access, and execution timeouts. Avoid generic command execution tools; instead, expose specific verbs like `git_status` or `read_file` with narrowly defined parameters to minimize attack surfaces and maintain strict control over resource access.