Implementing Function Calling with LLMs for Tool Use: A Production-Ready Registry
Function calling with LLMs requires a robust tool registry that validates arguments against JSON schemas, executes functions in parallel, and returns structured error observations—enabling safe, multi-turn agent workflows.
Function calling transforms large language models from passive text generators into active agents that can interact with external systems. This guide examines a production-grade implementation from the rohitg00/ai-engineering-from-scratch repository, specifically phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py, which provides a stdlib-only ToolRegistry that bridges LLM reasoning with reliable tool execution.
Core Architecture: The Tool Interface
The foundation of reliable LLM tool use lies in explicit schema definitions that allow models to select appropriate functions. Each tool is encapsulated in a ToolDef object containing a name, human-readable description, and JSON Schema-based input_schema.
JSON Schema Definitions
In main.py, schema definitions capture required fields, primitive types, enums, and numeric limits. This uniform description lets the model understand tool capabilities before generation. The ToolDef dataclass stores these definitions alongside the executor function, creating a declarative interface that supports complex validation rules including type constraints and allowed values.
The ToolRegistry Class
The registry is implemented using only Python standard library components, exposing three public methods:
register()– Accepts aToolDefand stores the executor, description, and schemacatalog()– Returns a JSON-serializable list of tool definitions ready for LLM consumptiondispatch_many()– Executes multiple tool calls in parallel while preserving correlation IDs
This lightweight design eliminates external dependencies while supporting the "Tool Interface → Function Calling → Execution → Observation" loop standard across OpenAI, Anthropic, and Gemini APIs.
Validation and Type Coercion
Before execution, incoming arguments undergo strict validation against their declared schemas. The private _coerce() helper in phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py performs safe type conversion—such as transforming the string "4" into integer 4—while enforcing enum constraints and numeric bounds.
When coercion fails, the system returns precise error messages (e.g., "validation error: missing required: a") rather than raising exceptions. This approach treats validation failures as observations that the LLM can read and correct in subsequent turns, maintaining conversation flow without breaking the agent loop.
Parallel Execution and Correlation
Modern LLMs support multiple tool calls within a single generation. The dispatch_many() method handles these parallel invocations by accepting a list of ToolCall objects, each carrying a unique tool_use_id that ties the request to its result.
The method returns a list of ToolResult objects containing the same correlation IDs, an ok boolean indicator, and string content. This design matches the parallel-call semantics described in docs/en.md, ensuring that results can be matched to their originating requests even when functions complete out of order.
Security and Sandboxing Best Practices
The implementation emphasizes per-tool sandbox policies rather than generic execution privileges. The lesson documentation specifically warns against implementing generic "run_shell" tools, recommending instead explicit verbs like git_status with narrowly constrained read/write surfaces, network access limits, and execution timeouts.
This granular approach minimizes attack surfaces while providing the LLM with specific, verifiable capabilities. Each tool should declare its resource requirements explicitly, allowing the agent framework to enforce least-privilege execution environments.
Alignment with BFCL V4 Standards
The registry design targets the Berkeley Function Calling Leaderboard V4 (BFCL V4) evaluation criteria, which tests agents across five axes: agentic reasoning, multi-turn dialogue, live function calling, non-live execution, and hallucination detection. The strict input_schema validation, parallel dispatch capabilities, and structured error reporting directly support these benchmarks, providing a solid foundation for tackling open problems in memory, dynamic decision-making, and long-horizon chaining.
Complete Implementation Example
The following runnable example demonstrates registration, parallel dispatch with intentional errors, and structured observation handling:
from dataclasses import dataclass
from typing import Any, Callable, Dict, List
# Simplified class definitions for clarity
@dataclass
class ToolDef:
name: str
description: str
input_schema: Dict[str, Any]
executor: Callable
@dataclass
class ToolCall:
tool_use_id: str
name: str
arguments: Dict[str, Any]
@dataclass
class ToolResult:
tool_use_id: str
ok: bool
content: str
class ToolRegistry:
# Simplified implementation showing the interface
def __init__(self):
self._tools: Dict[str, ToolDef] = {}
def register(self, tool: ToolDef) -> None:
self._tools[tool.name] = tool
def catalog(self) -> List[Dict]:
return [{"name": t.name, "description": t.description, "input_schema": t.input_schema}
for t in self._tools.values()]
def dispatch_many(self, calls: List[ToolCall]) -> List[ToolResult]:
results = []
for call in calls:
if call.name not in self._tools:
results.append(ToolResult(call.tool_use_id, False, f"unknown tool: {call.name}"))
continue
tool = self._tools[call.name]
try:
result = tool.executor(**call.arguments)
results.append(ToolResult(call.tool_use_id, True, str(result)))
except Exception as e:
results.append(ToolResult(call.tool_use_id, False, f"execution error: {e}"))
return results
# Register three example tools
reg = ToolRegistry()
reg.register(
ToolDef(
name="add",
description="Add two integers a and b. Use for any integer addition.",
input_schema={"type": "object",
"properties": {"a": {"type": "integer"},
"b": {"type": "integer"}},
"required": ["a", "b"]},
executor=lambda a, b: str(a + b),
)
)
reg.register(
ToolDef(
name="multiply",
description="Multiply two integers a and b. Prefer multiplication over looped addition.",
input_schema={"type": "object",
"properties": {"a": {"type": "integer"},
"b": {"type": "integer"}},
"required": ["a", "b"]},
executor=lambda a, b: str(a * b),
)
)
reg.register(
ToolDef(
name="classify",
description="Classify a status as one of the allowed labels.",
input_schema={"type": "object",
"properties": {"status": {"type": "string",
"enum": ["open", "closed", "pending"]}},
"required": ["status"]},
executor=lambda status: f"classified as {status}",
)
)
# Build a parallel batch of tool calls (including intentional errors)
calls = [
ToolCall("u01", "add", {"a": 2, "b": 3}),
ToolCall("u02", "multiply", {"a": "4", "b": 5}), # string will be coerced
ToolCall("u03", "classify", {"status": "in_progress"}), # enum violation
ToolCall("u04", "classify", {"status": "open"}),
ToolCall("u05", "subtract", {"a": 1, "b": 2}), # unknown tool
]
# Dispatch all calls in parallel
results = reg.dispatch_many(calls)
for r in results:
tag = "OK" if r.ok else "ERR"
print(f"{r.tool_use_id} {tag}: {r.content}")
Running this snippet produces a catalog of available tools, then processes each call—returning validation errors as readable observations that the LLM can consume and retry.
Summary
- ToolRegistry provides a stdlib-only foundation for LLM function calling with JSON Schema validation and parallel execution capabilities
- The
dispatch_many()method handles parallel tool calls while preserving correlation throughtool_use_idfields - Input validation and type coercion occur via
_coerce()logic before function execution, supporting enums and numeric bounds - Errors return as structured observation strings rather than exceptions, enabling LLM self-correction in multi-turn conversations
- Sandboxing policies should restrict tool capabilities to specific verbs and explicit resource limits, avoiding generic shell execution
- The implementation aligns with BFCL V4 evaluation standards for production agent systems
Frequently Asked Questions
What is the ToolRegistry in the ai-engineering-from-scratch repository?
The ToolRegistry is a lightweight Python class implemented in phases/14-agent-engineering/06-tool-use-and-function-calling/code/main.py that manages function registration, schema validation, and execution for LLM tool use. It provides the register(), catalog(), and dispatch_many() methods to bridge LLM outputs with actual function calls without requiring external dependencies beyond the Python standard library.
How does the registry handle invalid function arguments?
Instead of raising exceptions that would break the agent loop, the registry validates arguments against JSON schemas and returns structured error strings (e.g., "validation error: missing required: a") as observation content. This allows the LLM to receive feedback, reason about the failure, and retry with corrected parameters in the next turn, following the "error observations as structured strings" design principle documented in docs/en.md.
Can the registry execute multiple tools simultaneously?
Yes, the dispatch_many() method processes ToolCall objects in parallel, where each call includes a unique tool_use_id. The method returns a list of ToolResult objects containing the same correlation IDs, an ok status boolean, and string content. This enables the LLM to match specific results with their originating requests even when functions complete out of order, supporting the parallel tool call semantics found in modern LLM APIs.
What security measures are recommended when implementing LLM tool use?
The lesson documentation in phases/14-agent-engineering/06-tool-use-and-function-calling/docs/en.md emphasizes implementing per-tool sandbox policies that explicitly define read/write surfaces, network access, and execution timeouts. Avoid generic command execution tools; instead, expose specific verbs like git_status or read_file with narrowly defined parameters to minimize attack surfaces and maintain strict control over resource access.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →